Video compression sensing restoration method and system based on generative network and low-rank regularization

By generating a network and a low-rank regularized video compressed sensing restoration method, using a low-rank latent tensor and a space-frequency attention module, combined with inter-frame patch matching, the problem of low video clarity in the existing technology is solved and efficient video restoration is achieved.

CN116645280BActive Publication Date: 2025-09-09SEVNCE ROBOTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310485586.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2025-09-09
Estimated Expiration
2043-04-28

AI Technical Summary

Technical Problem

The existing generative networks based on deep image priors have limited feature capabilities in video compressed sensing restoration and fail to fully exploit the temporal correlation between video frame sequences, resulting in low clarity and poor visual effects of the restored video.

Method used

A video compressed sensing restoration method based on generative network and low-rank regularization is adopted. The low-rank latent tensor of the generative network is used for initial recovery and enhanced recovery. The spatial-frequency attention module is used to improve the generative network structure. The inter-frame patch matching and low-rank approximation technology are combined to fully utilize the spatiotemporal correlation of video sequences.

Benefits of technology

It achieves video restoration with high definition and good visual effects, improves the restoration performance of non-key frames, and avoids blocking effects without the need for training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116645280B_ABST
    Figure CN116645280B_ABST
Patent Text Reader

Abstract

The present invention provides a video compression sensing restoration method and system based on a generative network and low-rank regularization. The method obtains an image group and restores an image sequence by: the generative network uses a low-rank latent tensor for initial iteration to generate an initial restored image sequence; the generative network includes N downsampling modules, N upsampling modules, and N space-frequency attention modules; similar patch position information of each patch in a non-key frame is obtained; the generative network uses a low-rank latent tensor for enhanced iteration to enhance and restore the non-key frame, and in each enhanced iteration, the enhanced iteration loss is obtained by using the similar patch position information of the patch in the non-key frame; the initial restored image of the key frame and the enhanced restored image of the non-key frame are combined to obtain a restored image sequence. The key frame is restored through initial iteration using spatiotemporal correlation, and the non-key frame is enhanced and restored using inter-frame temporal correlation. The restored image does not have a block effect and is clear with good visual effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video restoration technology, and in particular to a video compression sensing restoration method and system based on a generative network and low-rank regularization. Background Art

[0002] Video signals typically exhibit significant spatiotemporal redundancy, making them well-suited for data acquisition and recovery using compressed sensing techniques. Current image compressed sensing methods primarily exploit spatial correlation within images and can be used to recover video sequences frame by frame. However, this approach often overlooks temporal correlation within video sequences. Some traditional video compressed sensing methods, such as KTSLR, design three-dimensional sparse prior models to simultaneously recover multiple consecutive frames. Other methods, such as MC / ME (Reference 1), Video-MH (Reference 2), and RRS (Reference 3), first perform initial frame-by-frame recovery and then improve performance through motion estimation and compensation. These methods effectively exploit the spatiotemporal correlation within video sequences, but their hand-crafted models and manually selected parameters often limit their performance. Deep learning-based methods often achieve superior performance. For example, VCSNet (Reference 4) achieves state-of-the-art performance by jointly training a measurement network and a recovery network in an end-to-end manner. However, deep learning-based methods typically require large training datasets, which can be challenging in some practical applications.

[0003] The deep image prior (DIP) shows that the structure of an untrained generative network itself can be used as a priori information for image restoration, providing a new approach to image restoration that requires neither training data nor too many manually designed parameters. In the existing field of image compression sensing, good restoration performance has been achieved by combining DIP with additional regularization. DIP bridges the gap between methods that require training and those that do not, and shows great potential in video compression sensing restoration. Although there are existing technologies that attempt to extend DIP for video sequence restoration, the existing generative networks based on deep image priors have limited feature extraction capabilities, and the regularization used does not deeply explore the temporal correlation between video frame sequences, resulting in low clarity and poor visual effects in the restored video.

[0004] Literature Description:

[0005] Document 1: S.Mun, JEFowler.Residual reconstruction for block-basedcompressed sensing of video[C], Proceedings of the IEEE Data CompressionConference(DCC).Snowbird, UT: IEEE, 2011: 183-192.

[0006] Document 2: EW Tramel, JEFowler. Video compressed sensing with multihypothesis[C], Proceedings of the IEEE Data Compression Conference (DCC). Snowbird, UT: IEEE, 2011: 193-202.

[0007] Document 3: C. Zhao, S. Ma, J. Zhang, et al.Video compressive sensingreconstruction via reweighted residual sparsity[J]. IEEE Transactions onCircuitsand Systems for Video Technology, 2017, 27(6): 1182-1195.

[0008] Document 4: W.Shi, S.Liu, F.Jiang, et al. Video compressed sensing using aconvolutional neural network [J]. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 31(2): 425-438. Summary of the Invention

[0009] The present invention aims to at least solve the technical problems existing in the prior art and provide a video compression sensing recovery method and system based on a generative network and low-rank regularization.

[0010] In order to achieve the above-mentioned purpose of the present invention, according to the first aspect of the present invention, the present invention provides a video compression sensing restoration method based on a generative network and low-rank regularization, comprising: obtaining compressed sensing data of a video, wherein the video is divided into a plurality of image groups, each image group comprising a key frame and a plurality of non-key frames; obtaining a restored image sequence of the image groups one by one based on the compressed sensing data of the video, and the process of obtaining the restored image sequence of a single image group comprises: the generative network uses a low-rank latent tensor to generate an initial restored image sequence of the image group based on the initial iteration of the compressed sensing data of the video; the generative network comprises N downsampling modules and N upsampling modules connected in sequence, and N Space-frequency attention module; N downsampling modules and N upsampling modules correspond to each other one by one and are connected by jumpers, and a space-frequency attention module jumper is provided on the jumper; N is a positive integer; based on the initial restored image sequence, inter-frame patch matching is performed to obtain similar patch position information of each patch in the non-key frame; the generative network continues to use the low-rank latent tensor for enhanced iteration to enhance and restore the non-key frames in the initial restored image sequence. In each enhanced iteration, the enhanced iteration loss is obtained by the similar patch position information of the patches in the non-key frames; the key frames of the initial restored image sequence and the enhanced restored non-key frames are combined to obtain the restored image sequence of the image group.

[0011] To achieve the above-mentioned object of the present invention, according to a second aspect of the present invention, a video restoration device is provided, comprising: a data acquisition module, configured to acquire compressed sensing data of a video, wherein the video is divided into a plurality of image groups, each image group including a key frame and a plurality of non-key frames; and a restoration module, configured to acquire restored image sequences of the image groups one by one based on the compressed sensing data of the video, wherein the process of acquiring the restored image sequence of a single image group comprises:

[0012] The generative network uses a low-rank latent tensor to generate an initial restored image sequence of an image group based on the compressed sensing data of the video through initial iteration; the generative network includes N downsampling modules and N upsampling modules connected in sequence, and N space-frequency attention modules; the N downsampling modules and the N upsampling modules correspond to each other one by one and are connected by jumpers, and a space-frequency attention module is provided on the jumper for jump connection; N is a positive integer; based on the initial restored image sequence, inter-frame patch matching is performed to obtain similar patch position information of each patch in the non-key frame; the generative network continues to use the low-rank latent tensor for enhanced iteration to enhance and restore the non-key frames in the initial restored image sequence to obtain enhanced restored images of the non-key frames, and in each enhanced iteration, the enhanced iteration loss is obtained through the similar patch position information of the patches in the non-key frames; the initial restored images of the key frames of the image group and the enhanced restored images of the non-key frames in the image group are combined to obtain a restored image sequence corresponding to the image group.

[0013] In order to achieve the above-mentioned purpose of the present invention, according to the third aspect of the present invention, the present invention provides a video compression perception recovery system, including: a video compression device, used to compress each frame image of the video in blocks to obtain video compression perception data; a video recovery device as described in the first aspect of the present invention; the video compression device and the video recovery device are connected and communicated.

[0014] The technical effect of the present application is as follows: video restoration is achieved through initial restoration and enhanced restoration. Both restoration processes are restored using a generative network through a low-rank latent tensor, without the need for training data training, and the low-rank latent tensor is used to realize the overall video sequence recovery of the image group based on the deep image prior, without blocking effects; in the initial restoration process, the space-frequency attention module is used to improve the generative network structure to improve the restoration performance, and the spatiotemporal correlation of the video sequence is fully utilized. Before entering the enhanced restoration, each patch of the non-key frame is similarly matched, and the matching result is applied to the loss calculation of each iteration of the enhanced restoration, and the temporal redundancy of the video sequence is used to enhance the performance. In the enhanced restoration process, the inter-frame low-rank approximation is further used to realize frame-by-frame non-local low-rank regularization, thereby improving the restoration performance of non-key frames. The restored video obtained by the present application has the advantages of high clarity and good visual effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 1 is a flow chart of a video compression sensing restoration method based on a generative network and low-rank regularization in Example 1 of the present invention;

[0016] Figure 2 Schematic diagram of the framework of the video restoration method proposed in Example 1 of the present invention;

[0017] Figure 3 2 is a schematic diagram of the structure of the space-frequency attention module of embodiment 1 of the present invention;

[0018] Figure 4 is a system block diagram of a video recovery device in Example 2 of the present invention;

[0019] Figure 5 This is a system block diagram of the video compression sensing restoration system in Example 3 of the present invention. DETAILED DESCRIPTION

[0020] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0021] In the description of the present invention, it should be understood that the terms "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention.

[0022] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal communication between two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.

[0023] Example 1

[0024] This embodiment discloses a video compression sensing recovery method based on a generative network and low-rank regularization, and its flow chart is as follows: Figure 1 Shown, including:

[0025] Step S1: Obtaining compressed sensing data for a video. The video is divided into multiple image groups. Preferably, the original image sequence of the video is divided into image groups in chronological order, with each image group including a key frame and multiple non-key frames. The key frame is preferably, but not limited to, the first frame or the last frame. Key frames and non-key frames have different measurement rates during compressed sensing. Image frames measured at a first measurement rate are key frames, and image frames measured at a second measurement rate are non-key frames. The first measurement rate is higher than the second measurement rate.

[0026] In order to reduce the computational complexity and make the measurement value acquisition process hardware-friendly, a block-based frame-by-frame compressed sensing method is used to obtain the measurement value of the original image sequence of the video. Each video frame is first divided into B×B blocks without overlap, and then each block is vectorized in a raster scan manner and the measurement matrix is ​​used. The measurement is performed, j represents the frame index, and the measurement rates of key frames and non-key frames are different, resulting in different measurement matrices for the two. The process of obtaining measurement values ​​based on block-by-frame compressed sensing can be expressed as: in, represents the i-th block in the j-th frame of the vectorization, express The corresponding measurement value. The measurement rate is defined as the ratio of the measurement value dimension to the signal dimension, that is, M / B 2 .

[0027] Step S2, based on the compressed sensing data of the video, the restored image sequence of the image groups is obtained one by one, and one image group is restored at a time. Each image group uses Figure 2 The framework shown is used for recovery. The framework proposed in this application is named LRR-VCSNet (Video Compressed Sensing Reconstruction Network with Low-Rank Regularization).

[0028] like Figure 2 As shown, the process of obtaining a restored image sequence for a single image group includes:

[0029] In step S21, the generative network uses a low-rank latent tensor to generate an initial restored image sequence for the image group based on the compressed sensing data of the video through initial iterations. The generative network includes N downsampling modules and N upsampling modules connected in sequence, as well as N space-frequency attention modules. The N downsampling modules and the N upsampling modules correspond to each other one-to-one and are connected by jumpers, and the jumpers are provided with a space-frequency attention module jumper. N is a positive integer, preferably, but not limited to, 5. Since the restoration process of each image group requires two key frames, the restoration of each image group requires the use of key frames of the preceding, following, or following adjacent image groups. When the key frame is the first frame of the image group, the key frame of the following image group is used. When the key frame is the last frame of the image group, the key frame of the previous image group is used.

[0030] like Figure 2 As shown, construct a low-rank hidden tensor Z LR :When initializing the low-rank hidden tensor, first generate a random hidden variable z1 with a resolution equal to the size of a video frame, and then expand it into three-dimensional space to form a low-rank hidden tensor Z LR The low-rank hidden tensor can be expressed as follows:

[0031]

[0032] Among them, Q represents the number of frames in the image group. Since the key frames of adjacent image groups are needed, the low-rank latent tensor Z LR It has Q+1 latent variables, H×W represents the resolution of the image frame. Different from the random latent tensor, the proposed low-rank latent tensor z LR It has a low-rank property and can achieve global low-rank regularization for the entire image group in the latent space. Therefore, in step S21, the generation network uses the spatiotemporal correlation of the video sequence to perform initial restoration on Q+1 frames to obtain Q+1 initial restored images.

[0033] In this embodiment, preferably, the structure of the generated network is as follows Figure 2As shown, a three-dimensional generation network is used. When N is equal to 5, the first downsampling module, the second downsampling module, the third downsampling module, the fourth downsampling module, the fifth downsampling module, the fifth upsampling module, the fourth upsampling module, the third upsampling module, the second upsampling module, and the first upsampling module in the generation network are connected in sequence. The corresponding upsampling modules and downsampling modules are connected by jumpers, and a space-frequency attention module is provided on the jumper, such as the first downsampling module corresponds to the first upsampling module. Figure 2 As shown, each downsampling module includes a first encoding 3D convolution (convolution kernel size is 3×3×3), a first normalization layer, a first Leaky ReLU activation layer, a second encoding 3D convolution (convolution kernel size is 3×3×3), a second batch normalization layer, and a second Leaky ReLU activation layer, which are connected in sequence. Each upsampling module includes a first decoding batch normalization layer, a first decoding 3D convolution (convolution kernel size is 3×3×3), a second decoding batch normalization layer, a third Leaky ReLU activation layer, a second decoding 3D convolution (convolution kernel size is 1×1×1), a third decoding batch normalization layer, a fourth Leaky ReLU activation layer, and an upsampling layer, which are connected in sequence. Preferably, a convolution layer with a kernel size of 1×1×1 and a sigmoid activation layer are also connected to the output end of the first upsampling module.

[0034] In this embodiment, in order to better utilize the generative network to explore the deep image prior, preferably, the spatial frequency attention module (SFA) is used to optimize the skip connection of the generative network, which has a non-local receptive field and can integrate multi-scale information. Figure 3 As shown in the figure, the channel is divided into a local branch and a global branch. The local branch uses the original convolution to capture local information, and the global branch uses the fast Fourier transform (FFT) to capture global information in the spectral domain. Then, information with different receptive fields is exchanged and fused internally. Therefore, the spatial frequency attention (SFA) module can optimize the feature extraction ability of the generated network.

[0035] More preferably, Figure 3As shown, the input features of the space-frequency attention module include global processing features and local processing features; the space-frequency attention module includes a global average pooling layer, a first fully connected layer, and a first activation layer (preferably a ReLU activation function) connected in sequence, as well as a global branch and a local branch respectively connected to the output end of the first activation layer; the global branch includes a second fully connected layer, a second activation layer (preferably a Sigmoid activation function), a first multiplication unit, a first three-dimensional convolution batch normalization activation module (3D Conv-BN-ReLU), Real-2D-FFT module (two-dimensional fast Fourier transform for real-valued signals), a second three-dimensional convolution batch normalization activation module, Inv-Real-2D-FFT module (two-dimensional inverse fast Fourier transform for real-valued signals), a first addition unit, a first three-dimensional convolution layer, a second addition unit and a first batch normalization activation module (BN-ReLU), the global branch also includes a second three-dimensional convolution layer connected to the output end of the first multiplication unit; the local branch includes a third fully connected layer, a third activation layer (preferably a Sigmoid activation function), a second multiplication unit, a third three-dimensional convolution layer, a third addition unit and a second batch normalization activation module connected in sequence, the local branch It also includes a fourth three-dimensional convolution layer connected to the output end of the second multiplication unit; the first multiplication unit is used to perform element-wise multiplication processing on the global processing features and the features output by the second activation layer; the first addition unit is used to perform element-wise addition processing on the output features of the Inv-Real-2D-FFT module and the output features of the first three-dimensional convolution batch normalization activation module; the second addition unit is used to perform element-wise addition processing on the output features of the fourth three-dimensional convolution layer and the output features of the first three-dimensional convolution layer; the second multiplication unit is used to perform element-wise multiplication processing on the local processing features and the features output by the third activation layer; and the third addition unit is used to perform element-wise addition processing on the output features of the second three-dimensional convolution layer and the output features of the third three-dimensional convolution layer.

[0036] Further preferably, in step S21, the generation network uses the low-rank latent tensor to generate an initial restored image sequence of the image group based on the compressed sensing data of the video by initial iteration, specifically comprising:

[0037] The generative network iteratively generates an initial restored image sequence using a low-rank latent tensor until an initial iteration stopping condition is reached. In each initial iteration, a measurement fidelity loss is calculated, and the network parameters of the low-rank latent tensor and the generative network are updated with the optimization objective of minimizing the measurement fidelity loss. The initial iteration stopping condition is preferably, but not limited to, when the number of initial iterations reaches a preset first iteration threshold.

[0038] The initial recovery process can be formulated as the following optimization problem:

[0039]

[0040] L LS Represents the measurement fidelity loss, and the measurement fidelity loss calculation formula is:

[0041]

[0042] Where Q represents the number of frames in the image group; N b Indicates the number of blocks that each frame of the video is divided into in block-based compressed sensing; Φ j Represents the block measurement matrix of the j-th frame image; represents the measurement value of the i-th block of the j-th frame image; represents the j-th frame image output by the generated network, and represents the j-th initial restored image in the initial iteration; P i The (·) function represents the operator that extracts and vectorizes the i-th block in the image frame. Represents the restored image group, and the generated network can be regarded as the original image group Parameterization x = g ω (Z LR ). Using the hidden tensor Z LR , the generative network maps the network parameters ω to the original image set.

[0043] In step S22, inter-frame patch matching is performed based on the initial restored image sequence to obtain similar patch location information for each patch in non-key frames. A patch represents a block, which is different from the block used in image compression sensing. Its size is S×S, and each non-key frame can be divided into multiple patches.

[0044] In step S22, inter-frame patch matching is performed based on the initial restored image to obtain groups of similar patches. Similar patches have similar structures, so the similar patch matrix has a low-rank characteristic, and low-rank approximation can be used to remove noise and improve the enhanced restoration effect.

[0045] Further preferably, step S22 includes:

[0046] Step S221, the initial restored image sequence of the image group and the initial restored images of the key frames of the adjacent image groups of the image group are combined into a matching sequence, and the matching sequence has a total of Q+1 frames of images;

[0047] Step S222, selecting adjacent frames of non-key frames in the matching sequence as reference frames of non-key frames. Specifically, including: the first frame and the last frame of the matching sequence are both key frames, the non-key frame close to the first frame uses the previous non-key frame as the reference frame, the non-key frame close to the last frame uses the next non-key frame as the reference frame, the non-key frame adjacent to the first frame uses the initial restored image of the first frame as the reference frame, and the non-key frame adjacent to the last frame uses the initial restored image of the last frame as the reference frame. If the number of frames between the non-key frame and the first frame (i.e., the interval distance) is less than the number of frames between the non-key frame and the last frame (i.e., the interval distance), the non-key frame is considered to be close to the first frame; otherwise, the key frame is considered to be close to the last frame. Since the temporal correlation decreases as the distance between frames increases, in practical applications, it is preferred to first divide the non-key frames into two groups in order according to the distance between them and the key frames. In the group closer to the backward key frames, each non-key frame is matched patch by patch with its backward adjacent frame as a reference; in the group closer to the forward key frames, each non-key frame is matched patch by patch with its forward adjacent frame as a reference.

[0048] Step S223: Divide the non-key frame into multiple patches. For each patch in the non-key frame, obtain k similar patches in the reference frame. Record the location information of each patch and its similar patches, where k is a positive integer. Sequentially obtain similar patches for all patches in each non-key frame. Preferably, the step of obtaining k similar patches in the reference frame for each patch in the non-key frame includes:

[0049] Step S2231: Obtain a matching patch in the reference frame using a sliding window. Slide the window in the reference frame, using the image within each sliding window as the matching patch. The sliding window size is equal to S×S, and the search window size is larger than S×S.

[0050] Step S2232: Calculate the similarity between the fth patch on the non-key frame and the matching patch according to the following formula:

[0051] Where S represents the side length of the patch; Represents the j-th frame image output from the generative network Extract and vectorize the operator of the fth patch of size S×S, at this time the jth frame image It is actually the initial restored image; express The reference frame of P w (·) represents the operator that obtains and vectorizes matching patches in a sliding window manner in the search window.

[0052] Step S2233: Sort the matching patches in descending order of similarity, and select the first k matching patches as similar patches.

[0053] The position information of each patch and its corresponding similar patch on the non-key frame is recorded as

[0054] In step S23, the generative network continues to use the low-rank latent tensor to perform enhancement iterations to enhance and restore the non-key frames in the initial restored image sequence, obtaining enhanced restored images of the non-key frames. In each enhancement iteration, the enhancement iteration loss is obtained through the similar patch position information of the patch in the non-key frame.

[0055] Specifically, in step S23, in each enhancement iteration, execute:

[0056] Step S231: Generate an enhanced restored image sequence of the image group through the generation network. In particular, during the first enhancement iteration, the initial restored image sequence of the image group is used as input and step S232 is directly executed by skipping step S231.

[0057] Step S232: vectorize each patch and its similar patches in the non-key frame and form a similar patch matrix for each patch. Use the recorded position information of each patch and its similar patches to generate the similar patch matrix for the patch:

[0058]

[0059] in, To extract and vectorize the position information into C j,f The patch operator is

[0060] In step S233, singular value decomposition is used to perform low-rank approximation on the similar patch matrix of each patch to obtain a low-rank approximation of each patch. The similar patch matrix can be decomposed into:

[0061]

[0062] Among them, σ i′ represents the i′th singular value of matrix T; u i′ represents the i′th left singular vector of the matrix T; v i′ represents the i′th right singular vector of matrix T, v i′ T Indicates v i′ The transpose of ; R represents the rank of the matrix T, and rank(T) means to find the rank of the matrix T.

[0063] According to the Eckart-Young-Mirsky theorem, only the largest r singular values ​​are retained, and the optimal rank r approximation matrix is:

[0064]

[0065] Thus, the low-rank approximation You can start from Obtained in.

[0066] In step S234, the non-local low-rank regularization loss is calculated according to the following formula:

[0067]

[0068] Among them, N p Indicates the number of patches in each frame; Q indicates the number of frames in the image group; It represents the fth patch in the jth frame image of the generated network output. At this time, it is in the enhanced restoration stage, and the generated network output image is the enhanced restoration image; express A low-rank approximation of .

[0069] Step S235, calculate the enhanced iteration loss L = L LS +λL R ,λ represents the regularization parameter; L LS Represents the loss of measurement fidelity.

[0070] Step S236, update the network parameters of the low-rank latent tensor and the generator network with the optimization goal of minimizing the enhancement iteration loss. The enhancement recovery process can be expressed as the following optimization problem:

[0071]

[0072] Step S24 : combining the initial restored images of the key frames of the image group and the enhanced restored images of the non-key frames in the image group to obtain a restored image sequence corresponding to the image group.

[0073] It can be seen that in the enhanced restoration process, the temporal redundancy of the video sequence is further utilized through inter-frame low-rank approximation to enhance the restoration effect of non-key frames.

[0074] The image restoration framework provided in this embodiment was experimentally verified, and the LRR-VCSNet proposed in this application was experimented based on the PyTorch framework. Using a Gaussian random matrix, the measurement value is obtained in a block-based frame-by-frame compressed sensing measurement method, and the block size is set to 32×32. The number of input channels of the latent tensor is set to 16, and additive Gaussian noise with zero mean and standard deviation of 0.03 is used for perturbation after each iteration. In the space-frequency attention module SFA, the proportion of the global branch in the input channel and the output channel is 0.5. The Adam optimizer is used to optimize the network parameters, and the learning rate is set to 0.001. The size of each image group is set to 8, that is, one key frame and seven non-key frames constitute an image group. The measurement rate of the key frame is set to 0.5, and the measurement rates of the non-key frames are set to 0.01, 0.05 and 0.1 respectively. The patch size is set to 16×16, and the search box padding size is set to 20. The regularization parameter of the enhanced recovery process is set to 0.01. The stopping criterion is set based on the maximum number of iterations. The number of iterations of the initial restoration process is set to 10,000, while the number of iterations of the enhanced restoration process is set to 500, 1000, and 1500, corresponding to the cases of non-keyframe measurement rates of 0.01, 0.05, and 0.1, respectively. Six standard CIF format video sequences are used to test the performance, namely Akiyo, Bus, Coastguard, Flower, Mother_daughter, and Paris.

[0075] The LRR-VCSNet of this application is compared with the current mainstream video compression sensing method. The methods used for comparison include MC / ME (Document 1), Video-MH (Document 2), RRS (Document 3) and VCSNet (Document 4), among which VCSNet is a method based on deep learning, which realizes the joint training of the measurement network and the recovery network in an end-to-end manner, and has achieved the current optimal effect (the VCSNet here refers to the dual key frame model in Document 4). The block size is set to 32×32, the image group size is set to 8, and the key frame measurement rate is set to 0.5. Since the Video-MH method is designed for one key frame and one non-key frame, in order to make the comparison fair, it is additionally designed here so that the non-key frame in an image group can refer to the backward and forward key frames at the same time.

[0076] Tables 1 and 2 compare the average PSNR and SSIM of different video compressed sensing methods on the first two image groups of six CIF video sequences, respectively, for non-keyframe measurement rates of 0.01, 0.05, and 0.1. The best results are marked in bold, and the suboptimal results are marked in bold italics. As can be seen, the proposed LRR-VCSNet achieves the best overall average PSNR and SSIM. Compared with MC / ME, Video-MH, RRS, and VCSNet, the proposed LRR-VCSNet achieves gains of approximately 4.00dB, 2.96dB, 5.36dB, and 0.24dB in overall average PSNR, and gains of approximately 0.0985, 0.0597, 0.1328, and 0.0083 in overall average SSIM. Compared with VCSNet using large training data, the proposed LRR-VCSNet achieves competitive performance. Specifically, compared to VCSNet, the proposed LRR-VCSNet achieved average PNSR gains of approximately 0.52dB and 0.87dB, and average SSIM gains of approximately 0.0282 and 0.0303, respectively, at non-keyframe measurement rates of 0.05 and 0.1. In terms of non-keyframe recovery, LRR-VCSNet surpassed traditional video compressed sensing methods and achieved comparable results to VCSNet using large training data. However, compared to other methods, the proposed LRR-VCSNet achieved clearer visual effects and recovered more texture details.

[0077] Table 1 Comparison of the average PSNR (unit: dB) of different video compression sensing methods on the first two image groups of six CIF video sequences

[0078]

[0079]

[0080] Table 2 Comparison of average SSIM of different video compression sensing methods on the first two image groups of six CIF video sequences

[0081]

[0082]

[0083] Example 2

[0084] This embodiment discloses a video recovery device, such as Figure 4 Shown, including:

[0085] A data acquisition module is configured to acquire compressed sensing data of a video, where the video is divided into multiple image groups, each of which includes a key frame and multiple non-key frames. The data acquisition module is preferably, but not limited to, a communication module for receiving compressed sensing data, or a data acquisition module (such as a digital communication interface) for reading the stored video compressed sensing data from a data storage module.

[0086] The restoration module obtains restored image sequences of image groups one by one based on compressed sensing data of the video. The process of obtaining the restored image sequence of a single image group includes: a generative network uses a low-rank latent tensor to initially iteratively generate the initial restored image sequence of the image group based on the compressed sensing data of the video; the generative network includes N downsampling modules and N upsampling modules connected in sequence, and N space-frequency attention modules; the N downsampling modules and the N upsampling modules correspond to each other and are connected by jumpers, and the jumpers are provided with a space-frequency attention module jumper; N is a positive integer; based on the initial restored image sequence, inter-frame patch matching is performed to obtain similar patch position information of each patch in non-key frames; the generative network continues to use the low-rank latent tensor to perform enhanced iterations to enhance and restore non-key frames in the initial restored image sequence, and in each enhanced iteration, the enhanced iteration loss is obtained based on the similar patch position information of patches in non-key frames; and the key frames of the initial restored image sequence and the enhanced and restored non-key frames are combined to obtain the restored image sequence of the image group. The restoration module is preferably, but not limited to, a microprocessor or a computer host.

[0087] In this embodiment, the specific steps executed by the recovery module can refer to Example 1 and will not be repeated here.

[0088] Example 3

[0089] This embodiment discloses a video compression sensing recovery system. Figure 5 As shown, the apparatus includes: a video compression device for compressing each frame of a video in blocks to obtain video compression perception data; and the video recovery device provided in Example 2; the video compression device and the video recovery device are connected and communicated. The video compression device and the video recovery device can be remotely located and connected and communicated with each other via wired or wireless means.

[0090] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0091] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

Claims

1. A video compression sensing restoration method based on a generative network and low-rank regularization, characterized in that: include: Obtaining compressed sensing data of a video, wherein the video is divided into a plurality of image groups, each image group including a key frame and a plurality of non-key frames; The restored image sequence of the image groups is obtained one by one based on the compressed sensing data of the video. The process of obtaining the restored image sequence of a single image group includes: The generative network uses a low-rank latent tensor to generate an initial restored image sequence of an image group based on compressed sensing data of the video through initial iteration; the generative network includes N downsampling modules and N upsampling modules connected in sequence, and N space-frequency attention modules; the N downsampling modules and the N upsampling modules correspond to each other one-to-one and are connected by jumpers, and a space-frequency attention module jumper is provided on the jumper; N is a positive integer; Based on the initial restored image sequence, inter-frame patch matching is performed to obtain the similar patch position information of each patch in the non-key frame; The generative network continues to use the low-rank latent tensor for enhancement iteration to enhance the non-key frames in the initial restored image sequence and obtain the enhanced restored images of the non-key frames. In each enhancement iteration, the enhancement iteration loss is obtained through the similar patch position information of the patch in the non-key frame. The initial restored images of the key frames of the image group and the enhanced restored images of the non-key frames in the image group are combined to obtain a restored image sequence corresponding to the image group.

2. The video compression sensing restoration method based on generative network and low-rank regularization according to claim 1, characterized in that The input features of the space-frequency attention module include global processing features and local processing features; The spatial-frequency attention module includes a global average pooling layer, a first fully connected layer, and a first activation layer connected in sequence, as well as a global branch and a local branch respectively connected to the output of the first activation layer; The global branch includes a second fully connected layer, a second activation layer, a first multiplication unit, a first three-dimensional convolution batch normalization activation module, a Real-2D-FFT module, a second three-dimensional convolution batch normalization activation module, an Inv-Real-2D-FFT module, a first addition unit, a first three-dimensional convolution layer, a second addition unit and a first batch normalization activation module, which are connected in sequence. The global branch also includes a second three-dimensional convolution layer connected to the output end of the first multiplication unit; The local branch includes a third fully connected layer, a third activation layer, a second multiplication unit, a third three-dimensional convolutional layer, a third addition unit, and a second batch normalization activation module connected in sequence, and the local branch also includes a fourth three-dimensional convolutional layer connected to the output end of the second multiplication unit; The first multiplication unit is used to perform element-wise multiplication of the global processing feature and the feature output by the second activation layer; The first adding unit is used to perform element-wise addition processing on the output features of the Inv-Real-2D-FFT module and the output features of the first three-dimensional convolution batch normalization activation module; The second adding unit is used to perform element-wise addition processing on the output features of the fourth three-dimensional convolutional layer and the output features of the first three-dimensional convolutional layer; The second multiplication unit is used to perform element-wise multiplication of the local processing features and the features output by the third activation layer; The third adding unit is used to perform element-wise addition processing on the output features of the second three-dimensional convolutional layer and the output features of the third three-dimensional convolutional layer.

3. The video compression sensing restoration method based on generative network and low-rank regularization according to claim 1, characterized in that: The generative network uses a low-rank latent tensor to generate an initial restored image sequence of the image group based on the initial iteration of the compressed sensing data of the video, specifically including: The generative network uses the low-rank hidden tensor to iteratively generate the initial restored image sequence until the initial iteration stopping condition is reached. In each initial iteration, the measurement fidelity loss is calculated, and the network parameters of the low-rank hidden tensor and the generative network are updated with the minimization of the measurement fidelity loss as the optimization goal.

4. The video compression sensing restoration method based on a generative network and low-rank regularization according to any one of claims 1 to 3, characterized in that: The measurement fidelity loss calculation formula is: Where Q represents the number of frames in the image group; N b Indicates the number of blocks that each frame of the video is divided into in block-based compressed sensing; Φ j Represents the block measurement matrix of the j-th frame image; represents the measurement value of the i-th block of the j-th frame image; represents the j-th frame image output by the generated network; P i The (·) function represents the operator that extracts and vectorizes the i-th block in the image frame.

5. The video compression sensing restoration method based on a generative network and low-rank regularization according to any one of claims 1 to 3, characterized in that: Based on the initial restored image sequence, inter-frame patch matching is performed to obtain similar patch position information of each patch in the non-key frame, specifically including: Combining an initial restored image sequence of an image group and initial restored images of key frames of an adjacent image group of the image group into a matching sequence; Selecting adjacent frames of non-key frames in the matching sequence as reference frames of non-key frames; Multiple patches are divided on non-key frames. Each patch of the non-key frame obtains k similar patches in the reference frame, and the position information of each patch and its similar patches is recorded, where k is a positive integer.

6. The video compression sensing restoration method based on generative network and low-rank regularization according to claim 5, characterized in that: Select adjacent frames of non-key frames in the matching sequence as reference frames of non-key frames, including: The first and last frames of the matching sequence are both key frames. The non-key frame close to the first frame uses the previous non-key frame as the reference frame, and the non-key frame close to the last frame uses the next non-key frame as the reference frame. The non-key frame adjacent to the first frame uses the initial restored image of the first frame as the reference frame, and the non-key frame adjacent to the last frame uses the initial restored image of the last frame as the reference frame.

7. The video compression sensing restoration method based on generative network and low-rank regularization according to claim 5, characterized in that: Each patch of a non-key frame obtains k similar patches in the reference frame, including: Obtain matching patches in the reference frame using a sliding window approach; The similarity between the fth patch and the matching patch on the non-key frame is calculated according to the following formula: Where S represents the side length of the patch; Represents the j-th frame image output from the generative network An operator that extracts and vectorizes the fth patch of size S×S from ; express The reference frame of P w (·) represents the operator that obtains and vectorizes matching patches in a sliding window manner in the search window; Sort the matching patches in descending order of similarity, and select the top k matching patches as similar patches.

8. The video compression sensing restoration method based on a generative network and low-rank regularization according to claim 1, 2, 3, 6 or 7, characterized in that: In each enhancement iteration, perform: The generative network generates an enhanced restored image sequence of the image group; Vectorize each patch and its similar patches in non-key frames and form a similar patch matrix for each patch; Use singular value decomposition to perform low-rank approximation on the similar patch matrix of each patch to obtain a low-rank approximation of each patch; The non-local low-rank regularization loss is obtained according to the following formula: Among them, N p Indicates the number of patches in each frame; Q indicates the number of frames in the image group; Represents the fth patch in the jth frame image of the generated network output; express A low-rank approximation of ; Obtain enhanced iterative loss L = L LS +λL R ,λ represents the regularization parameter; L LS represents the loss of measurement fidelity; The network parameters of the low-rank hidden tensor and the generator network are updated with the optimization objective of minimizing the enhanced iterative loss.

9. A video restoration device, characterized in that: include: A data acquisition module is used to acquire compressed sensing data of a video, where the video is divided into a plurality of image groups, each image group including a key frame and a plurality of non-key frames; The restoration module obtains restored image sequences of image groups one by one based on the compressed sensing data of the video. The process of obtaining the restored image sequence of a single image group includes: The generative network uses a low-rank latent tensor to generate an initial restored image sequence of an image group based on compressed sensing data of the video through initial iteration; the generative network includes N downsampling modules and N upsampling modules connected in sequence, and N space-frequency attention modules; the N downsampling modules and the N upsampling modules correspond to each other one-to-one and are connected by jumpers, and a space-frequency attention module jumper is provided on the jumper; N is a positive integer; Based on the initial restored image sequence, inter-frame patch matching is performed to obtain the similar patch position information of each patch in the non-key frame; The generative network continues to use the low-rank latent tensor for enhancement iteration to enhance the non-key frames in the initial restored image sequence and obtain the enhanced restored images of the non-key frames. In each enhancement iteration, the enhancement iteration loss is obtained through the similar patch position information of the patch in the non-key frame. The initial restored images of the key frames of the image group and the enhanced restored images of the non-key frames in the image group are combined to obtain a restored image sequence corresponding to the image group.

10. A video compression sensing recovery system, characterized in that: include: Video compression equipment, used for compressing each frame of the video into blocks to obtain video compression perception data; The video restoration device according to claim 9; The video compression device and the video restoration device are connected for communication.

Citation Information

Patent Citations

  • Adaptive 3D video coding and decoding method based on compressed sensing

    CN107509074A

  • Hyperspectral image compression method based on deep learning and distributed information source coding

    CN111145276A