Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

123 results about "Video reconstruction" patented technology

Generative video coding and decoding method based on multi-modal large model

The invention relates to the technical field of video coding and decoding, and discloses a multi-mode large model-based generative video coding and decoding, which comprises a key frame selection module for determining a key frame by analyzing the semantic and motion characteristics of a video frame; the multi-modal semantic description generation module is used for generating semantic description according to the key frame and the video clip; the key frame compression module is used for realizing efficient compression through latent variable modeling and entropy coding; the key frame reconstruction module reconstructs a key frame by using a conditional diffusion model in combination with the compressed data and the semantic description information; and the video generation module generates a non-key frame by using the semantic description and the key frame, and reconstructs a complete video. Through key frame screening combining semantic and motion information, key frame compression and reconstruction based on a conditional latent variable diffusion model, and frame supplementation and frame insertion generation based on semantic description, efficient compression and high-quality video reconstruction can be realized under a low code rate, and the video storage efficiency and the visual quality are effectively improved.
Owner:上海芯开技术有限公司

Video snapshot compression imaging reconstruction method based on space-time deformable attention

The invention provides a video snapshot compression imaging reconstruction method based on spatio-temporal deformable attention, which improves the reconstruction quality and efficiency, and comprises the following steps: inputting a single frame compression measurement value and a measurement matrix into an initial reconstruction module to obtain an initial reconstruction video frame; inputting the initial reconstructed video frame into a feature extraction encoder, mapping the initial reconstructed video frame to a high-dimensional feature space through multi-layer 3D convolution, and outputting a feature map; the feature map is input into a plurality of stacked DenseRNet Blocks, and the number of the DenseRNet Blocks is one; the DenseRNet Block internally comprises a plurality of DeT Blocks, after the DenseRNet Block dynamically divides an input feature channel, grouping progressive processing and feature fusion are carried out through the plurality of DeT Blocks, and the DeT Blocks comprise a deformable space convolution branch used for modeling local deformation perception, a time self-attention branch used for modeling global time sequence dependence and a feature interaction module used for cross-channel information interaction; and the features processed by the DenseRNet Block are input into a video reconstruction decoder, and a reconstructed video sequence is output through up-sampling of transposition convolution and refining of multilayer 3D convolution.
Owner:DALIAN UNIV

Machine learning models for reconstruction and synthesis of dynamic scenes from video

In various examples, systems and methods are disclosed relating to reconstruction and synthesis of dynamic scenes from video, such as to generate a four-dimensional (4D) representation of one or more scenes based on one or more videos (e.g., two-dimensional (2D) videos) of the one or more scenes. A system may determine, using a neural network and based on a three-dimensional (3D) representation of one or more scenes, a 4D representation of the one or more scenes, the 3D representation generated by a featurizer using a plurality of first image frames from video data of the one or more scenes. The system may determine, from the 4D representation, a target image having a target pose and a target time.
Owner:NVIDIA CORP

Structure large displacement estimation method based on Canny-Hough transformation and KLT optical flow

The invention discloses a structure large displacement estimation method based on Canny-Hough transformation and KLT optical flow, and relates to the related field of signal processing and vision measurement, and the method comprises the steps: collecting and preprocessing structure vibration data; carrying out integer pixel displacement estimation on the basis of Canny-Hough transformation; reconstructing the video; performing sub-pixel displacement estimation based on the KLT optical flow; and physical displacement. According to the method, the limitation of a single identification method in identifying different pixel levels (pixel level and sub-pixel level) is overcome, the measurement precision of the large displacement of the structure is remarkably improved, and an estimation error caused when an image pyramid is used for large displacement tracking is avoided. The research provides a new target-free large-displacement motion estimation scheme for structural vibration measurement based on computer vision, and has important reference value.
Owner:CENT SOUTH UNIV

HDR video reconstruction method based on standardized stream

The invention discloses an HDR video reconstruction method based on a standardized stream, and belongs to the technical field of high dynamic range image processing. The method comprises the following steps of: firstly, constructing a convolution optical flow estimation module with a self-adaptive normalized structure, wherein the convolution optical flow estimation module is used for accurately acquiring optical flow information between adjacent frames in an alternative exposure LDR video image sequence; then, carrying out multi-level feature alignment on the image sequence through an image alignment module so as to reduce alignment errors caused by illumination difference and movement; and finally, inputting the aligned and fused multi-level LDR image features into a standardized flow reconstruction network to realize high-quality HDR video image reconstruction. Aiming at the video reconstruction problem under the alternate exposure condition, the invention designs a standardized flow modeling structure considering the optical flow estimation precision and the feature alignment effect, and effectively improves the HDR video reconstruction quality in a complex dynamic scene.
Owner:BEIHANG UNIV

Video snapshot compression imaging reconstruction method and system

The invention relates to a video snapshot compression imaging reconstruction method and system. The method comprises the following steps: inputting a video frame sequence and a time-varying mask set thereof into a measurement model to obtain initial estimation; constructing a reconstruction network which comprises a feature extraction module, a gating residual network module and a video reconstruction module; the feature extraction module comprises two three-dimensional convolution layers, each three-dimensional convolution layer is connected with an activation function, and the feature extraction module extracts initial features from the initial estimation; inputting the initial features into a gating residual network module, and outputting reconstruction information features; and the video reconstruction module fuses the reconstruction information features, and performs up-sampling and detail refining to reconstruct a video sequence. According to the method, on the premise that parameters and computing power are hardly increased, ghosting and flickering are effectively restrained, the stability of long-time reconstruction is improved, and an effective scheme is provided for SCI reconstruction with the high compression ratio, the super-definition resolution ratio and the long sequence.
Owner:GUANGDONG UNIV OF TECH

Latent Geodesic Traversal Across Multi-Axis Hyperspaces for Real-Time Video Reconstruction and Augmentation

A system and method for latent geodesic traversal across multi-axis hyperspaces for real-time video reconstruction and augmentation. Spatiotemporal video data are compressed into navigable latent representations using hierarchical and Lorentzian autoencoders that preserve geometric and temporal structure. A geodesic traversal engine computes paths across spatial, temporal, spectral, and semantic axes, guided by symbolic anchors and spatiotemporal routing protocols. A correlation network restores fine detail, while an augmentation generator synthesizes additional or counterfactual content to enable infinite zoom, continuous multi-scale exploration, and temporally coherent augmentation. A strategy caching system preserves successful traversal patterns for reuse, supporting persistent learning and adaptive real-time performance.
Owner:ATOMBEAM TECH INC

Cloud edge-end collaborative monitoring video transmission method, system and device based on key semantic guidance video super-resolution and medium

The invention provides a cloud side-end collaborative monitoring video transmission method, system and device based on key semantics guiding video super-resolution and a medium, and relates to the field of super-resolution. The method comprises the following steps: a terminal sends each low-resolution video segment and two corresponding high-resolution key frames to an edge server; the edge server performs key object extraction on the plurality of high-resolution key frames to obtain the high-resolution key frames after the plurality of key objects are extracted, and sends each low-resolution video segment and the corresponding high-resolution key frames after the two key objects are extracted to a cloud server; and the cloud server performs reconstruction by using each low-resolution video segment and the high-resolution key frames extracted from the corresponding two key objects based on a video reconstruction model guided by the key objects to obtain a reconstructed high-resolution video sequence, so that the high-resolution video sequence can be obtained in a monitoring scene with limited bandwidth. And the video transmission bit rate is reduced, and meanwhile, the key semantic information of the reconstructed video is kept.
Owner:TSINGHUA UNIVERSITY +1

Method for simultaneously reconstructing dynamic and static scenes based on event camera

A method for simultaneously reconstructing a dynamic scene and a static scene based on an event camera belongs to the field of image processing, and comprises the following steps: dividing an original event stream into space-time grids, setting an event number threshold value, judging events exceeding the threshold value as dynamic events triggered by motion, and judging events not exceeding the threshold value as static events triggered by background; calculating an initial static reconstruction image of the static event through a convolution integral method, and further optimizing through a static reconstruction noise reduction network to obtain a static background image; dividing a dynamic event into voxel grids through a space-time window, and fusing the static background image and the voxel grids on the premise of introducing an event tag tensor to obtain a fusion tensor; and inputting the fusion tensor into a dynamic and static video reconstruction network to obtain a dynamic video with a static background. According to the method, the quality of the static reconstructed image is improved, the phenomenon of visual inconsistency caused by fusion after independent reconstruction is eliminated, and model overfitting caused by monotonous data set elements is avoided.
Owner:PEKING UNIV

HDR video reconstruction method based on brightness alignment

The invention discloses an HDR video reconstruction method based on brightness alignment, and the method comprises the steps: carrying out the gamma correction of continuous frames of an input LDR video, generating an HDR image domain, and splicing the HDR image domain with an original frame to form an input tensor; constructing a brightness alignment network model, and generating brightness alignment features through a brightness attention module; generating detail features through a detail synthesis module; performing dynamic weighted fusion on the brightness alignment features and the detail features through an adaptive mixing layer, and performing up-sampling to obtain final alignment features; and finally, generating an HDR video frame by fusing the final alignment features. Through collaborative optimization of the brightness attention mechanism and the time domain alignment module, the problem of brightness inconsistency in a motion scene is effectively solved, and the HDR reconstruction quality is remarkably improved.
Owner:DALIAN NEUSOFT UNIV OF INFORMATION

Efficient video compression method based on multivariate space-time entropy network

The invention discloses an efficient video compression method based on adaptive masks, which comprises the following steps of: firstly, introducing an adaptive channel pruning technology, dynamically generating a channel importance mask by fusing multivariate space-time prior, and autonomously identifying and removing channel information with low contribution in a reconstructed image; then, the initial coding content is optimized based on a multivariable spatio-temporal information entropy model, and important information is further screened through adaptive mask in combination with optical flow information and spatio-temporal priori of a current frame and a previous frame; a Swin Transform module is introduced to enhance the prediction capability of the context information; and finally, expanding a hyper-parameter range by adopting a unified training strategy, and realizing smooth code rate adjustment and wider adaptability. The channel-level dynamic optimization is realized through the adaptive mask, the video reconstruction quality is ensured while the compression efficiency is remarkably improved, and the method can be widely applied to video conferences, streaming media transmission and other scenes with high requirements on the compression efficiency and quality.
Owner:NANTONG MARINE ADVANCED RESEARCH INSTITUTE SOUTHEAST UNIVERSITY +1

Video processing method and device, computer readable storage medium and computer program product

The invention relates to a video processing method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises: acquiring an original video; tracking and detecting a target object in the original video to obtain an original motion track of the target object in the original video; obtaining a target picture size and a target shooting style; determining a picture extension parameter according to the original motion track, the target picture size and the original picture size of the original video; performing picture expansion on the original video according to the picture expansion parameter to obtain an expanded video; determining a cutting frame sequence according to the target movement track and the mapped movement track; and according to the cutting frame sequence, performing picture cutting on the expanded video according to the target picture size to obtain a target video of the original video in the target shooting style. By adopting the method, the efficiency and the effect of video reconstruction images can be improved.
Owner:XIAMEN MEITUZHIJIA TECH

Machine learning models for generative human motion simulation

In various examples, systems and methods are disclosed relating to receive at least one of a text prompt or a kinematic constraint and determine first human motion data using a motion model by applying the at least one of the text prompt or the kinematic constraint to the motion model. The motion model is updated by generating, using the motion model, second human motion data by applying motion capture (mocap) data and video reconstruction data as inputs to the motion model, receiving user feedback information for the second human motion data, and updating the motion model based on the user feedback information. The video reconstruction data is generated by reconstructing human motions from a plurality of videos. Physically implausible artifacts are filtered from the video reconstruction data using a motion imitation controller. The motion imitation controller is updated using at least one of Reinforced Learning (RL) or physics-based character simulations.
Owner:NVIDIA CORP

High-fidelity generation type video stream transmission system based on visual base model

The invention relates to a high-fidelity generative video stream transmission system based on a visual base model, which belongs to the field of image communication, and is characterized in that a visual enhancement-oriented generative codec is designed, and high-fidelity video reconstruction under a high compression ratio is realized through an asymmetric space-time compression strategy and time sequence consistency enhancement; a resolution scaling module is provided, the calculation complexity is remarkably reduced through a video super-resolution recovery module of adaptive resolution control and joint optimization, and real-time high-definition video processing is achieved; and constructing a network adaptive video stream transmission controller, an intelligent token discarding mechanism based on semantic importance and a mixed packet loss processing strategy to realize code rate scalable control and robust transmission under network fluctuation.
Owner:THE CHINESE UNIV OF HONG KONG (SHENZHEN) +1

Video processing method and apparatus, device, storage medium, and computer program product

The present application discloses a video processing method and apparatus, a device, a storage medium, and a computer program product. The method comprises: acquiring a video to be processed captured by an event camera and event data corresponding to the video to be processed; performing event feature extraction on the event data to obtain a motion region feature corresponding to the video to be processed; performing frame feature extraction on the video to be processed and the event data to obtain a motion holistic feature corresponding to the video to be processed, wherein the motion holistic feature is used for representing a dependency relationship between time and space of the video to be processed; performing feature fusion on the motion region feature and the motion holistic feature to obtain a fused feature; and performing video reconstruction on the basis of the fused feature to obtain a target video. The method provided in the present application can improve the resolution and frame rate of a restored video, thereby enhancing video quality.
Owner:THE HONG KONG UNIV OF SCI & TECH (GUANGZHOU)

Online volume video reconstruction method and device, equipment and medium

The invention relates to the technical field of videos, and provides an online volume video reconstruction method and device, equipment and a medium. Volume video reconstruction is performed on an original image of a target scene at the current moment by using an initialized Gaussian element model; the initialized Gaussian point attribute value is a parameter required by the Gaussian primitive model to obtain the volume video of the target scene at the previous moment, and the parameter of the initialized Gaussian primitive model is a parameter updated by online incremental training after the Gaussian primitive model obtains the volume video of the target scene at the previous moment; according to the method, the video at the current moment can be reconstructed by directly utilizing the reconstruction key information at the previous moment, the reconstruction quality of the volume video of the target scene at the current moment is improved, the Gaussian primitive model does not need to be retrained from the beginning to update the parameters of the model due to the incremental training mode, the number of iterations required for training the model is reduced, and the reconstruction efficiency is improved. And calculation redundancy and reconstruction time are reduced.
Owner:SHENZHEN UNIV

Bidirectional adaptive video super-resolution method based on frame difficulty index

The invention discloses a bidirectional adaptive video super-resolution method based on frame difficulty index, and belongs to the field of computer vision. The invention provides a frame-level dynamic reconstruction method for solving the problems that simple frame calculation is redundant and difficult frame reconstruction is insufficient due to the fact that an existing model adopts a fixed calculation strategy for video frames with different difficulties. According to the method, a motion detail decoupling propagation network is constructed, motion information is efficiently transmitted by utilizing a shallow forward propagation branch, and a deep backward propagation branch focuses on recovering texture details so as to decouple a time sequence propagation task; meanwhile, a frame reconstruction difficulty evaluation network is introduced to generate a global difficulty index, so that the receptive field weight of the adaptive time sequence fusion network and the refining depth of the dynamic refining network are regulated and controlled. According to the method, through explicit modeling frame-level reconstruction difficulty, adaptive matching of the model capacity and the video frame feature complexity is realized, and the video reconstruction performance is remarkably improved under limited computing power.
Owner:GUILIN UNIV OF ELECTRONIC TECH

A method for video coding based on implicit neural representation considering saliency

The application provides a video coding method based on implicit neural representation considering saliency, comprising: original video preprocessing; constructing a video implicit neural representation network based on a multi-scale feature grid, comprising a multi-scale feature grid and a decoder; optimizing the model through a saliency-guided training strategy; compressing the multi-scale feature grid and the decoder as compressed data to obtain a video code stream; sending and decompressing the video code stream, generating feature embedding through the feature grid according to the frame index of each frame, inputting the feature embedding into the decoder to output a corresponding reconstructed image, arranging the reconstructed images in sequence to obtain a decoded video. The application codes the video in an implicit neural network, proposes a multi-scale feature grid and a decoder based on a light-weight convolutional neural network, significantly improves the objective quality of video reconstruction, and introduces saliency preprocessing and a saliency-guided training method to comprehensively improve the visual quality of video reconstruction.
Owner:TONGJI UNIV

A video reconstruction method based on state-space equations and driven by neuromorphic signals.

This invention discloses a video reconstruction method driven by neuromorphic signals based on state-space equations, comprising: 1. constructing a video reconstruction network based on state-space equations; 2. introducing a random window displacement Mamba module designed for the spatial characteristics of neuromorphic signals; 3. introducing a Hilbert-filled Mamba module designed for the spatiotemporal characteristics of neuromorphic signals; 4. training the hybrid super-resolution network through backpropagation and continuously optimizing it until the loss function converges. The resulting video reconstruction model is used to reconstruct the neuromorphic signals to be processed, thereby generating corresponding high-quality videos. This invention achieves efficient operation and excellent visual effects through the linear global modeling capability of state-space equations and targeted network module design.
Owner:UNIV OF SCI & TECH OF CHINA

An Image and Video Reconstruction Method, System, Terminal, and Storage Medium Based on a Diffusion Model

The present invention discloses an image and video reconstruction method, system, terminal and storage medium based on a diffusion model. The method includes: obtaining original image data and an original text description, and inputting them into the diffusion model to obtain noise data and intermediate features; determining a time step prediction model, and performing an inverse generation process on the noise data, intermediate features and the original text description through the time step prediction model and the diffusion model to obtain initial image data; updating the time step prediction model according to the initial image data to obtain an updated time step prediction model; obtaining an updated text description, and performing an inverse generation process on the noise data, intermediate features and the updated text description through the updated time step prediction model and the diffusion model to obtain target image data. The present invention can effectively improve the reconstruction ability of the diffusion model and achieve accurate reconstruction of the original image data.
Owner:GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)

End-to-end video compression method based on foreground and background segmentation and sub-channel coding

The invention discloses an end-to-end video compression method based on foreground and background segmentation and sub-channel coding, and the method comprises the following steps: obtaining a target frame which comprises a current frame and a reconstruction frame of a previous frame; constructing an encoding end-decoding end video compression model, and training based on the target frame to obtain a target end-to-end video compression model; constructing a multi-loss function, and carrying out training optimization on the target coding end-decoding end video compression model to obtain an optimized target coding end-decoding end video compression model; and compressing the video to be compressed based on the target coding end-decoding end video compression model. According to the invention, foreground and background segmentation is realized by using the motion estimation module, the background processing flow is simplified, the calculation complexity is reduced, the coding data volume is reduced, and the compression efficiency is improved; foreground optical flow passes through the motion compensation module and the residual error generation module, foreground motion features are effectively extracted, and the video reconstruction quality is improved while accurate transmission of foreground motion information is ensured.
Owner:HUBEI ELECTRIC POWER TRANSMISSION & DISTRIBUTION ENG

Video coding method, video decoding method and device

PendingCN121771396AReduce bit rateGuaranteed reconstruction qualityBiological modelsDigital video signal modificationVideo encodingTheoretical computer science
The invention provides a video coding method and device and a video decoding method and device, and relates to the technical field of video coding and decoding. Comprises: obtaining a student model; the student model is a model obtained by guiding a first network model to train through a teacher model, the teacher model is a model obtained by training a second network model, and when the same video is coded based on the first network model and the second network model respectively, the first network model and the second network model are selected; the code rate of the coding result of the first network model is smaller than that of the coding result of the second network model; and obtaining a coding result of the video to be coded according to the student model. According to some embodiments of the invention, in a video coding scheme for carrying out video coding by using a deep learning network model, the video reconstruction quality is ensured, and the code rate of a video coding result is reduced at the same time.
Owner:HISENSE VISUAL TECH CO LTD

Unified classification method and device in loop filtering

The invention discloses a loop filtering method and device for reconstructing a video. The method receives input data of a current block, wherein the input data includes reconstructed samples of the current block. At least two loop filters are applied to a current block, where the at least two loop filters belong to a loop filter bank comprising a bilateral filter (BIF) and an adaptive loop filter (ALF), and the classification processes of the at least two loop filters share one input source, one or more classification rules, one or more processing units or a combination thereof. A filtered output generated by applying the at least two loop filters to the current block is provided.
Owner:MEDIATEK INC

FPGA-based infrared image data parallel processing method and circuit

This invention relates to the field of infrared imaging and low-level hardware processing technology, and discloses a parallel processing method and circuit for infrared image data based on FPGA. The method includes: establishing a pixel spatiotemporal mapping coordinate system and extracting pixel mapping point coordinate pairs; performing hardware topology sensing and spatial decoupling to output a spatial topology matrix; performing real-time pixel-level non-uniformity correction; implementing phase-locked noise reduction through a virtual offset field to obtain an equalized corrected bitstream; generating a quantization mapping curve based on saliency sensing; and performing real-time video reconstruction combined with adaptive power consumption control to output a video signal. This invention solves the problems of latency and instantaneous heat accumulation in large-area infrared data processing, achieving high-fidelity restoration of spatial topology and suppression of non-uniform noise; closed-loop energy management eliminates image blurring caused by thermal drift, comprehensively improving the system's imaging clarity and hardware operating efficiency.
Owner:HANGZHOU ZHIPU TECHNOLOGY CO LTD

Construction method, system and equipment of video reconstruction system of joint information source channel, and medium

The invention discloses a construction method of a video reconstruction system of a joint information source channel, the video reconstruction system, equipment and a medium. At a transmitting end, the method comprises the following steps: performing multi-frame joint semantic coding on a video sequence, extracting a potential representation simultaneously containing spatial information, time information and high-level semantic information, and directly mapping the potential representation to a wireless channel for transmission; and at a receiving end, a diffusion generation model is constructed based on the denoising network and the diffusion model, and semantic features are introduced to carry out condition guidance on the diffusion denoising process, so that generative reconstruction of the potential representation damaged by noise is realized. According to the method, the diffusion generation model is integrated into a deep joint source channel coding framework, so that the video reconstruction quality and semantic consistency are remarkably improved in a low signal-to-noise ratio and complex channel environment, the cliff effect is effectively weakened, and higher robustness and adaptive ability are achieved.
Owner:SHENZHEN UNIV

Systems and methods for generative video reconstruction using multimodal latent sensor data

A system and method for generating synthetic video from diverse sensor inputs within a unified computational framework. The system receives heterogeneous data such as acoustic, thermal, and textual streams, encodes each into modality-specific latent representations, and projects them into a shared geometric manifold. Within this manifold, convergence points known as multimodal landmarks are established and used to compute geodesic trajectories that describe relationships among the inputs. The trajectories are verified for reversibility to ensure that forward and reverse mappings remain consistent. A Lorentzian autoencoder then decodes the validated trajectories into temporally coherent video sequences derived from the multimodal evidence rather than reconstructed imagery. The system records geometric states for auditability and persistently stores the resulting landmarks and trajectories for reuse, enabling reversible, verifiable generation of synthetic video that accurately reflects the integrated sensor data.
Owner:ATOMBEAM TECH INC

A configuration decision optimization system and method for film and television rendering

The application relates to the field of film and television rendering, in particular to a configuration decision optimization system and method for film and television rendering, which comprises a shooting calibration module, an intra-frame rendering module, a scene reconstruction module, a video rendering module and a configuration compression module; the shooting calibration module is used for video de-jittering and adjusting a picture center; the intra-frame rendering module is used for area light rendering; the scene reconstruction module is used for generating a point cloud model; the video rendering module is used for merging a rendered video; and the configuration compression module is used for compressing a video stream; the application can ensure the definition and consistency of an output picture, reduce high-frequency artifacts and fuzzy sawteeth caused by shooting angle problems, reduce video rendering result jitter, improve image rendering quality, solve a sampling rate deficiency problem, improve the scale of multi-path video rendering, improve video reconstruction speed and rendering efficiency, and realize efficient video compression and transmission.
Owner:DIGITAL INTELLIGENCE CLOUD LIBRARY (BEIJING) TECHNOLOGY CO LTD

A video compression transmission method and related device

The application discloses a video compression transmission method and related equipment, the method comprises the following steps: obtaining a video sequence to be compressed, taking the first frame and the last frame as conditional frames; inputting the video sequence into the encoder of a pre-trained variational autoencoder to obtain a target latent space representation; processing the target latent space representation through the downsampling module of the compressor to generate an extreme compression representation; transmitting the conditional frames and the extreme compression representation to the receiving end to enable the receiving end to reconstruct the video through the upsampling module of the compressor, the generation model and the decoder of the variational autoencoder to obtain a reconstructed video sequence; the application provides time sequence boundary information through the conditional frames, and in combination with the reconstruction capability of the generation model, can effectively reduce the block effect, blur and high-frequency detail loss commonly seen in traditional methods; the conditional frames and the extreme compression representation significantly reduce the transmission code rate and bandwidth demand, can significantly improve the rate distortion performance, and can be widely applied to the technical field of video compression.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Error-resistant network transmission method for auxiliary stream based on deep learning video codec

This application provides an error-resistant network transmission method for an auxiliary stream based on a deep learning video codec, relating to the field of video transmission technology. The decoding end uses an erroneous bitstream for encoding and decoding. During the encoding and decoding process, the decoding result corresponding to the erroneous bitstream is set to all zeros, resulting in a slightly distorted decoded reconstructed frame APn-1. The decoding end uses the decoded reconstructed frame APn-1 as a reference image and refreshes its decoding buffer. Encoding and decoding are performed with all reference content except the reference image set to None, resulting in a correctly decoded reconstructed frame APn. This method is a low-error network transmission method that reduces transmission bandwidth requirements while ensuring video reconstruction quality, thus enhancing the error-resistant robustness of the deep learning-based video codec.
Owner:TSINGHUA UNIVERSITY +1

Video reconstruction method and system based on prior features and global frequency domain filtering

This invention proposes a video reconstruction method and system based on prior features and global frequency domain filtering, relating to the field of video coding technology. The method involves inputting video frames into a pre-trained encoder to extract multi-scale features, and progressively fusing these deep features to obtain the first prior feature. Hybrid residual features are then extracted from the video frames based on a hybrid residual grid. The first prior feature and the hybrid residual feature are input into convolutional blocks respectively to obtain the second prior feature and residual grid features. These two features are then fused based on adaptive weights to obtain a fused feature containing both prior and grid information. The fused feature is input into a coupled mapping RNN module to obtain state features, which are then input into a global frequency domain filtering upsampling module for video frame reconstruction, resulting in reconstructed video frames containing information at different frequencies. This invention improves the high-quality reconstruction performance of the implicit neural video representation model by enhancing the representational power of intermediate features and the decoding power of the upsampling decoding module.
Owner:SHANDONG UNIV