A method and related equipment for locating video restoration areas

By extracting high-frequency residual features from video restoration detection and fusing them with RGB stream data in the early stage, and combining 3D Swin-UNet and spatiotemporal attention mechanisms, and employing full-time or sparse-time supervision, the problem of low-contrast detection failure and optical flow dependence in existing technologies is solved, and efficient and accurate localization of video restoration areas is achieved.

CN121789019BActive Publication Date: 2026-05-26HUNAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN UNIV
Filing Date
2026-03-09
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing video restoration and detection methods fail to detect in low-contrast scenes, cannot effectively utilize temporal consistency, result in high computational costs and reliance on expensive frame-by-frame fully supervised annotation, thus limiting practical applications.

Method used

By extracting high-frequency residual features from the training noise stream data and fusing them with the RGB stream data in the early stage, and combining 3D Swin-UNet and spatiotemporal attention mechanisms, the network parameters are updated by calculating the loss function in full-time or sparse-time supervision mode, thereby achieving efficient and accurate localization under sparse-time supervision.

Benefits of technology

It effectively solves the problems of low-contrast area detection failure and optical flow feature dependence, reduces the dependence on massive frame-by-frame annotation data, and achieves efficient and accurate video restoration area positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789019B_ABST
    Figure CN121789019B_ABST
Patent Text Reader

Abstract

This invention provides a video restoration region localization method and related equipment, relating to the field of video restoration technology. It extracts high-frequency residual features from training noise stream data and performs early fusion of these features with the original pixel features in the training RGB stream data to capture the correlation features between content and noise. A loss function is calculated based on a set supervision mode and a dense prediction mask sequence, and the network parameters of the 3D Swin-UNet with a spatiotemporal attention mechanism are updated based on the loss function value. This eliminates optical flow dependence and effectively solves the problems of existing video restoration detection methods when facing modern deep learning restoration algorithms, such as failure to detect low-contrast regions, loss of small target forgery traces, low inference efficiency due to excessive reliance on optical flow features, and excessive dependence on massive frame-by-frame labeled data in fully supervised training. It achieves efficient and accurate localization under sparse temporal supervision.
Need to check novelty before this filing date? Find Prior Art