Wire Erasing Method, Device and Storage Medium Based on Deep Generation Model

Through a deep generation model-based method, using self-attention mechanism and convolutional neural network to perform wire-width erasure on videos, the problem of insufficient content organization and coherence in traditional methods is solved, and a more natural and efficient video repair effect is achieved.

CN114742698BActive Publication Date: 2025-05-30TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210433690.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-24
Publication Date
2025-05-30
Estimated Expiration
2042-04-24

AI Technical Summary

Technical Problem

The traditional wire erasing method lacks reasonable organizational scheduling and inter-frame coherence considerations, resulting in the generated content being unnatural enough and prone to artifacts and chromatic aberrations.

Method used

A video depth generation model based on the deep generation model is used to construct a video depth generation model of self-attention mechanism and convolutional neural network. The effective content in the video clip is matched and reasonable content is generated through pre-training models, and cropped according to the confidence level to cover the Wia line frame by frame.

Benefits of technology

A more natural video repair is achieved, time coherence is enhanced, artifacts and chromatic aberrations are reduced, and the automation efficiency of Wia wire erasing is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114742698B_ABST
    Figure CN114742698B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for erasing wire ropes in a video based on a deep generation model. A video deep generation model based on a self-attention mechanism and a convolutional neural network is constructed and trained, or a suitable pre-trained video deep generation model is selected. Using any video segment in the target video as the basic erasing unit, with the pre-trained video deep generation model, first match all valid contents within the video segment to obtain valid information to generate reasonable content, and then crop the generated content according to the confidence level, and cut out the area with a confidence level higher than the preset value to cover the position of the wire rope to achieve the erasure of the wire rope. Repeat this process until all the wire ropes in all video segments are erased. The present invention utilizes the correlation between the content generated by the generation model and other contents and the inter-frame coherence within the video to gradually and iteratively complete video repair segment by segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video processing method, and particularly to a method, device and storage medium for erasing wire ropes for wire walking based on a deep generation model. Background Art

[0002] Currently, in the traditional film and television industry, for the need of plot design, actors usually rely on wire ropes to complete high - difficulty actions such as "flying over eaves and walking on walls". Subsequently, in the post - processing of the captured video, professional technicians are required to erase the wire ropes. This process is very cumbersome because technicians need to erase the wire ropes appearing in the video frame by frame and from the outside to the inside with the naked eye.

[0003] In recent years, with the continuous development of computing power, deep - learning methods represented by deep neural networks have been successfully applied to a large number of computer vision tasks and achieved good results, including target removal. With the help of deep neural networks, the specified area in the video will be automatically and quickly covered by the content generated by the deep network frame by frame to achieve the removal of the specified target. However, most of the existing methods directly use the content generated by the generation model to cover the original content frame by frame at one time, lacking reasonable organization and scheduling of the content and consideration of frame - to - frame coherence. For wire - rope erasure, due to the viewing requirements of film and television dramas, wire - rope erasure not only requires generating reasonable content to cover the wire ropes, but also requires the repaired video to have stronger temporal coherence to avoid artifacts and color differences. Therefore, directly applying traditional target - removal methods frame by frame in the field of video wire erasure to achieve automatic wire erasure faces many problems. Summary of the Invention

[0004] The present invention provides a method, device and storage medium for erasing wire ropes for wire walking based on a deep generation model to solve the technical problems existing in the known technology.

[0005] The technical solution adopted by the present invention to solve the technical problems existing in the known technology is as follows: A method for erasing wire ropes for wire walking based on a deep generation model, constructing and training a video deep generation model based on self - attention mechanism and convolutional neural network, or selecting a suitable pre - trained video deep generation model; taking any video segment in the target video as the basic erasure unit, using the pre - trained video deep generation model, first matching all valid content in the video segment to obtain valid information to generate reasonable content, and then cropping the generated content according to the confidence level, and cutting out the area with a confidence level higher than the preset value to cover the position of the wire rope to achieve the erasure of the wire rope; repeating this process until all the wire ropes in all video segments are erased.

[0006] Further, it includes the following specific steps:

[0007] Step 1: Extract a video clip from the original video. Through a pre-trained video depth generation model, search for the area to be erased and other valid areas within the video clip, so as to match and generate appropriate content to cover the wire to be erased in the video clip.

[0008] Step 2: Use the distance between the edge and the center of the generated content as a basis to distinguish the credibility of the generated content. Crop the generated content according to the credibility. The high-credibility area is used to cover the corresponding part of the wire, and the low-credibility area is discarded.

[0009] Step 3: Cover the wire from the outside to the inside with the cropped generated content, and insert the updated video clip into the original video clip to complete one update.

[0010] Step 4: Repeat Step 1 to Step 3 until the erasure of the wire in the entire video is completed.

[0011] Furthermore, Step 1 includes the following sub-steps:

[0012] Step A1: Select a video clip from the complete video with the wire to be erased in sequence.

[0013] Step A2: Select a suitable pre-trained video depth generation model for generating content to cover the wire.

[0014] Step A3: For each video frame in the clip, perform the following operations in sequence: wire recognition and extraction, binarization processing, dilation operation, and inversion operation to obtain the frame mask corresponding to each video frame in the clip.

[0015] Step A4: After normalizing the video clip, multiply it by the inversion of the mask to obtain the valid area of the video clip. Input the valid area of the video clip into the video depth generation model, and generate a preliminarily erased video clip through the following formula

[0016] R i ’ = G[(255 - V i ) ÷ 255) × (1 - M i )];

[0017] represents an arbitrary video clip with a length of t 0 frames;

[0018] represents the mask obtained through binarization processing of each video frame;

[0019] G(·) represents any pre-trained video depth generation model;

[0020] represents the real number space;

[0021] t 0 represents the length of the video segment to be erased;

[0022] h represents the height of the video to be erased;

[0023] w represents the width of the video to be erased.

[0024] Furthermore, step two includes the following sub-steps:

[0025] Step B1, according to the mask of each frame in the segment, calculate the distance matrix D corresponding to each frame in the video segment by the following formula i ∈ {1,..., t 0}:

[0026]

[0027] Step B2, according to the given confidence threshold l, calculate the confidence matrix I corresponding to each frame in the segment by the following formula i ∈ {1,..., t 0};

[0028]

[0029] t 0 represents the length of the video sequence to be erased;

[0030] a represents the abscissa of any point in the distance matrix D i ;

[0031] b represents the ordinate of any point in the distance matrix D i ;

[0032] a' represents the abscissa of any point on the mask M other than a i ;

[0033] b' represents the ordinate of any point on the mask M other than b i ;

[0034] Furthermore, in step three, generate the updated video segment according to the following formula:

[0035] V i ' = V i × (1 - M i ) + (M i - I i ) × R i ;

[0036] V i ' represents the video segment after one iteration of update;

[0037] V i represents an arbitrary video clip with a length of t 0 frames;

[0038] M i represents a mask obtained by binarizing each video frame;

[0039] I i represents a confidence matrix;

[0040] R i represents the video clip that has been erased.

[0041] The present invention also provides a device for implementing a wire removal method based on a deep generative model, including a memory and a processor. The memory is used to store a computer program. The processor is used to execute the computer program and implement the steps of the above-mentioned wire removal method based on the deep generative model when executing the computer program.

[0042] The present invention also provides a storage medium. The storage medium stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned wire removal method based on the deep generative model are implemented.

[0043] The advantages and positive effects of the present invention are as follows: The present invention makes full use of the correlation between the content generated by the generative model and other content, as well as the inter-frame coherence within the video, and has the following advantages:

[0044] Novelty: For the first time, it is proposed to organize and retrieve the content generated by the generative model according to the confidence level, and instead of the traditional one-time repair and frame-by-frame repair, a progressive iterative method is proposed to complete video repair segment by segment.

[0045] Effectiveness: Experiments prove that compared with other existing object removal methods, the intelligent wire removal method based on the generative model designed by the present invention has improved performance on both traditional object removal datasets and wire removal datasets, indicating the effectiveness of the present invention.

[0046] Generality: The present invention mainly focuses on the algorithm perspective, so it is not limited to the generative model itself. It can be applied as a "plug-and-play" module to any generative model and achieve a certain performance improvement, indicating the generality of the present invention. Brief Description of the Drawings

[0047] Figure 1 is a schematic diagram of the working process of a wire removal method based on a deep generative model of the present invention.

[0048] Figure 2 is a schematic diagram of the operation steps for obtaining the frame mask corresponding to each video frame within the segment. Detailed Implementation Modes

[0049] To further understand the content, features and effects of the present invention, the following embodiments are listed and described in detail in conjunction with the accompanying drawings as follows:

[0050] Please refer to Figures 1 to 2 , a wire removal method based on a deep generation model, constructs a video deep generation model based on a self-attention mechanism and a convolutional neural network and trains it, or selects a suitable pre-trained video deep generation model; taking any video segment in the target video as the basic erasure unit, using the pre-trained video deep generation model, first match all valid contents in the video segment to obtain the matched valid information to generate reasonable content, and then crop the generated content according to the confidence level, and cut out the area with a confidence level higher than the preset value to cover the position of the wire to achieve the removal of the wire; repeat this process until all the wires in all video segments are erased.

[0051] For any video segment, the method of removing the wire from the outside to the inside can be adopted, that is, gradually covering from the outside of the wire to the inside of the wire. Each video segment may require one or more iterations to be completely erased. After a video segment is completely erased, the erasure of the video segment is completed, and then the next video segment is erased until all the wires in all video segments are erased.

[0052] The video deep generation model can adopt a suitable video deep generation model in the prior art; it can also adopt software or components in the prior art and be constructed by conventional technical means.

[0053] Preferably, a wire removal method based on a deep generation model may include the following specific steps:

[0054] Step 1: A video segment can be intercepted from the original video, and the pre-trained video deep generation model can be used to complete the search for the area to be erased in the video segment and other valid areas in the segment, so as to match and generate appropriate content to cover the wire to be erased in the video segment.

[0055] Step 2: The distance between the edge and the center of the generated content can be used as a basis to distinguish the credibility of the generated content. The generated content can be cropped according to the credibility. The high-credibility area is used to cover the corresponding part of the wire, and the low-credibility area can be discarded.

[0056] Step 3: The cropped generated content can be covered on the wire from the outside to the inside, that is, the cropped generated content is gradually covered from the outside of the wire to the inside of the wire, and then the updated video segment is inserted into the original video segment to complete one update.

[0057] Step 4: Repeat Steps 1 to 3 until the wire removal of the entire video is completed.

[0058] Preferably, Step 1 may include the following sub-steps:

[0059] Step A1: Select a video clip in sequence from the complete video with the wire to be removed.

[0060] Step A2: Select a suitable pre-trained video depth generation model for generating content to cover the wire.

[0061] Step A3: For each video frame in the clip, perform the following processing in sequence: wire recognition and extraction, binarization processing, dilation operation, and inversion operation to obtain the frame mask corresponding to each video frame in the clip.

[0062] Step A4: After normalizing the video clip, multiply it by the inversion of the mask to obtain the effective region of the video clip. Input the effective region of the video clip into the video depth generation model, and a preliminarily erased video clip can be generated through the following formula

[0063] R i ’ = G[(255 - V i ) ÷ 255) × (1 - M i )];

[0064] represents an arbitrary video clip with a length of t 0 frames;

[0065] represents the mask obtained through binarization processing of each video frame;

[0066] G(·) represents any pre-trained video depth generation model;

[0067] represents the real number space;

[0068] t 0 represents the length of the video clip to be removed;

[0069] h represents the height of the video to be removed;

[0070] w represents the width of the video to be removed.

[0071] For wire recognition and extraction, binarization processing, dilation operation, inversion operation, and normalization processing, applicable modules in the prior art can be used; or software or modules in the prior art can be used and constructed by conventional technical means.

[0072] Further, Step 2 may include the following sub-steps:

[0073] Step B1. According to the mask of each frame in the segment, the distance matrix D corresponding to each frame in the video segment can be calculated by the following formula i ∈ {1,..., t 0}:

[0074]

[0075] Step B2. According to the given confidence threshold l, the confidence matrix I corresponding to each frame in the segment can be calculated by the following formula i ∈ {1,..., t 0};

[0076]

[0077] The min function represents the function composed of the minimum values on the common domain in the function group of the contained elements.

[0078] t 0 represents the length of the video sequence to be erased;

[0079] a represents the abscissa of any point in the distance matrix D i ;

[0080] b represents the ordinate of any point in the distance matrix D i ;

[0081] a′ represents the abscissa of any non-a point on the mask M i ;

[0082] b′ represents the ordinate of any non-b point on the mask M i ;

[0083] l represents the confidence threshold.

[0084] Preferably, in step three, the updated video segment can be generated according to the following formula

[0085] V i ′ = V i × (1 - M i ) + (M i - I i ) × R i ;

[0086] V i ′ represents the video segment after one iteration of update;

[0087] V i represents an arbitrary video segment with a length of t 0 frames;

[0088] M i represents the mask obtained by binarizing each video frame;

[0089] I i represents a confidence matrix;

[0090] R i represents the video segments that have been erased.

[0091] The present invention also provides a device for implementing the wire removal method based on a deep generation model, including a memory and a processor. The memory is used to store a computer program; the processor is used to execute the computer program and implement the steps of the above-mentioned wire removal method based on a deep generation model when executing the computer program.

[0092] The present invention also provides a storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned wire removal method based on a deep generation model are implemented.

[0093] The working process and working principle of the present invention will be further described below with a preferred embodiment of the present invention:

[0094] A wire removal method based on a deep generation model of the present invention uses video segments as the basic erasure unit. By using a pre-trained video deep generation model, it first matches all valid contents within the video segment to obtain valid information to generate reasonable contents; then uses an iterative erasure algorithm to organize the generated contents to gradually complete the erasure.

[0095] A wire removal method based on a deep generation model of the present invention mainly has two stages: the task of the first stage is to obtain contents by using a pre-trained video deep generation model, and the second stage is mainly to organize the generated contents to cover the wires in the original video video segment by video segment to complete the wire removal.

[0096] The video deep generation model first extracts a video segment from the original video for operation. The operation process mainly completes the search for the area to be erased in the video segment and other valid areas within the segment, so as to match suitable contents to cover the wires to be erased in the video segment. Subsequently, the generated content is cropped from the outside to the inside according to a certain thickness. The area with a certain thickness on the outside is considered to have high confidence and is used to cover the corresponding part of the wire, and the area on the inside is considered to have low confidence and is discarded. The updated video segment is inserted into the original video to complete one wire removal, and this step is repeated several times until the entire video is erased.

[0097] The specific steps of a wire removal method based on a deep generation model are as follows:

[0098] 1) First, select a trained deep generation model G(·) for generating content, and the complete video to be repaired Then sample the complete video and select a video clip

[0099] 2) According to Figure 2 As shown, for each video frame in the sampled video clip, successively perform wire identification and extraction, binarization processing, dilation operation, and inversion operation to obtain the frame mask corresponding to each video frame in the clip

[0100] 3) After normalizing the video clip V i Multiply by the inverse of the mask to obtain the effective region of the video clip and input it into the generation model. Through formula (1), generate a video clip after "rough erasing"

[0101] We define to represent an original video with a length of t frames, to represent an arbitrary video clip with a length of t 0 frames. represents the mask obtained by binarizing each video frame, represents the result of wire erasing. G(·) represents any trained deep generation model. Then, in each iteration process, a wire erasing result will be generated first:

[0102] R i ’ = G[(255 - V i ) ÷ 255) × (1 - M i )] (1);

[0103] 4) According to the mask of each frame in the clip, calculate the distance matrix D corresponding to each frame in the video clip by formula (2) i ∈ {1,..., t 0}.

[0104]

[0105] 5) According to the given confidence threshold l, calculate the confidence matrix I corresponding to each frame in the clip by formula (3) i ∈ {1,..., t 0}. Then, the erasing structure V i ′ of this iteration can be obtained by formula (4). Insert V i ′ back into V i in order to complete the wire erasing of this time.

[0106]

[0107] l represents the confidence threshold;

[0108] t 0 represents the length of the video sequence to be erased;

[0109] a represents the abscissa of any point on the distance matrix D i ;

[0110] b represents the ordinate of any point on the distance matrix D i ;

[0111] a' represents the abscissa of any point on the mask M other than a i ;

[0112] b' represents the ordinate of any point on the mask M other than b i ;

[0113] Let V i ' represent the video segment after one iteration of update; then there is:

[0114] V i ' = V i ×(1 - M i ) + (M i - I i ) × R i (4);

[0115] V i represents an arbitrary video segment with a length of t 0 frames;

[0116] M i represents the mask obtained by binarizing each video frame;

[0117] I i represents the confidence matrix;

[0118] R i represents the video segment that has been erased.

[0119] 6) Repeat steps 1) to 5) until all video segments are erased, and the erasure result of the video is obtained.

[0120] The above-described embodiments are only used to illustrate the technical ideas and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The patent scope of the present invention cannot be limited only by these embodiments, that is, any equivalent changes or modifications made according to the spirit disclosed by the present invention still fall within the patent scope of the present invention.

Claims

1. A method for erasing wire ropes in aerial work scenes based on a deep generative model, characterized in that, a video deep generative model based on the self-attention mechanism and convolutional neural network is constructed and trained, or a suitable pre-trained video deep generative model is selected; taking any video segment in the target video as the basic erasure unit, using the pre-trained video deep generative model, first match all valid contents in the video segment to obtain valid information to generate reasonable content, and then crop the generated content according to the confidence level, and cut out the area with a confidence level higher than the preset value to cover the position of the wire rope to achieve the erasure of the wire rope; repeat this process until all the wire ropes in all video segments are erased; including the following specific steps: Step 1, extract a video segment from the original video, and through the pre-trained video deep generative model, complete the search for the area to be erased in the video segment and other valid areas in the segment, so as to match and generate suitable content to cover the wire rope to be erased in the video segment; Step 2, use the distance between the edge and the center of the generated content as the basis to distinguish the credibility of the generated content, crop the generated content according to the credibility, the high-credibility area is used to cover the corresponding part of the wire rope, and the low-credibility area is discarded; Step 3, cover the wire rope with the cropped generated content from the outside to the inside, and insert the updated video segment into the original video segment to complete one update; Step 4, repeat Step 1 to Step 3 until the wire ropes in the entire video are erased; Step 1 includes the following sub-steps: Step A1, sequentially select a video segment from the complete video with the wire rope to be erased; Step A2, select a suitable pre-trained video deep generative model for generating content to cover the wire rope; Step A3, for each video frame in the segment, perform the following processing in sequence: wire rope recognition and extraction, binarization processing, dilation operation and inversion operation to obtain the frame mask corresponding to each video frame in the segment; Step A4, after normalizing the video clip, multiply it by the inverse of the mask to obtain the valid region of the video clip, and input the valid region of the video clip into the video depth generation model to generate a preliminarily erased video clip through the following formula R i ’ = G[(255 - V i ) ÷ 255) × (1 - M i )]; Denotes an arbitrary video segment of length t 0 frames; denotes the mask obtained by performing binaryzation processing on each video frame; G(·) represents any pre-trained video deep generative model; denote the real number space; t 0 Indicates the length of the video segment to be erased; h represents the height of the video to be erased; w represents the width of the video to be erased.

2. The method for erasing wire ropes in aerial work scenes based on a deep generative model according to claim 1, characterized in that, Step 2 includes the following sub-steps: Step B1. According to the mask of each frame in the video clip, calculate the distance matrix D corresponding to each frame in the video clip by the following formula i ∈ {1, …, t 0}: Step B2: According to the given confidence threshold \(l\), calculate the confidence matrix \(I\) corresponding to each frame within the segment according to the following formula i \(\in \{1,\ldots,t\}\) 0 \}\); t 0 Indicates the length of the video sequence to be erased; a represents the distance matrix D i The abscissa of any point; b represents the distance matrix D i The ordinate of any point; a' represents the mask M i Any abscissa other than a on it; b' represents the mask M i Any ordinate other than b on it.

3. The method for erasing wire ropes in aerial work scenes based on a deep generative model according to claim 1, characterized in that, In Step 3, generate the updated video segment according to the following formula: V i ′ = V i × (1 - M i ) + (M i - I i ) × R i ; V i ′ represents a video segment that has completed one iteration update; V i represents any video segment of length t 0 frames; M i represents the mask obtained by binarizing each video frame I i represents a confidence matrix; R i Indicates an erased video segment.

4. A device for implementing the method for erasing wire ropes in aerial work scenes based on a deep generative model, including a memory and a processor, characterized in that, the memory is used for storing a computer program; the processor is used for executing the computer program and implementing the steps of the method for erasing wire ropes in aerial work scenes based on a deep generative model according to any one of claims 1 to 3 when executing the computer program.

5. A storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, the steps of the method for erasing wire ropes in aerial work scenes based on a deep generative model according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Multi-dimensional classroom quantification system and method based on computer vision

    CN110334610A

  • Picture visual effect grade evaluation method and system, electronic device and storage medium

    CN112017179A