Intelligent traceless erasing method and system for video picture characters
Through the space-time dual-domain attention model and dynamic weight allocation technology, the flickering and compatibility problems in text erasing of video images are solved, and efficient and stable traceless text repair effect is achieved.
Patent Information
- Application Number
- CN202510443725.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-09
AI Technical Summary
There are flickering problems caused by the lack of time consistency in the existing video screen text erasure technology, insufficient compatibility problems caused by fast moving text, and the contradiction between computing efficiency and repair quality.
The space-time dual-domain attention model is used to combine dynamic weight allocation, and the characteristics of the spatial domain and the temporal domain are integrated through cross-attention calculation, and the weight is dynamically adjusted to adapt to the speed of text movement. The sliding window averaging method and frequency domain compensation algorithm are combined to achieve traceless erasure of text on the video screen.
Effectively reduce the flicker frequency, improve the timing coherence of the repair area, reduce the residual rate of the drag, improve the repair efficiency and quality, and meet the real-time needs.
Smart Images

Figure CN120298265A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of video stream processing, and particularly to an intelligent seamless erasure method and system for text in video frames. Background Art
[0002] The method for erasing text in video frames is to identify and remove the text area in the video through algorithms and reconstruct the background by synchronously referring to the spatio-temporal information of the previous and subsequent frames. However, in the existing video text erasure technologies, the following technical bottlenecks mainly exist:
[0003] The single-frame image restoration method based on the generative adversarial network only relies on the texture information of the current frame and does not consider the temporal correlation of the previous and subsequent frames, which results in flickering phenomena such as background jitter and brightness mutation when the restored video is played continuously. Especially in the static text erasure scenario, the flickering frequency is as high as 30%.
[0004] Moreover, for the optical flow estimation method based on ProPainter, although time information is introduced, it is sensitive to the movement speed of the text. When the text movement speed exceeds 30 pixels / frame, such as bullet screens and scrolling subtitles, the optical flow estimation error increases significantly, with the error > 10 pixels, resulting in trailing or residual artifacts in the restored area. In addition, the public technology does not design a dynamic weight allocation mechanism and cannot adaptively adjust the contribution ratio of spatio-temporal domain features to the restoration result.
[0005] In addition, models relying on optical flow calculation or 3D convolution consume a large amount of computing resources and do not optimize the feature fusion method for the text erasure scenario, which results in low restoration efficiency and is difficult to meet the real-time requirement. Summary of the Invention
[0006] This application provides an intelligent seamless erasure method and system for text in video frames, aiming to solve the flickering problem caused by the lack of temporal consistency, the compatibility problem caused by fast-moving text, and the contradiction between computing efficiency and restoration quality in the related technologies.
[0007] To achieve the above object, this application provides an intelligent seamless erasure method for text in video frames, including the following:
[0008] Obtain the first video frame sequence of the video stream to be processed and perform preprocessing.
[0009] Extract spatio-temporal dual-domain features from the preprocessed first video frame sequence to obtain the first spatio-temporal features.
[0010] Establish a spatio-temporal dual-domain attention model according to the first spatio-temporal features. The spatio-temporal dual-domain attention model calculates and fuses the spatial domain and time domain features through cross-attention and dynamically allocates fusion weights based on the text movement speed.
[0011] Obtain the second video frame sequence of the video stream to be processed and perform preprocessing to extract the second spatio-temporal features.
[0012] Input the second spatio-temporal features into the spatio-temporal dual-domain attention model to execute the matching process.
[0013] The matching process includes main matching and auxiliary determination.
[0014] Main matching: According to the dynamic propagation weight matrix, match the second spatio-temporal features with the spatio-temporal dual-domain features in the model. If the matching degree is higher than the matching threshold, directly generate the repaired second video frame.
[0015] Auxiliary determination: If the matching degree is lower than the matching threshold, trigger the temporal consistency compensation mechanism, calculate the SSIM similarity of the repaired regions of adjacent frames. When the SSIM is lower than the auxiliary determination threshold, smooth the temporal mutation by the sliding window averaging method and re-repair to generate the corrected second video frame.
[0016] Generate a seamless erased video stream based on the repaired second video frame.
[0017] As a preferred solution of the present invention, obtain the first video frame sequence of the video stream to be processed and perform preprocessing, specifically including:
[0018] The first video frame sequence includes the current frame and its adjacent front and rear frames. The preprocessing includes detecting the text region through a multi-scale feature pyramid network to generate an initial text mask and a motion trajectory prediction map. Extract N consecutive frames from the video stream to be processed as the first video frame sequence, which includes the current frame, the first M frames and the last M frames, and eliminate the offset error caused by camera movement between adjacent frames through a time alignment module.
[0019] Input the first video frame sequence into the multi-scale feature pyramid network and perform the following operations:
[0020] Extract the hierarchical feature maps of each frame through the backbone network to generate a feature pyramid of four scales P2 - P5. Deploy deformable convolutional layers at the P3 - P5 levels to adaptively match the geometric deformation features of the text region, and output multi-scale feature maps based on the matching.
[0021] Based on the output multi-scale feature maps, generate initial text candidate boxes at the P3 level, and use the candidate box regression network to adjust the coordinates; upsample the feature maps at the P4 and P5 levels, fuse them with the P3 level features, and then screen the final candidate boxes through the non-maximum suppression algorithm.
[0022] Perform refined processing on the screened final candidate boxes: Use the mask branch network to predict the binary text mask, and the mask branch network consists of 3 convolutional layers of 3×3 and 1 bilinear interpolation upsampling layer.
[0023] Based on the temporal continuity of the first video frame sequence, optical flow estimation is performed on the candidate boxes of adjacent frames, the displacement vectors of the text regions in the front and back frames are calculated, and the displacement vectors are fitted through a trajectory smoothing algorithm to generate a motion trajectory prediction map, where the trajectory map contains the coordinate sequence and velocity vector of the text center point; the velocity vector is normalized to obtain the text motion speed value for subsequent weight assignment.
[0024] As a preferred solution of the present invention, spatio-temporal dual-domain feature extraction is performed on the preprocessed first video frame sequence to obtain the first spatio-temporal feature, specifically including:
[0025] Based on the generated initial text mask and multi-scale feature map, spatial domain feature extraction is performed, and deformable convolutional layers are deployed within the text region. The convolutional kernel contains K offset sampling points, and the offset is calculated through the following formula:
[0026] Δp k = W k · F text
[0027] In the formula, F text is the text region feature map, W k is the learnable weight matrix, and Δp k is the coordinate offset of the k-th sampling point.
[0028] The background edge features are extracted using the offset sampling points, and the spatial features at the P3 level are cross-scale fused with the context features at the P4 and P5 levels through a 3-layer residual convolutional network to generate a spatial domain feature map. Based on the generated motion trajectory prediction map, temporal domain feature extraction is performed.
[0029] Time series construction: The text velocity vectors in the trajectory prediction map and the displacement vectors of adjacent frames are arranged in chronological order to form a time series with a length of 2M + 1.
[0030] Bidirectional LSTM modeling: The time series is input into a bidirectional LSTM network, which contains two hidden layers, each with a dimension of 128, and the temporal dependence relationship is calculated through forget gates, input gates, and output gates.
[0031] Dynamic weight generation: The forward and backward hidden states output by the LSTM are concatenated and mapped to a dynamic propagation weight matrix through a fully connected layer, and the matrix dimension is consistent with the spatial domain feature Figure 1 consistent.
[0032] The spatial domain feature map and the dynamic propagation weight matrix are multiplied element by element to generate the fused first spatio-temporal feature, and the calculation formula is as follows:
[0033] F ST = F spatial⊙σ(W temporal )
[0034] where σ is the Sigmoid activation function, used to constrain the weight range to [0, 1], F ST represents the first spatio-temporal feature after fusion, W temporal represents the dynamic propagation weight matrix, F spatial represents the spatial domain feature map.
[0035] As a preferred solution of the present invention, a spatio-temporal dual-domain attention model is established based on the first spatio-temporal feature. The spatio-temporal dual-domain attention model calculates and fuses the spatial domain and temporal domain features through cross-attention, and dynamically allocates fusion weights based on the text movement speed, specifically including:
[0036] Taking the spatial domain feature map as the query vector, the temporal domain dynamic propagation weight matrix as the key-value and content vector; using the multi-head attention mechanism to calculate the correlation weights between the spatial domain and the temporal domain, where the dimension of each attention head is a preset value, and normalizing the weights through the Softmax function; merging the calculation results of multiple attention heads to generate an initial spatio-temporal correlation weight matrix.
[0037] Based on the normalized text movement speed value, perform the following operations:
[0038] When the text movement speed is lower than the threshold α, set the spatial domain weight ratio to be greater than 70%.
[0039] When the text movement speed is higher than or equal to the threshold α, set the temporal domain weight ratio to be greater than 60%.
[0040] Convert the speed value into the proportional coefficient of the spatial domain and temporal domain weights through a mapping function, and the mapping function includes a trainable speed-weight mapping coefficient and a bias term.
[0041] Perform weighted fusion of the dynamic weight allocation rule and the initial spatio-temporal correlation weight matrix to generate the final weight parameters of the spatio-temporal dual-domain attention model, and store them in the model database.
[0042] As a preferred solution of the present invention, obtain the second video frame sequence of the video stream to be processed and perform preprocessing to extract the second spatio-temporal feature, specifically including:
[0043] Receiving the video stream to be processed in real time, extracting continuous T frames as the second video frame sequence, including the current processing frame and its adjacent front and rear frames, eliminating the pixel offset caused by camera movement or object displacement between adjacent frames through time alignment, and outputting the aligned second video frame sequence.
[0044] Input the aligned second video frame sequence into a multi-scale feature pyramid network to generate text region candidate boxes and initial text masks for each frame, and calculate the displacement vectors and normalized velocity values of the text regions in the second video frame sequence based on the motion trajectory prediction map generation method.
[0045] According to the initial text mask, use the spatial domain feature extraction method to extract the spatial domain feature map of the second video frame sequence; based on the text displacement vectors, generate the dynamic propagation weight matrix of the second video frame sequence through the time domain feature extraction method; multiply the spatial domain feature map and the dynamic propagation weight matrix element by element according to the fusion rule to generate the second spatio-temporal feature.
[0046] As a preferred solution of the present invention, generating a traceless erasure video stream based on the repaired second video frame specifically includes:
[0047] Perform frequency domain separation on the text erasure region in the repaired second video frame to extract low-frequency structure information and high-frequency texture details.
[0048] Compare the spectrum of the repaired region with the spectrum of the corresponding region of the original video. If the spectrum distribution error exceeds the distribution threshold, adjust the frequency components of the repaired region through the frequency domain compensation algorithm.
[0049] Arrange the second video frames after dynamic frequency constraint processing in the time order of the original video stream, and perform SSIM similarity verification on the repaired regions of adjacent frames. If the SSIM values of consecutive L frames are all higher than the similarity threshold, then retain the synthesis result; otherwise, trigger the sliding window averaging method for local smoothing.
[0050] Encode the video frame sequence after temporal coherence synthesis into a complete video stream, remove all text mask marks and temporary data, and output a text-free video stream with the same resolution and frame rate as the original video.
[0051] This application further provides an intelligent traceless erasure system for video screen text, including the following:
[0052] A preprocessing unit that obtains the first video frame sequence of the video stream to be processed and performs preprocessing.
[0053] A feature extraction unit that performs spatio-temporal dual-domain feature extraction on the preprocessed first video frame sequence to obtain the first spatio-temporal feature.
[0054] A model construction unit for establishing a spatio-temporal dual-domain attention model according to the first spatio-temporal feature. The spatio-temporal dual-domain attention model calculates and fuses spatial domain and time domain features through cross-attention and dynamically allocates fusion weights based on the text motion speed.
[0055] The secondary extraction unit is used to obtain the second video frame sequence of the video stream to be processed, perform preprocessing, and extract the second spatio-temporal features.
[0056] The matching unit is used to input the second spatio-temporal features into the spatio-temporal dual-domain attention model to execute the matching process.
[0057] The matching process includes main matching and auxiliary determination.
[0058] The main determination unit is used for main matching. According to the dynamic propagation weight matrix, it matches the second spatio-temporal features with the spatio-temporal dual-domain features in the model. If the matching degree is higher than the matching threshold, the repaired second video frame is directly generated.
[0059] The auxiliary determination unit is used for auxiliary determination. If the matching degree is lower than the matching threshold, it triggers the temporal consistency compensation mechanism, calculates the SSIM similarity of the repaired regions of adjacent frames. When the SSIM is lower than the auxiliary determination threshold, it smooths the temporal mutation and re-repairs through the sliding window averaging method to generate the corrected second video frame.
[0060] The output unit is used to generate a seamless erasure video stream based on the repaired second video frame.
[0061] The beneficial effects of this application are as follows:
[0062] 1. By using the spatio-temporal dual-domain attention model to synchronously reference the texture and motion features of the front and back frames, the temporal coherence of the repaired region is improved. Combining with the SSIM-driven temporal consistency compensation mechanism, the flicker frequency is reduced. Through the dynamic weight allocation strategy, the spatio-temporal domain weights are adaptively adjusted according to the text speed to support stable erasure of fast-moving text. At the same time, according to the deformable convolution and trajectory smoothing algorithm, the text deformation and motion trajectory are captured, reducing the smear residue rate in the repaired region.
[0063] 2. By using dynamic frequency constraints and the sliding window averaging method to reduce redundant calculations, the number of model parameters is reduced, the inference speed is accelerated, and through the frequency domain compensation algorithm, the spectral distribution error of the repaired region is controlled within ≤3dB, improving the high-frequency detail reconstruction quality.
[0064] To make the above objects, features, and advantages of this application more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0065] To more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required in the embodiments of this application. It should be understood that the following drawings only show some embodiments of this application, so they should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.
[0066] Figure 1 This is a flowchart of the intelligent seamless erasure method for video frame text provided by the embodiments of the present application.
[0067] Figure 2 This is an overall framework diagram of the intelligent seamless erasure system for video frame text provided by the embodiments of the present application. Detailed implementation manners
[0068] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.
[0069] Please refer to Figure 1 , Figure 1 This is a flowchart of the intelligent seamless erasure method for video frame text provided by the embodiments of the present application.
[0070] In this embodiment, the intelligent seamless erasure method for video frame text may include step S10, step S20, step S30, step S40, step S50, and step S60.
[0071] Step S10: Obtain the first video frame sequence of the video stream to be processed and perform preprocessing, specifically including:
[0072] The first video frame sequence includes the current frame and its adjacent front and rear frames. The preprocessing includes detecting the text area through a multi-scale feature pyramid network, generating an initial text mask and a motion trajectory prediction map, extracting N consecutive frames from the video stream to be processed as the first video frame sequence, including the current frame, the first M frames, and the last M frames, and eliminating the offset error caused by camera movement between adjacent frames through a temporal alignment module.
[0073] It should be noted that N≥3 and M≥1.
[0074] Input the first video frame sequence into the multi-scale feature pyramid network and perform the following operations:
[0075] Extract the hierarchical feature maps of each frame through the backbone network, generate a feature pyramid of four scales P2 - P5, deploy deformable convolutional layers at the P3 - P5 levels, adaptively match the geometric deformation features of the text area, and output a multi-scale feature map based on the matching.
[0076] It should be noted that P2 has the highest resolution, followed by P3, and so on to P5.
[0077] Based on the output multi-scale feature map, generate an initial text candidate box at the P3 level, and use the candidate box regression network to adjust the coordinates; upsample the feature maps at the P4 and P5 levels, fuse them with the P3-level features, and then screen the final candidate boxes through the non-maximum suppression algorithm.
[0078] It should be noted that the confidence threshold here is set to θ (θ≥0.7).
[0079] Perform refinement processing on the finally selected candidate boxes: use the mask branch network to predict the binary text mask, and the mask branch network is composed of 3 3×3 convolutional layers and 1 bilinear interpolation upsampling layer.
[0080] It should be noted that the L1 loss constraint is applied to the mask edge area to ensure that the text boundary positioning error is less than 2 pixels.
[0081] Based on the temporal continuity of the first video frame sequence, perform optical flow estimation on the candidate boxes of adjacent frames, calculate the displacement vectors of the text regions in the front and back frames, fit the displacement vectors through the trajectory smoothing algorithm, and generate a motion trajectory prediction map. The trajectory map contains the coordinate sequence and velocity vector of the text center points; perform normalization processing on the velocity vectors to obtain the text motion speed values for subsequent weight assignment.
[0082] Step S20: Perform spatio-temporal dual-domain feature extraction on the preprocessed first video frame sequence to obtain the first spatio-temporal features, specifically including:
[0083] Based on the generated initial text mask and multi-scale feature maps, perform spatial-domain feature extraction, deploy deformable convolutional layers inside the text regions, and the convolutional kernel contains K deformable sampling points. It should be noted that K≥9, and the offset is calculated through the following formula:
[0084] Δp k =W k ·F text
[0085] In the formula, F text is the text region feature map, W k is the learnable weight matrix, and Δp k is the coordinate offset of the kth sampling point.
[0086] Extract background edge features using the offset sampling points, and perform cross-scale fusion of the spatial features at the P3 level and the context features at the P4 and P5 levels through a 3-layer residual convolutional network to generate a spatial-domain feature map.
[0087] It should be noted that each layer of the residual convolutional network contains 64 3×3 convolutional kernels.
[0088] Based on the generated motion trajectory prediction map, perform temporal-domain feature extraction.
[0089] Time series construction: Arrange the text velocity vectors in the trajectory prediction map and the displacement vectors of adjacent frames in chronological order to form a time series with a length of 2M+1.
[0090] Bidirectional LSTM modeling: Input the time series into a bidirectional LSTM network, which consists of two hidden layers with a dimension of 128 for each layer. Calculate the temporal dependence relationship through forget gates, input gates, and output gates.
[0091] Dynamic weight generation: Concatenate the forward and backward hidden states output by the LSTM, and map them to a dynamic propagation weight matrix through a fully connected layer. The dimension of the matrix is consistent with the spatial domain features Figure 1 Consistent.
[0092] Multiply the spatial domain feature map element-wise with the dynamic propagation weight matrix to generate the first spatio-temporal feature after fusion. The calculation formula is as follows:
[0093] F ST = F spatial ⊙ σ(W temporal )
[0094] In the formula, σ is the Sigmoid activation function, which is used to constrain the weight range to [0, 1]. F ST represents the first spatio-temporal feature after fusion, W temporal represents the dynamic propagation weight matrix, and F spatial represents the spatial domain feature map.
[0095] Step S30: Establish a spatio-temporal dual-domain attention model based on the first spatio-temporal feature. The spatio-temporal dual-domain attention model calculates the fusion of spatial domain and time domain features through cross-attention and dynamically allocates fusion weights based on the text movement speed. Specifically, it includes:
[0096] Use the spatial domain feature map as the query vector, and the time domain dynamic propagation weight matrix as the key and content vectors; adopt the multi-head attention mechanism to calculate the correlation weights between the spatial domain and the time domain, where the dimension of each attention head is a preset value, and normalize the weights through the Softmax function; merge the calculation results of multiple attention heads to generate the initial spatio-temporal correlation weight matrix.
[0097] Based on the normalized text movement speed value, perform the following operations:
[0098] When the text movement speed is lower than the threshold α, set the spatial domain weight ratio to be greater than 70%.
[0099] When the text movement speed is higher than or equal to the threshold α, set the time domain weight ratio to be greater than 60%.
[0100] Convert the speed value into the proportional coefficients of the spatial domain and time domain weights through a mapping function, and the mapping function includes trainable speed-weight mapping coefficients and bias terms.
[0101] The dynamic weight assignment rule is weighted and fused with the initial spatio-temporal correlation weight matrix to generate the final weight parameters of the spatio-temporal dual-domain attention model, which are stored in the model database.
[0102] It should be noted that the threshold α is set as follows:
[0103] Statistically analyze the distribution of the motion speeds of all text regions in the training dataset, and select the maximum speed value covering 95% of the samples as the initial threshold.
[0104] Adjust the initial threshold on the validation set, and select the one that meets the following conditions as the final threshold α:
[0105] The decline rate of the flicker frequency exceeds the preset index.
[0106] The mis-erasion rate of non-text regions is lower than 5%.
[0107] Set the value range of α to be from 0.3 to 0.5, and the unit is the same as the normalized speed value.
[0108] Step S40: Obtain the second video frame sequence of the video stream to be processed and perform preprocessing to extract the second spatio-temporal features, specifically including:
[0109] Receive the video stream to be processed in real time, extract continuous T frames as the second video frame sequence, including the current frame being processed and its adjacent front and back frames, eliminate the pixel offset between adjacent frames caused by camera movement or object displacement through time alignment, and output the aligned second video frame sequence.
[0110] It should be noted that T≥3.
[0111] Input the aligned second video frame sequence into the multi-scale feature pyramid network to generate candidate bounding boxes for text regions and initial text masks for each frame, and calculate the displacement vectors and normalized speed values of text regions in the second video frame sequence based on the motion trajectory prediction map generation method.
[0112] According to the initial text mask, use the spatial domain feature extraction method to extract the spatial domain feature map of the second video frame sequence; based on the text displacement vector, generate the dynamic propagation weight matrix of the second video frame sequence through the time domain feature extraction method; multiply the spatial domain feature map and the dynamic propagation weight matrix element by element according to the fusion rule to generate the second spatio-temporal feature.
[0113] Step S50: Input the second spatio-temporal feature into the spatio-temporal dual-domain attention model to execute the matching process.
[0114] The matching process includes main matching and auxiliary determination.
[0115] Primary matching: According to the dynamic propagation weight matrix, match the second spatio-temporal feature with the spatio-temporal dual-domain feature in the model. If the matching degree is higher than the matching threshold, directly generate the repaired second video frame.
[0116] Auxiliary determination: If the matching degree is lower than the matching threshold, trigger the temporal consistency compensation mechanism, calculate the SSIM similarity of the repaired areas of adjacent frames. When the SSIM is lower than the auxiliary determination threshold, smooth the temporal mutation and re-repair through the sliding window averaging method to generate the corrected second video frame.
[0117] Step S60: Generate a seamless erasure video stream based on the repaired second video frame, specifically including:
[0118] Perform frequency domain separation on the text erasure area in the repaired second video frame, and extract low-frequency structure information and high-frequency texture details.
[0119] Compare the spectrum of the repaired area with the spectrum of the corresponding area of the original video. If the spectrum distribution error exceeds the distribution threshold, adjust the frequency components of the repaired area through the frequency domain compensation algorithm.
[0120] It should be noted that the distribution threshold ≤ 3dB.
[0121] Arrange the second video frames after dynamic frequency constraint processing in the time order of the original video stream, verify the SSIM similarity of the repaired areas of adjacent frames. If the SSIM values of consecutive L frames are all higher than the similarity threshold, retain the synthesis result; otherwise, trigger the sliding window averaging method for local smoothing.
[0122] It should be noted that in the training stage, count the SSIM similarity of all repaired frames and non-text original background areas, select the 95% percentile value as the initial similarity threshold, and test the influence of different similarity thresholds on the flicker frequency and mis-erasure rate on the validation set, and select the final similarity threshold that simultaneously meets the following conditions:
[0123] The flicker frequency reduction rate ≥ 85%; that is, the number of flicker frames after repair / the number of original flicker frames ≤ 0.15.
[0124] The mis-erasure rate of non-text areas ≤ 3%.
[0125] Value range limitation: Set γ ∈ [0.88, 0.92], γ represents the final similarity threshold, and the unit is the same as the SSIM index.
[0126] Encode the video frame sequence after temporal coherence synthesis into a complete video stream, remove all text mask marks and temporary data, and output a text-free video stream with the same resolution and frame rate as the original video.
[0127] So far, the intelligent seamless erasure method for video frame text is completed.
[0128] Please also refer to Figure 2 , Figure 2 An intelligent seamless erasure system for video frame text is provided, including the following:
[0129] A preprocessing unit that obtains the first video frame sequence of the video stream to be processed and performs preprocessing.
[0130] A feature extraction unit that performs spatio-temporal dual-domain feature extraction on the preprocessed first video frame sequence to obtain the first spatio-temporal feature.
[0131] A model construction unit that is used to establish a spatio-temporal dual-domain attention model based on the first spatio-temporal feature. The spatio-temporal dual-domain attention model fuses spatial-domain and time-domain features through cross-attention calculation and dynamically allocates fusion weights based on the text movement speed.
[0132] A secondary extraction unit that is used to obtain the second video frame sequence of the video stream to be processed and perform preprocessing, and extract the second spatio-temporal feature.
[0133] A matching unit that is used to input the second spatio-temporal feature into the spatio-temporal dual-domain attention model to execute the matching process.
[0134] The matching process includes main matching and auxiliary determination.
[0135] A main determination unit for main matching. According to the dynamic propagation weight matrix, it matches the second spatio-temporal feature with the spatio-temporal dual-domain feature in the model. If the matching degree is higher than the matching threshold, the repaired second video frame is directly generated.
[0136] An auxiliary determination unit for auxiliary determination. If the matching degree is lower than the matching threshold, it triggers the temporal consistency compensation mechanism, calculates the SSIM similarity of the repaired regions of adjacent frames. When the SSIM is lower than the auxiliary determination threshold, it smooths the temporal mutation through the sliding window averaging method and re-repairs to generate the corrected second video frame.
[0137] An output unit that is used to generate a seamless erasure video stream based on the repaired second video frame.
[0138] So far, the intelligent seamless erasure system for video frame text is completed.
[0139] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the embodiments of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the embodiments of the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the embodiments of the present invention. Therefore, the protection scope of the embodiments of the present invention should be subject to the protection scope of the claims.
Claims
1. An intelligent seamless erasure method for text in video frames, characterized in that, It includes the following: Obtain the first video frame sequence of the video stream to be processed and perform preprocessing; Extract spatio-temporal dual-domain features from the preprocessed first video frame sequence to obtain first spatio-temporal features; Establish a spatio-temporal dual-domain attention model according to the first spatio-temporal features. The spatio-temporal dual-domain attention model fuses spatial-domain and temporal-domain features through cross-attention calculation and dynamically assigns fusion weights based on the text movement speed; Obtain the second video frame sequence of the video stream to be processed and perform preprocessing, and extract second spatio-temporal features; Input the second spatio-temporal features into the spatio-temporal dual-domain attention model to execute the matching process; The matching process includes main matching and auxiliary determination; Main matching: According to the dynamic propagation weight matrix, match the second spatio-temporal features with the spatio-temporal dual-domain features in the model. If the matching degree is higher than the matching threshold, directly generate the repaired second video frame; Auxiliary determination: If the matching degree is lower than the matching threshold, trigger the temporal consistency compensation mechanism, calculate the SSIM similarity of the repaired regions of adjacent frames. When the SSIM is lower than the auxiliary determination threshold, smooth the temporal mutation by the sliding window averaging method and re-repair to generate the corrected second video frame; Generate a seamless erasure video stream based on the repaired second video frame.
2. The intelligent seamless erasing method for video picture text according to claim 1, characterized in that, Obtain the first video frame sequence of the video stream to be processed and perform preprocessing, specifically including: The first video frame sequence includes the current frame and its adjacent front and rear frames. The preprocessing includes detecting the text region through a multi-scale feature pyramid network to generate an initial text mask and a motion trajectory prediction map. Extract N consecutive frames from the video stream to be processed as the first video frame sequence, which includes the current frame, the first M frames, and the last M frames, and eliminate the offset error caused by camera movement between adjacent frames through a time alignment module; Input the first video frame sequence into the multi-scale feature pyramid network and perform the following operations: Extract the hierarchical feature maps of each frame through the backbone network to generate a feature pyramid of four scales P2 - P5. Deploy deformable convolutional layers at the P3 - P5 levels to adaptively match the geometric deformation features of the text region, and output multi-scale feature maps based on the matching; Based on the output multi-scale feature maps, generate initial text candidate boxes at the P3 level, and use the candidate box regression network to adjust the coordinates; Upsample the feature maps at the P4 and P5 levels, fuse them with the P3-level features, and then screen the final candidate boxes through the non-maximum suppression algorithm; Perform refined processing on the screened final candidate boxes: Use the mask branch network to predict the binary text mask. The mask branch network consists of 3 3×3 convolutional layers and 1 bilinear interpolation upsampling layer; Based on the time continuity of the first video frame sequence, perform optical flow estimation on the candidate boxes of adjacent frames, calculate the displacement vectors of the text regions in the front and rear frames, fit the displacement vectors through the trajectory smoothing algorithm to generate a motion trajectory prediction map. The trajectory map contains the coordinate sequence and velocity vector of the text center point; Normalize the velocity vector to obtain the text movement speed value for subsequent weight assignment.
3. The intelligent seamless erasing method for video picture text according to claim 1, characterized in that, Extract spatio-temporal dual-domain features from the preprocessed first video frame sequence to obtain first spatio-temporal features, specifically including: Based on the generated initial text mask and multi-scale feature maps, perform spatial domain feature extraction, deploy a deformable convolutional layer inside the text region, where the convolutional kernel contains K offset sampling points, and calculate the offset through the following formula: Δp k = W k ·F text Where, F text is the text region feature map, W k is the learnable weight matrix, and Δp k is the coordinate offset of the k-th sampling point; Extract background edge features using the offset sampling points, perform cross-scale fusion of the spatial features at the P3 level and the context features at the P4 and P5 levels through a 3-layer residual convolutional network to generate a spatial domain feature map, and based on the generated motion trajectory prediction map, perform temporal domain feature extraction: Time series construction: Arrange the text velocity vectors in the trajectory prediction map and the displacement vectors of adjacent frames in chronological order to form a time series of length 2M + 1; Bidirectional LSTM modeling: Input the time series into a bidirectional LSTM network, which contains two hidden layers, each with a dimension of 128, and calculate the temporal dependence relationship through forget gates, input gates, and output gates; Dynamic weight generation: Concatenate the forward and backward hidden states output by the LSTM, and map them to a dynamic propagation weight matrix through a fully connected layer, where the matrix dimension is the same as that of the spatial domain feature map; Multiply the spatial domain feature map and the dynamic propagation weight matrix element by element to generate the first fused spatio-temporal feature, and the calculation formula is as follows: F ST = F spatial ⊙σ(W temporal ) where σ is the Sigmoid activation function, which is used to constrain the weight range to [0, 1], and F ST represents the first spatio-temporal feature after fusion, and W temporal represents the dynamic propagation weight matrix, and F spatial represents the spatial domain feature map.
4. The intelligent seamless erasing method for video picture text according to claim 1, characterized in that, Establish a spatio-temporal dual-domain attention model based on the first spatio-temporal feature. The spatio-temporal dual-domain attention model calculates the fusion of spatial domain and temporal domain features through cross-attention and dynamically assigns fusion weights based on the text motion speed, specifically including: Use the spatial domain feature map as the query vector, and the temporal domain dynamic propagation weight matrix as the key-value and content vector; calculate the correlation weights between the spatial domain and the temporal domain using the multi-head attention mechanism, where the dimension of each attention head is a preset value, and normalize the weights through the Softmax function; merge the calculation results of multiple attention heads to generate an initial spatio-temporal correlation weight matrix; Based on the normalized text motion speed value, perform the following operations: When the text motion speed is lower than the threshold α, set the spatial domain weight ratio to be greater than 70%; When the text motion speed is higher than or equal to the threshold α, set the temporal domain weight ratio to be greater than 60%; Convert the speed value to the proportional coefficients of the spatial domain and temporal domain weights through a mapping function, and the mapping function contains trainable speed-weight mapping coefficients and bias terms; Perform weighted fusion of the dynamic weight allocation rule and the initial spatio-temporal correlation weight matrix to generate the final weight parameters of the spatio-temporal dual-domain attention model, and store them in the model database.
5. The intelligent seamless erasing method for video picture text according to claim 1, characterized in that, Obtain the second video frame sequence of the video stream to be processed and perform preprocessing to extract the second spatio-temporal feature, specifically including: Receive the video stream to be processed in real time, extract continuous T frames as the second video frame sequence, including the current processing frame and its adjacent front and back frames, eliminate the pixel offset caused by camera movement or object displacement between adjacent frames through time alignment, and output the aligned second video frame sequence; Input the aligned second video frame sequence into a multi-scale feature pyramid network to generate candidate text region boxes and initial text masks for each frame, and calculate the displacement vector and normalized speed value of the text region in the second video frame sequence based on the motion trajectory prediction map generation method; According to the initial text mask, use the spatial domain feature extraction method to extract the spatial domain feature map of the second video frame sequence; based on the text displacement vector, generate the dynamic propagation weight matrix of the second video frame sequence through the temporal domain feature extraction method; multiply the spatial domain feature map and the dynamic propagation weight matrix element by element according to the fusion rule to generate the second spatio-temporal feature.
6. The intelligent seamless erasing method for video picture text according to claim 1, characterized in that, Generate a traceless erasure video stream based on the repaired second video frame, specifically including: Perform frequency domain separation on the text erasure area in the repaired second video frame, and extract low-frequency structure information and high-frequency texture details; Compare the spectrum of the repaired area with the spectrum of the corresponding area of the original video. If the spectrum distribution error exceeds the distribution threshold, adjust the frequency components of the repaired area through the frequency domain compensation algorithm; Arrange the second video frames after dynamic frequency constraint processing in the time order of the original video stream, and perform SSIM similarity verification on the repaired areas of adjacent frames. If the SSIM values of consecutive L frames are all higher than the similarity threshold, retain the synthesis result; otherwise, trigger the sliding window averaging method for local smoothing; Encode the video frame sequence after temporal coherence synthesis into a complete video stream, remove all text mask marks and temporary data, and output a text-free video stream with the same resolution and frame rate as the original video.
7. An intelligent seamless erasure system for video frame text, characterized in that, Specifically including: A preprocessing unit that obtains the first video frame sequence of the video stream to be processed and performs preprocessing; A feature extraction unit that performs spatio-temporal dual-domain feature extraction on the preprocessed first video frame sequence to obtain the first spatio-temporal feature; A model construction unit for establishing a spatio-temporal dual-domain attention model according to the first spatio-temporal feature. The spatio-temporal dual-domain attention model fuses spatial domain and temporal domain features through cross-attention calculation and dynamically allocates fusion weights based on the text movement speed; A secondary extraction unit for obtaining the second video frame sequence of the video stream to be processed and performing preprocessing, and extracting the second spatio-temporal feature; A matching unit for inputting the second spatio-temporal feature into the spatio-temporal dual-domain attention model to execute the matching process; The matching process includes main matching and auxiliary determination; A main determination unit for main matching, which matches the second spatio-temporal feature with the spatio-temporal dual-domain feature in the model according to the dynamic propagation weight matrix. If the matching degree is higher than the matching threshold, directly generate the repaired second video frame; An auxiliary determination unit for auxiliary determination. If the matching degree is lower than the matching threshold, trigger the temporal consistency compensation mechanism, calculate the SSIM similarity of the repaired areas of adjacent frames. When the SSIM is lower than the auxiliary determination threshold, smooth the temporal mutation and re-repair through the sliding window averaging method to generate the corrected second video frame; An output unit for generating a traceless erasure video stream based on the repaired second video frame.
Citation Information
Patent Citations
End-to-end behavior recognition method and system based on self-adaptive space-time attention mechanism
CN111401177A
Controllable video generation method and system based on multi-modal fusion
CN119091362A
System and method for human action recognition and intensity indexing from video stream using fuzzy attention machine learning
US20210312183A1
Video text tracking method and electronic device
WO2023115838A1
Cited By
Intelligent cherry irrigation decision-making method based on multi-mode and space-time prediction
CN121525882A
Subtitle elimination method and device, electronic equipment and storage medium
CN121864926A
A subtitle elimination method and device, electronic equipment and storage medium
CN121864926B