A method, apparatus, device and medium for erasing video subtitles
By acquiring the fidelity stream and computational stream of the video, the target area is detected and structural texture restoration is performed to generate a high-quality video after subtitle erasure. This solves the problem of unstable video subtitle erasure in existing technologies and improves computational efficiency and video quality.
Patent Information
- Application Number
- CN202610773374.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-07-03
AI Technical Summary
Existing technologies cannot efficiently obtain high-quality video after subtitle erasure. Traditional methods suffer from problems such as inconsistent textures between video frames and screen flickering. Methods based on deep learning models have high computational complexity and cause memory overflow in high-definition videos.
By acquiring the fidelity stream and computational stream of the video to be processed, the target detection region is detected, an initial repair mask is obtained and structural texture repair is performed, and the fidelity stream and repair mask are used to generate a video after subtitle erasure, the video resolution is reduced and upsampling is performed.
It improves the computational efficiency of the video after subtitle erasure under the premise of low video resolution, reduces the computational cost, avoids image quality degradation, and improves video quality.
Smart Images

Figure CN122340306A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, and in particular to a method, apparatus, device, and medium for erasing video subtitles. Background Technology
[0002] With the rapid development of multimedia technology, the demand for video content editing and post-processing continues to rise. Among these, the editing and modification of local pixels in video content, such as erasing subtitles or restoring background pixels obscured by subtitles, is particularly important.
[0003] In existing technologies, on the one hand, traditional video subtitle erasure methods, such as those based on image block matching or single-frame generative adversarial networks, can be used to erase the objects to be erased frame by frame in a video frame. On the other hand, deep learning model-based video subtitle erasure methods, such as those based on Transformer models or diffusion models, can be used to erase the objects to be erased in a video frame.
[0004] However, traditional video subtitle erasure methods do not fully consider the temporal correlation between video frames, resulting in temporal instability such as texture inconsistencies, screen flickering, or jitter in the area to be erased during video playback. Deep learning-based video subtitle erasure methods, due to their high computational complexity and memory consumption, suffer from memory overflow issues when processing high-definition video, failing to meet the stability and efficiency requirements of real-world production environments. In summary, existing technologies cannot efficiently obtain high-quality video with erased subtitles. Summary of the Invention
[0005] The purpose of this invention is to provide a video subtitle erasure method to solve the problem in the prior art that it is impossible to efficiently obtain high-quality video after subtitle erasure.
[0006] To achieve the above objectives, in a first aspect, embodiments of the present invention provide a video subtitle erasure method, the method comprising: Obtain the video fidelity stream and video computation stream corresponding to the video to be processed, wherein the video computation stream consists of multiple video computation frames; each video computation frame includes: a target detection region; For multiple video calculation frames to be processed, detect whether there is a target subtitle to be erased in each target detection region. If so, obtain the initial repair mask of the target subtitle to be erased. Structural texture restoration is performed on multiple initial repair masks to obtain target repair masks corresponding to each of the multiple initial repair masks; Based on the fidelity stream of the video to be processed, multiple initial repair masks, and the target repair masks corresponding to the multiple initial repair masks, obtain the subtitle-erased video corresponding to the video to be processed.
[0007] In one embodiment, obtaining the fidelity stream and computational stream of the video to be processed corresponding to the video to be processed includes: Obtain the pixel value of the long side of the video resolution corresponding to the video to be processed; Determine whether the pixel value of the longer side is greater than a preset threshold; If the pixel value of the longer side is greater than a preset threshold, the video to be processed is downsampled to obtain the video computation stream and the video to be processed is determined to be the video fidelity stream.
[0008] In one embodiment, the method further includes: If the pixel value of the longer side is not greater than a preset threshold, the video to be processed is determined to be both the video computation stream and the video fidelity stream.
[0009] In one embodiment, obtaining the initial repair mask for the target subtitle to be erased includes: The target subtitle to be erased is binarized to obtain the binarized mask of the target subtitle to be erased; The binarized mask is subjected to morphological dilation to obtain the initial repair mask for the target subtitle to be erased.
[0010] In one embodiment, performing structural texture restoration on a plurality of initial restoration masks to obtain target restoration masks corresponding to the plurality of initial restoration masks includes: Perform structural repair processing on the multiple initial repair masks to obtain the structural repair masks corresponding to the multiple initial repair masks respectively; Texture restoration processing is performed on multiple structural restoration masks to obtain target restoration masks corresponding to the multiple initial restoration masks.
[0011] In one embodiment, obtaining the subtitle-erased video corresponding to the video to be processed based on the video fidelity stream to be processed, multiple initial repair masks, and target repair masks corresponding to the multiple initial repair masks respectively includes: Based on the video resolution of the video fidelity stream to be processed, the multiple target restoration masks are upsampled respectively to obtain multiple upsampled target restoration masks; Gaussian blurring is applied to multiple initial repair masks to obtain the soft masks corresponding to each initial repair mask; Based on the fidelity stream of the video to be processed, multiple upsampled target repair masks, and multiple soft masks, obtain the subtitle-erased video corresponding to the video to be processed.
[0012] In one embodiment, the video fidelity stream to be processed consists of multiple video fidelity frames to be processed, each video fidelity frame to be processed corresponding one-to-one with each video computation frame to be processed. The step of obtaining the subtitle-erased video corresponding to the video to be processed based on the video fidelity stream to be processed, multiple upsampling target restoration masks, and multiple soft masks includes: For multiple video calculation frames to be processed that contain target subtitles to be erased, the target video fidelity frame corresponding to each video calculation frame to be processed is determined among the multiple video fidelity frames to be processed. Alpha mixing is performed on multiple target video fidelity frames to be processed, multiple upsampled target repair masks corresponding to the target subtitles to be erased, and multiple soft masks to obtain the subtitle-erased video corresponding to the video to be processed.
[0013] Secondly, embodiments of the present invention provide a video subtitle erasing device, the device comprising: The video processing module is used to acquire the video fidelity stream and the video computation stream corresponding to the video to be processed, wherein the video computation stream consists of multiple video computation frames; each video computation frame includes: a target detection region; The initial repair mask acquisition module is used to calculate frames for multiple videos to be processed, detect whether there are target subtitles to be erased in each target detection area, and if so, acquire the initial repair mask of the target subtitles to be erased; The target repair mask acquisition module is used to perform structural texture repair on multiple initial repair masks and acquire the target repair mask corresponding to each of the multiple initial repair masks. The subtitle-erased video acquisition module is used to acquire the subtitle-erased video corresponding to the video to be processed based on the fidelity stream of the video to be processed, multiple initial repair masks, and target repair masks corresponding to the multiple initial repair masks.
[0014] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the video subtitle erasure method described in the first aspect.
[0015] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the video subtitle erasure method described in the first aspect.
[0016] The technical solution provided by the embodiments of the present invention has the following advantages compared with the prior art: This invention provides a video subtitle erasure method. It acquires a video fidelity stream and a video computation stream corresponding to the video to be processed. The video computation stream consists of multiple video computation frames, each including a target detection region. For each of the multiple video computation frames, it detects whether a target subtitle to be erased exists in each target detection region. If so, it acquires an initial repair mask for the target subtitle. It then performs structural texture repair on the multiple initial repair masks to acquire target repair masks corresponding to each initial repair mask. Based on the video fidelity stream, the multiple initial repair masks, and the target repair masks corresponding to each initial repair mask, it acquires the subtitle-erased video corresponding to the video to be processed. This method effectively improves the computational efficiency of acquiring the subtitle-erased video and reduces computational costs, even with a low-resolution video computation stream. Furthermore, by performing structural and texture restoration only on multiple initial restoration masks, the entire video calculation frame to be processed can be restored, thus solving the problem of image quality degradation caused by re-encoding or scaling the entire video calculation frame to be processed and improving the quality of the video after subtitle erasure. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 A flowchart illustrating a video subtitle erasure method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a video subtitle erasing device provided in an embodiment of the present invention. Detailed Implementation
[0018] To facilitate a clear description of the technical solutions in the embodiments of the present invention, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, the first threshold and the second threshold are merely used to distinguish different thresholds and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" are not necessarily different.
[0019] It should be noted that in this invention, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0020] In this invention, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between the associated objects, indicating that three relationships can exist.
[0021] like Figure 1 As shown, Figure 1 This is a flowchart illustrating a video subtitle erasure method provided in an embodiment of the present invention, which specifically includes the following steps: S10: Obtain the fidelity stream and computational stream of the video to be processed.
[0022] The video fidelity stream to be processed refers to a video with high video resolution, used to ensure that the subsequent video after subtitle erasure has high video resolution. The video fidelity stream to be processed consists of multiple video fidelity frames.
[0023] The video computation stream to be processed refers to a video with low resolution, used to reduce the computational complexity of obtaining the video after subtitle erasure. The video computation stream consists of multiple video computation frames; each video fidelity frame corresponds one-to-one with each video computation frame. Each video computation frame includes a target detection region. It should be noted that the video to be processed includes multiple video frames, and the target detection region refers to the area within these multiple video frames that needs to be detected to contain subtitles that need to be erased. The target subtitles to be erased are the subtitles that need to be removed.
[0024] Specifically, after obtaining the video to be processed, the corresponding video fidelity stream and video computation stream are obtained.
[0025] Optionally, based on the above embodiments, in some embodiments of the present invention, one implementation of S10 may be: S101: Obtain the long side pixel value of the video resolution corresponding to the video to be processed.
[0026] The longer side pixel value refers to the number of pixels contained in the longer side of the video frame to be processed.
[0027] Specifically, obtain the long side pixel value of the video resolution corresponding to the video to be processed.
[0028] S102: Determine whether the pixel value of the longer side is greater than the preset threshold.
[0029] The preset threshold refers to the value set to determine the fidelity stream and computational stream of the video to be processed. The preset threshold can be, for example, 720, but is not limited thereto. This invention does not impose specific limitations, and those skilled in the art can set it according to the actual situation.
[0030] Specifically, the obtained long-side pixel value is compared with a preset threshold to determine whether the long-side pixel value is greater than the preset threshold.
[0031] S103: If the pixel value of the long side is greater than the preset threshold, the video to be processed is downsampled to obtain the video computation stream and the video to be processed is determined to be the video fidelity stream.
[0032] Specifically, when it is determined that the pixel value of the long side of the video resolution corresponding to the video to be processed is greater than a preset threshold, the video to be processed is downsampled to obtain a low-resolution video processing stream, and at the same time, the video to be processed is determined to be a high-fidelity video stream.
[0033] Optionally, based on the above embodiments, in some embodiments of the present invention, another implementation of S10 may be: S104: If the pixel value of the longer side is not greater than the preset threshold, determine the video to be processed as the video computation stream and the video fidelity stream.
[0034] Specifically, when it is determined that the pixel value of the longer side of the video resolution corresponding to the video to be processed is not greater than a preset threshold, the video to be processed is determined to be both the video computation stream and the video fidelity stream.
[0035] Thus, this embodiment compares the pixel value of the long side of the video resolution with a preset threshold to obtain a high-resolution video fidelity stream that ensures the subtitle-erased video corresponding to the subsequent video to be processed has high video resolution, and reduces the computational complexity of obtaining the subtitle-erased video corresponding to the video to be processed, resulting in a low-resolution video computational stream.
[0036] It should be noted that, in order to reduce the amount of computation in the process of obtaining the video after the subtitles have been erased, the video computation stream to be processed is segmented to ensure that the subsequent calculations are performed using multiple segments included in the video computation stream.
[0037] S11: For multiple video frames to be processed, detect whether there are target subtitles to be erased in each target detection area. If so, obtain the initial repair mask of the target subtitles to be erased.
[0038] Specifically, for the multiple video computation frames included in the video computation stream to be processed, a text detection model is used to detect whether there are target subtitles to be erased in the target detection area included in each video computation frame. When it is determined that there are target subtitles to be erased in the target detection area, the initial repair mask of the target subtitles to be erased is obtained.
[0039] Optionally, based on the above embodiments, in some embodiments of the present invention, one way to obtain the initial repair mask of the target subtitle to be erased may be: S20: Perform binarization processing on the target subtitle to be erased to obtain the binarized mask of the target subtitle to be erased.
[0040] Specifically, for each video calculation frame containing the target subtitle to be erased, the target subtitle to be erased in the video calculation frame is binarized to obtain the binarized mask of the target subtitle to be erased.
[0041] S21: Perform morphological dilation on the binarized mask to obtain the initial repair mask for the target subtitle to be erased.
[0042] Morphological dilation is a fundamental operation in image processing, primarily used to expand the white (foreground) regions in an image. It is achieved through the sliding of structuring elements (convolution kernels) and the calculation of local maxima. In this embodiment, morphological dilation is applied to the binarized mask to cover any residual artifacts that may arise at the edges of the subtitles to be erased, thereby ensuring the acquisition of a high-quality video after subtitle erasure.
[0043] Specifically, after obtaining the binarized mask of the target subtitle to be erased, the binarized mask is further subjected to morphological dilation to obtain the initial repair mask of the target subtitle to be erased.
[0044] In this embodiment, the target subtitle to be erased is binarized to obtain a binarized mask. The binarized mask is then further subjected to morphological dilation to obtain an initial repair mask for the target subtitle to be erased. This covers any residual artifacts that may be generated at the edges of the target subtitle to be erased, thereby ensuring that a high-quality video after subtitle erasure can be obtained.
[0045] S12: Perform structural texture restoration on multiple initial restoration masks to obtain the target restoration mask corresponding to each initial restoration mask.
[0046] Among them, structural texture restoration refers to the restoration of the structural information and texture detail information of the initial restoration mask corresponding to the subtitle to be erased.
[0047] Specifically, after obtaining multiple initial repair masks, structural texture repair processing is performed on the multiple initial repair masks to obtain the target repair masks corresponding to the multiple initial repair masks.
[0048] Optionally, based on the above embodiments, in some embodiments of the present invention, S12 may be implemented as follows: S121: Perform structural repair processing on multiple initial repair masks to obtain the structural repair masks corresponding to the multiple initial repair masks respectively.
[0049] Specifically, after obtaining multiple initial repair masks, structural repair processing is performed on the multiple initial repair masks to obtain the structural repair masks corresponding to the multiple initial repair masks.
[0050] Optionally, based on the above embodiments, in some embodiments of the present invention, one way to perform structural repair processing on multiple initial repair masks is to perform structural repair processing on multiple initial repair masks through an optical flow estimation network to obtain structural repair masks corresponding to the multiple initial repair masks respectively.
[0051] Optical flow estimation networks refer to deep learning models used to calculate the motion vector of each pixel in multiple initial restoration masks. Their core objective is to achieve dense optical flow estimation, which involves assigning a two-dimensional motion vector to each pixel in the initial restoration mask, calculating bidirectional optical flow information between adjacent initial restoration masks, and mapping and filling unoccluded background pixels from previous and subsequent frames into the initial restoration mask of the current frame. Mainstream optical flow estimation networks include the following categories: FlowNet series: As the first end-to-end model to apply convolutional neural networks to optical flow estimation, its architecture consists of an encoder and a decoder, fusing information from two frames through concatenation or correlation operations. PWCnet series: Employing three core components—Pyramid, Warping, and Cost Volume—it constructs an image feature pyramid and performs optical flow estimation at multiple scales. LiteFlowNet series: As a lightweight network, it combines sub-pixel correction layers with feature-driven local convolutional regularization layers through a three-level cascaded structure to progressively optimize optical flow results. However, this invention is not limited to these specific features; those skilled in the art can set the appropriate settings based on actual conditions.
[0052] S122: Perform texture restoration processing on multiple structural restoration masks to obtain the target restoration masks corresponding to the multiple initial restoration masks.
[0053] Specifically, after obtaining multiple structural repair masks, texture repair processing is performed on the multiple structural repair masks to obtain the target repair masks corresponding to the multiple initial repair masks.
[0054] Optionally, based on the above embodiments, in some embodiments of the present invention, one way to perform texture restoration processing on multiple structural restoration masks is to perform texture restoration processing on multiple structural restoration masks through a video diffusion model to obtain target restoration masks corresponding to multiple initial restoration masks respectively.
[0055] The video diffusion model refers to a generative model that combines the advantages of variational autoencoders and diffusion models. The core idea is to perform the diffusion process in the latent space, rather than directly operating in the high-dimensional data space. This generates high-frequency texture details consistent with the surrounding background distribution, which are then used to denoise and compensate for details in the initial inpainting mask, resulting in a final low-resolution target inpainting mask. This reduces the computational complexity and cost of obtaining a high-quality subtitle-erased video of the original video.
[0056] S13: Based on the fidelity stream of the video to be processed, multiple initial repair masks, and the target repair masks corresponding to the multiple initial repair masks, obtain the subtitle-erased video corresponding to the video to be processed.
[0057] Specifically, after obtaining multiple initial repair masks and the target repair masks corresponding to each initial repair mask, the video with subtitles erased corresponding to the video to be processed is obtained based on the fidelity stream of the video to be processed, the multiple initial repair masks, and the target repair masks corresponding to each initial repair mask.
[0058] Optionally, based on the above embodiments, since the multiple target restoration masks are obtained based on the video computation stream to be processed, and the multiple target restoration masks and the video computation stream to be processed have the same video resolution, in order to obtain the subtitle-erased video corresponding to the video to be processed with high video resolution, in some embodiments of the present invention, one implementation of S13 may be: S131: Based on the video resolution of the video fidelity stream to be processed, upsample the multiple target restoration masks respectively to obtain multiple upsampled target restoration masks.
[0059] Specifically, the video resolution of the video fidelity stream to be processed is obtained, and based on the video resolution of the video fidelity stream to be processed, multiple target restoration masks are upsampled to obtain multiple upsampled target restoration masks.
[0060] It should be noted that the fidelity stream of the video to be processed has the same video resolution as the video to be processed. Since multiple target restoration masks are obtained based on the computational stream of the video to be processed, and these multiple target restoration masks have the same video resolution as the computational stream of the video to be processed, the computational stream of the video to be processed may be obtained by downsampling the video to be processed, or it may be determined that the video to be processed is the computational stream of the video to be processed. Therefore, when the computational stream of the video to be processed is the video to be processed, it is not necessary to perform upsampling processing on the multiple target restoration masks separately; determining the multiple target restoration masks is equivalent to determining multiple upsampled target restoration masks.
[0061] S132: Perform Gaussian blur processing on multiple initial repair masks to obtain the soft masks corresponding to each initial repair mask.
[0062] Specifically, Gaussian blurring is applied to multiple initial repair masks to generate soft masks with gradient alpha channels corresponding to each initial repair mask.
[0063] It should be noted that the Gaussian blur kernel size is dynamically adjusted according to the video resolution.
[0064] S133: Based on the fidelity stream of the video to be processed, multiple upsampled target repair masks, and multiple soft masks, obtain the subtitle-erased video corresponding to the video to be processed.
[0065] Specifically, after obtaining multiple upsampled target restoration masks and multiple initial restoration masks, the subtitle-erased video corresponding to the video to be processed is obtained based on the fidelity stream of the video to be processed, the multiple upsampled target restoration masks, and the multiple soft masks.
[0066] Optionally, based on the above embodiments, since not all target detection areas included in multiple video calculation frames to be processed contain target subtitles to be erased, in some embodiments of the present invention, S133 can be implemented as follows: S1331: For multiple video calculation frames to be processed that contain target subtitles to be erased, determine the target video fidelity frame corresponding to each video calculation frame to be processed among the multiple video fidelity frames to be processed.
[0067] Specifically, for multiple video calculation frames to be processed where the target subtitle to be erased exists in the target detection region, the corresponding target video fidelity frame to be processed is determined among the multiple video fidelity frames to be processed.
[0068] S1332: Perform Alpha mixing processing on multiple target video fidelity frames to be processed, multiple upsampled target repair masks corresponding to the target subtitles to be erased, and multiple soft masks respectively to obtain the subtitle-erased video corresponding to the video to be processed.
[0069] Specifically, Alpha mixing is performed on multiple target video fidelity frames to be processed, multiple upsampled target repair masks corresponding to the target subtitles to be erased, and multiple soft masks to obtain the subtitle-erased video corresponding to the video to be processed.
[0070] Here, this embodiment can perform alpha mixing processing based on the high-resolution target video fidelity frame to be processed, the upsampled target repair mask, and multiple soft masks to obtain a high-quality video with subtitle erasure corresponding to the video to be processed.
[0071] Thus, the video subtitle erasure method provided in this embodiment obtains the video fidelity stream and the video computation stream corresponding to the video to be processed. The video computation stream consists of multiple video computation frames, each including a target detection region. For each of the multiple video computation frames, it detects whether a target subtitle to be erased exists in each target detection region. If so, it obtains an initial repair mask for the target subtitle. It then performs structural texture repair on the multiple initial repair masks to obtain the target repair masks corresponding to each initial repair mask. Based on the video fidelity stream, the multiple initial repair masks, and the target repair masks corresponding to each initial repair mask, it obtains the subtitle-erased video corresponding to the video to be processed. In this way, under the premise of a video computation stream with low video resolution, it obtains the subtitle-erased video, effectively improving the computational efficiency of obtaining the subtitle-erased video and reducing computational costs. Furthermore, by performing structural and texture restoration only on multiple initial restoration masks, the entire video calculation frame to be processed can be restored, thus solving the problem of image quality degradation caused by re-encoding or scaling the entire video calculation frame to be processed and improving the quality of the video after subtitle erasure.
[0072] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0073] In one embodiment, such as Figure 2 As shown, Figure 2This is a schematic diagram of a video subtitle erasing device provided in an embodiment of the present invention, including: a video processing module 10, an initial repair mask acquisition module 11, a target repair mask acquisition module 12, and a video acquisition module 13 after subtitle erasure.
[0074] The video processing module 10 is used to acquire the video fidelity stream and the video computation stream corresponding to the video to be processed. The video computation stream consists of multiple video computation frames to be processed. Each video computation frame to be processed includes a target detection region.
[0075] The initial repair mask acquisition module 11 is used to calculate frames for multiple videos to be processed, detect whether there are target subtitles to be erased in each target detection area, and if so, acquire the initial repair mask of the target subtitles to be erased.
[0076] The target repair mask acquisition module 12 is used to perform structural texture repair on multiple initial repair masks and obtain the target repair mask corresponding to each of the multiple initial repair masks.
[0077] The subtitle erasure video acquisition module 13 is used to acquire the subtitle erasure video corresponding to the video to be processed based on the fidelity stream of the video to be processed, multiple initial repair masks, and the target repair masks corresponding to the multiple initial repair masks.
[0078] Thus, the video subtitle erasure device provided in this embodiment can utilize the video processing module to obtain the fidelity stream and computational stream of the video to be processed, wherein the computational stream of the video to be processed consists of multiple computational frames; each computational frame includes a target detection region. The initial repair mask acquisition module detects whether there are target subtitles to be erased in each target detection region for the multiple computational frames. If so, it acquires the initial repair mask of the target subtitles to be erased. The target repair mask acquisition module performs structural texture restoration on the multiple initial repair masks to acquire the target repair masks corresponding to each of the multiple initial repair masks. The subtitle-erased video acquisition module acquires the subtitle-erased video corresponding to the video to be processed based on the fidelity stream of the video to be processed, the multiple initial repair masks, and the target repair masks corresponding to the multiple initial repair masks. In this way, under the premise of a computational stream of the video to be processed with low video resolution, acquiring the subtitle-erased video effectively improves the computational efficiency of acquiring the subtitle-erased video and reduces the computational cost. Furthermore, by performing structural and texture restoration only on multiple initial restoration masks, the entire video calculation frame to be processed can be restored, thus solving the problem of image quality degradation caused by re-encoding or scaling the entire video calculation frame to be processed and improving the quality of the video after subtitle erasure.
[0079] Specific limitations regarding the video subtitle erasure device can be found in the limitations of the video subtitle erasure method described above, and will not be repeated here. Each module in the aforementioned server can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independent of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the corresponding operations of each module.
[0080] This invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement a video subtitle erasing method provided in this invention. For example, when the processor executes the computer program, it can implement... Figure 1 The technical solutions of the method embodiments shown are similar in principle and in effect, and will not be described again here.
[0081] This invention provides a computer-readable storage medium storing at least one program, which is executed by a processor to implement... Figure 1 The technical solutions of the method embodiments shown are similar in principle and in effect, and will not be described again here.
[0082] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM), etc.
[0083] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0084] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A method for erasing a video caption, characterized by, The method includes: Obtain the video fidelity stream and video computation stream corresponding to the video to be processed, wherein the video computation stream consists of multiple video computation frames; each video computation frame includes: a target detection region; For multiple video calculation frames to be processed, detect whether there is a target subtitle to be erased in each target detection region. If so, obtain the initial repair mask of the target subtitle to be erased. Structural texture restoration is performed on multiple initial repair masks to obtain target repair masks corresponding to each of the multiple initial repair masks; Based on the fidelity stream of the video to be processed, multiple initial repair masks, and the target repair masks corresponding to the multiple initial repair masks, obtain the subtitle-erased video corresponding to the video to be processed.
2. The method of claim 1, wherein, The process of obtaining the fidelity stream and computational stream of the video to be processed, corresponding to the video to be processed, includes: Obtain the pixel value of the long side of the video resolution corresponding to the video to be processed; Determine whether the pixel value of the longer side is greater than a preset threshold; If the pixel value of the longer side is greater than a preset threshold, the video to be processed is downsampled to obtain the video computation stream and the video to be processed is determined to be the video fidelity stream.
3. The method of claim 2, wherein, The method further includes: If the pixel value of the longer side is not greater than a preset threshold, the video to be processed is determined to be both the video computation stream and the video fidelity stream.
4. The method according to claim 3, characterized in that, The process of obtaining the initial repair mask for the target subtitle to be erased includes: The target subtitle to be erased is binarized to obtain the binarized mask of the target subtitle to be erased; The binarized mask is subjected to morphological dilation to obtain the initial repair mask for the target subtitle to be erased.
5. The method according to claim 4, characterized in that, The step of performing structural texture restoration on multiple initial restoration masks to obtain target restoration masks corresponding to each of the multiple initial restoration masks includes: Perform structural repair processing on the multiple initial repair masks to obtain the structural repair masks corresponding to the multiple initial repair masks respectively; Texture restoration processing is performed on multiple structural restoration masks to obtain target restoration masks corresponding to the multiple initial restoration masks.
6. The method according to claim 5, characterized in that, The step of obtaining the subtitle-erased video corresponding to the video to be processed based on the fidelity stream of the video to be processed, multiple initial repair masks, and target repair masks corresponding to the multiple initial repair masks includes: Based on the video resolution of the video fidelity stream to be processed, the multiple target restoration masks are upsampled respectively to obtain multiple upsampled target restoration masks; Gaussian blurring is applied to multiple initial repair masks to obtain the soft masks corresponding to each initial repair mask; Based on the fidelity stream of the video to be processed, multiple upsampled target repair masks, and multiple soft masks, obtain the subtitle-erased video corresponding to the video to be processed.
7. The method according to claim 6, characterized in that, The video fidelity stream to be processed consists of multiple video fidelity frames to be processed, each of which corresponds one-to-one with a video calculation frame to be processed. The step of obtaining the subtitle-erased video corresponding to the video to be processed based on the video fidelity stream to be processed, multiple upsampling target restoration masks, and multiple soft masks includes: For multiple video calculation frames to be processed that contain target subtitles to be erased, the target video fidelity frame corresponding to each video calculation frame to be processed is determined among the multiple video fidelity frames to be processed. Alpha mixing is performed on multiple target video fidelity frames to be processed, multiple upsampled target repair masks corresponding to the target subtitles to be erased, and multiple soft masks to obtain the subtitle-erased video corresponding to the video to be processed.
8. A video subtitle erasing device, characterized in that, The device includes: The video processing module is used to acquire the video fidelity stream and the video computation stream corresponding to the video to be processed, wherein the video computation stream consists of multiple video computation frames; each video computation frame includes: a target detection region; The initial repair mask acquisition module is used to calculate frames for multiple videos to be processed, detect whether there are target subtitles to be erased in each target detection area, and if so, acquire the initial repair mask of the target subtitles to be erased; The target repair mask acquisition module is used to perform structural texture repair on multiple initial repair masks and acquire the target repair mask corresponding to each of the multiple initial repair masks. The subtitle-erased video acquisition module is used to acquire the subtitle-erased video corresponding to the video to be processed based on the fidelity stream of the video to be processed, multiple initial repair masks, and target repair masks corresponding to the multiple initial repair masks.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the video subtitle erasure method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the video subtitle erasure method according to any one of claims 1 to 7.