Video subtitle erasing method and device, equipment and storage medium

By splitting video subtitle segments, determining reference frames, and generating text masks, the accuracy problem of video subtitle erasure is solved, the erasure performance and picture quality are improved, and it is suitable for multilingual remastering and viewing experience.

CN121509729APending Publication Date: 2026-02-10MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511761223.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately erase video subtitles, resulting in a decrease in visual integrity and viewing immersion, and increasing the cost and time spent on multilingual remastering.

Method used

By acquiring the video with subtitles to be erased, the subtitle segments are split using a preset split-scene algorithm, the target reference frame is determined, an expanded independent shot segment is generated and the text mask is segmented, and the erased segment is generated by combining the preset erasure algorithm and post-processing optimization is performed. The subtitle-free segments are then integrated to generate the target video.

Benefits of technology

It achieves accurate erasure of video subtitles, improves erasure performance, avoids the impact of transition effects on transition frames between shot segments, and ensures natural and coherent visual integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509729A_ABST
    Figure CN121509729A_ABST
Patent Text Reader

Abstract

The invention discloses a video subtitle erasing method, device and equipment and a storage medium, and relates to the field of digital image processing, and the method comprises the steps: detecting subtitles of a subtitle video to be erased, merging timestamps of the same subtitles in the subtitle video to be erased, and determining a subtitle fragment set and a subtitle-free fragment set; splitting the subtitle segment into independent shot segments by using a preset lens splitting algorithm, and analyzing video frames of the independent shot segments to obtain a first frame and a tail frame; determining a target reference frame based on the first frame and the tail frame, generating an expanded independent shot segment according to the target reference frame and the independent shot segment, and segmenting a target character mask; and generating an erased independent lens segment according to the expanded independent lens segment and the target character mask through a preset erasure algorithm, performing a preset post-processing optimization operation on the erased independent lens segment to obtain a target independent lens segment, and integrating the target independent lens segment and the subtitle-free segment set to generate a target video. According to the method and the device, the video subtitles can be accurately erased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital image processing, and in particular to a method, apparatus, device, and storage medium for erasing video subtitles. Background Technology

[0002] In the booming trend of short dramas, subtitle erasure algorithms play a crucial strategic role. By accurately identifying and seamlessly removing original subtitles, they provide a clean canvas for content localization, enabling the same short drama to quickly adapt to different markets and significantly reducing the high costs and time consumption of traditional multilingual remastering. More importantly, high-quality subtitle erasure ensures the visual integrity and immersive viewing experience, avoiding the decline in viewing quality caused by jarring obstructions or image damage, and providing different users with an authentic yet linguistically appropriate viewing experience.

[0003] In conclusion, how to accurately erase video subtitles is a problem that urgently needs to be solved. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a method, apparatus, device, and storage medium for erasing video subtitles, capable of accurately erasing video subtitles. The specific solution is as follows:

[0005] Firstly, this application provides a method for erasing video subtitles, including:

[0006] The process involves acquiring a video with subtitles to be erased, detecting the subtitles in the video, and merging the timestamps corresponding to the same subtitles in the video to be erased to determine the set of segments with subtitles and the set of segments without subtitles in the video.

[0007] Using a preset segmentation algorithm, each subtitled segment in the set of subtitled segments is split into independent shot segments, and video frame parsing is performed on the independent shot segments to obtain the first and last frames of the independent shot segments;

[0008] A target reference frame is determined based on the first and last frames of the independent shot segment. An extended independent shot segment is generated based on the target reference frame and the independent shot segment. The target text mask of the extended independent shot segment is then segmented.

[0009] A preset erasure algorithm is used to generate corresponding erased independent shot segments based on the extended independent shot segments and the target text mask. The erased independent shot segments are then subjected to preset post-processing optimization operations to obtain target independent shot segments that meet the edge blending conditions. The target independent shot segments are then integrated with the set of segments without subtitles to generate the target video.

[0010] Optionally, the step of using a preset shot-by-shot algorithm to split each subtitled segment in the set of subtitled segments into independent shot segments includes:

[0011] The target transition video frames of each subtitled segment in the set of subtitled segments are identified using a preset segmentation algorithm;

[0012] The subtitled segment is split into independent shot segments based on the target transition video frame.

[0013] Optionally, determining the target reference frame based on the first and last frames of the independent shot segment includes:

[0014] Obtain the first set of video frames that are a preset number of frames prior to the first frame of the independent shot segment;

[0015] Obtain a second set of video frames that are a preset number of frames after the last frame of the independent shot segment;

[0016] Determine the first similarity between the first frame and the video frames in the first video frame set;

[0017] Determine the second similarity between the tail frame and video frames in the second video frame set;

[0018] The target reference frame is determined based on the first similarity and the second similarity.

[0019] Optionally, determining the target reference frame based on the first similarity and the second similarity includes:

[0020] Determine whether the first similarity and / or the second similarity are greater than a preset threshold;

[0021] If the first similarity is greater than the preset threshold, then the video frame corresponding to the first similarity is taken as the first target reference frame corresponding to the first frame.

[0022] If the second similarity is greater than the preset threshold, then the video frame corresponding to the second similarity is designated as the second target reference frame corresponding to the tail frame.

[0023] Optionally, generating the expanded independent shot segment based on the target reference frame and the independent shot segment includes:

[0024] The first target reference frame is used as the new first frame of the independent shot segment, and the second target reference frame is used as the new last frame of the independent shot segment.

[0025] The extended independent shot segments are determined based on the new first frame and the new last frame.

[0026] Optionally, the step of generating corresponding erased independent shot segments based on the expanded independent shot segments and the target text mask using a preset erasure algorithm includes:

[0027] The target diffusion model is determined by a preset erasure algorithm;

[0028] The target diffusion model is used to generate corresponding erased independent shot segments based on the expanded independent shot segments and the target text mask.

[0029] Optionally, the step of performing a preset post-processing optimization operation on the erased independent shot fragment to obtain a target independent shot fragment that meets the edge blending conditions includes:

[0030] A preset edge blending operation is performed on the erased independent shot segments to obtain edge-blended video segments;

[0031] Perform a preset color equalization operation on the edge-blended video clip to generate a color-equalized video clip;

[0032] Perform a preset temporal filtering operation on the edge-fused video segment to generate a temporally filtered video segment;

[0033] A preset video quality enhancement operation is performed on the edge-blended video segment to obtain a target independent shot segment that meets the edge-blending conditions.

[0034] Secondly, this application provides a video subtitle erasing device, comprising:

[0035] The set determination module is used to acquire the video with subtitles to be erased, detect the subtitles in the video with subtitles to be erased, and merge the timestamps corresponding to the same subtitles in the video with subtitles to be erased, so as to determine the set of subtitled segments and the set of subtitle-free segments in the video with subtitles to be erased.

[0036] The video frame acquisition module is used to split each subtitled segment in the set of subtitled segments into independent shot segments using a preset shot segmentation algorithm, and to perform video frame parsing on the independent shot segments to obtain the first frame and the last frame of the independent shot segment.

[0037] The mask segmentation module is used to determine the target reference frame based on the first and last frames of the independent shot segment, generate an extended independent shot segment based on the target reference frame and the independent shot segment, and segment the target text mask of the extended independent shot segment.

[0038] The video generation module is used to generate corresponding erased independent shot segments based on the extended independent shot segments and the target text mask using a preset erasure algorithm, and to perform preset post-processing optimization operations on the erased independent shot segments to obtain target independent shot segments that meet the edge blending conditions, and to integrate the target independent shot segments with the set of segments without subtitles to generate a target video.

[0039] Thirdly, this application provides an electronic device, comprising:

[0040] Memory, used to store computer programs;

[0041] A processor is used to execute the computer program to implement the video subtitle erasure method as described above.

[0042] Fourthly, this application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the video subtitle erasure method described above.

[0043] In summary, this application first acquires the video with subtitles to be erased, detects the subtitles in the video, and merges the timestamps corresponding to the same subtitles in the video to determine the set of subtitled segments and the set of subtitle-free segments in the video. A preset segmentation algorithm is used to split each subtitled segment in the set of subtitled segments into independent shot segments. Video frame parsing is performed on each independent shot segment to obtain its first and last frames. A target reference frame is determined based on the first and last frames of the independent shot segments. An extended independent shot segment is generated based on the target reference frame and the independent shot segment, and the target text mask of the extended independent shot segment is segmented. A preset erasure algorithm is used to generate corresponding erased independent shot segments based on the extended independent shot segments and the target text mask. A preset post-processing optimization operation is performed on the erased independent shot segments to obtain target independent shot segments that meet edge blending conditions. The target independent shot segments are then integrated with the set of subtitle-free segments to generate the target video. As described above, this application first acquires the video with subtitles to be erased, determines the set of subtitled segments and the set of subtitle-free segments through subtitle detection and timestamp merging, then uses a preset segmentation algorithm to split the subtitled segments into independent shot segments and parses their first and last frames. Based on these frames, target reference frames are determined to generate expanded independent shot segments. Simultaneously, the target text mask of the shot segment is segmented. Then, a preset erasure algorithm is used to generate erased independent shot segments, which are post-processed to meet edge blending conditions. Finally, the optimized target independent shot segments and the set of subtitle-free segments are integrated to generate the target video. In this way, video segments are divided into subtitled and non-subtitled segments, and only subtitled segments are erased, improving erasure performance. Simultaneously, a video segmentation algorithm is introduced to segment subtitle segments, and subtitle erasure is performed on independent shot segments, thus avoiding the reduction in erasure effect of transition frames between two shot segments due to transition effects. Furthermore, an expanded reference frame strategy is introduced, providing clean reference information by expanding reference frames that are similar to each independent shot segment and sequentially continuous, improving the erasure effect. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0045] Figure 1 This is a flowchart of a video subtitle erasure method disclosed in this application;

[0046] Figure 2 This is a schematic diagram illustrating a specific target reference frame determination disclosed in this application;

[0047] Figure 3 This application discloses a specific video subtitle erasure method flowchart;

[0048] Figure 4 This is a schematic diagram of the structure of a video subtitle erasing device disclosed in this application;

[0049] Figure 5 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] Currently, in the booming short drama trend, subtitle erasure algorithms play a crucial strategic role. By accurately identifying and seamlessly removing original subtitles, they provide a clean canvas for content localization, enabling the same short drama to quickly adapt to different markets and significantly reducing the high costs and time consumption of traditional multilingual remastering. More importantly, high-quality subtitle erasure ensures the visual integrity and immersive viewing experience, avoiding the decline in viewing quality caused by abrupt obstruction or image damage, and providing different users with an authentic yet linguistically appropriate viewing experience. To address the aforementioned technical issues, this application discloses a video subtitle erasure method, apparatus, device, and storage medium capable of accurately erasing video subtitles.

[0052] See Figure 1 As shown, this embodiment of the invention discloses a video subtitle erasure method, including:

[0053] Step S11: Obtain the video with subtitles to be erased, detect the subtitles in the video with subtitles to be erased, and merge the timestamps corresponding to the same subtitles in the video with subtitles to be erased, so as to determine the set of subtitled segments and the set of subtitle-free segments in the video with subtitles to be erased.

[0054] In this embodiment, the video containing the subtitles to be erased is acquired. The source of this video can be the original source material from a film production company, film clips distributed on online platforms, short videos with subtitles created by self-media creators, or recorded videos of online courses. Next, object detection and image segmentation technologies are used to automatically identify the specific position, font style, color, and background features of the subtitles in the video frames. For example, for white bottom subtitles in a movie, the object detection and image segmentation technology first scans the lower area of ​​each frame, using the brightness difference between the subtitles and the background, and edge contour features to locate the text and any possible semi-transparent black overlay areas around it. Even if the subtitles occasionally shift or the font size changes, feature matching can accurately capture the changes.

[0055] Then, since the same line of subtitles in a video usually remains unchanged across multiple consecutive frames, each frame corresponds to an independent timestamp. Directly processing a single frame would result in a large amount of repetitive calculations. Taking the line "The weather is really nice today" as an example, it is continuously displayed from frame 2:15:03 to frame 2:18:12 in the video, corresponding to dozens of consecutive timestamps. At this point, by comparing text content and matching image features, it is determined that these timestamps correspond to the same subtitle, and then they are merged into a complete time segment.

[0056] Finally, all the merged consecutive time segments containing subtitles, such as the multiple merged segments from 2 minutes 15 seconds 03 frames to 2 minutes 18 seconds 12 frames and 5 minutes 02 seconds 20 frames to 5 minutes 05 seconds 08 frames in the aforementioned movie, are integrated into a set of subtitled segments. The set of subtitle-free segments includes all content in the video excluding these time segments, such as the opening establishing shots of the movie, scenes without dialogue, and transitional segments between scenes. This division allows the subsequent subtitle erasure algorithm to accurately focus on subtitled segments, eliminating the need for additional processing on subtitle-free segments. This reduces computational resource consumption, avoids misprocessing subtitle-free areas, and ensures that the overall video quality remains unaffected.

[0057] Step S12: Using a preset segmentation algorithm, each subtitled segment in the set of subtitled segments is split into independent shot segments, and video frame parsing is performed on the independent shot segments to obtain the first and last frames of the independent shot segments.

[0058] In this embodiment, after dividing the set of subtitled segments, a preset segmentation algorithm is introduced to split the shots. This algorithm can identify the target transition video frames for each subtitled segment in the set of subtitled segments. Based on the target transition video frames, the subtitled segments are split into independent shot segments. Specifically, based on video image features such as scene abrupt changes, motion vector changes, color distribution differences, and inter-frame similarity, intelligent recognition rules are set to automatically capture the target transition video frames for shot transitions, such as common cut-in / cut-out, fade-in / fade-out, push-pull, pan, and tilt transitions. Using the identified target transition video frames as the splitting benchmark, each subtitled segment in the set of subtitled segments is split into independent shot segments. For example, a subtitled segment of a school-themed video, from 3 minutes 05 seconds 10 frames to 3 minutes 08 seconds 05 frames, depicts a classroom dialogue between students and a teacher. This segment includes two shots: the first is a medium shot of the teacher (3 minutes 05 seconds 10 frames to 3 minutes 06 seconds 20 frames), with the subtitles located at the bottom center of the frame; the second is a close-up of the students (3 minutes 06 seconds 21 frames to 3 minutes 08 seconds 05 frames), where the subtitles remain in the same position but the background changes from a blackboard to books. A pre-defined segmentation algorithm analyzes the inter-frame differences of this subtitled segment frame by frame. When it detects that in the frame at 3 minutes 06 seconds 21, the subject changes from the teacher to the students, the background color abruptly changes from a dark gray blackboard to light beige books, and the inter-frame similarity is below a preset threshold, it determines this as a shot transition point. Consequently, the original subtitled segment from 3 minutes 05 seconds 10 frames to 3 minutes 08 seconds 05 frames is split into two independent shot segments. Subsequently, video frames are analyzed for each independent shot segment. For each independent shot segment, all video frames within the segment are scanned, and the frame with the earliest timestamp is determined as the first frame, while the frame with the latest timestamp is determined as the last frame.

[0059] Step S13: Determine the target reference frame based on the first and last frames of the independent shot segment, generate an extended independent shot segment based on the target reference frame and the independent shot segment, and segment the target text mask of the extended independent shot segment.

[0060] In this embodiment, to improve the erasure effect of the algorithm, it is necessary to determine the target reference frame. This can be achieved by obtaining a first set of video frames preceding the first frame of the independent shot segment with a preset number of frames; obtaining a second set of video frames following the last frame of the independent shot segment with a preset number of frames; determining a first similarity between the first frame and the video frames in the first set of video frames; determining a second similarity between the last frame and the video frames in the second set of video frames; and determining the target reference frame based on the first and second similarities. Specifically, a preset number of video frames preceding the first frame are determined as the first set of video frames, and a preset number of video frames following the last frame are determined as the second set of video frames. It should be noted that the preset number of frames is a reasonable value preset based on the stability of the shot scene, the pattern of subtitle appearance, and the video frame rate. It is usually one to a dozen frames, and it must be ensured that the video frames within this range belong to the same scene as the independent shot segment, without shot transitions, large object movements, or sudden changes in lighting, to avoid introducing irrelevant interference frames that affect the quality of the reference frame. After obtaining the first and second video frame sets, the first frame, the last frame, and the corresponding video frame sets are transformed into feature vectors through key dimensions such as color distribution histogram, image texture structure, light intensity distribution, and scene object layout. Then, using common calculation methods such as cosine similarity and Euclidean distance, the first similarity between the first frame and each frame in the first video frame set, and the second similarity between the last frame and each frame in the second video frame set are quantified.

[0061] Furthermore, since a higher similarity indicates that the background of the corresponding video frame is closer to the first or last frame, the background information it contains has greater reference value for subtitle erasure. Therefore, after obtaining the first similarity and the second similarity, it is necessary to determine the target reference frame based on the similarity. It is necessary to determine whether the first similarity and / or the second similarity is greater than a preset threshold. If the first similarity is greater than the preset threshold, the video frame corresponding to the first similarity is designated as the first target reference frame corresponding to the first frame. If the second similarity is greater than the preset threshold, the video frame corresponding to the second similarity is designated as the second target reference frame corresponding to the last frame.

[0062] In one specific implementation, if there is a video frame in the first video frame set with a first similarity higher than a preset threshold, it means that the frame has the same background height as the first frame and is not obscured by subtitles. The frame with the highest first similarity can be selected as the first target reference frame. Similarly, if there is a video frame in the second video frame set with a second similarity higher than a threshold, the frame with the highest second similarity in the set is selected as the second target reference frame.

[0063] In another specific implementation, if there are video frames in the first video frame set with a first similarity higher than a threshold, then the video frame with the time closest to the first frame is selected from the video frames with the first similarity higher than the preset threshold as the first target reference frame. Similarly, if there are video frames in the second video frame set with a second similarity higher than the preset threshold, then the video frame with the time closest to the last frame is selected from the video frames with the second similarity higher than the preset threshold as the second target reference frame.

[0064] In the third specific implementation, such as Figure 2 As shown, the frame preceding the first frame is directly used as the first target reference frame, and the frame following the last frame is used as the second target reference frame.

[0065] In addition, if there are no video frames that meet the threshold in either the first or second set of video frames, it indicates that there is scene interference in the frames before the first frame and after the last frame. In this case, the first and last frames will be reverted, and the target reference frame will be determined through the previous feature complementation mechanism.

[0066] Furthermore, the first target reference frame is used as the new first frame of the independent shot segment, and the second target reference frame is used as the new last frame of the independent shot segment; the expanded independent shot segment is determined based on the new first frame and the new last frame. Specifically, the first target reference frame and the second target reference frame are updated to the new first frame and the new last frame of the independent shot segment, respectively, to obtain the expanded independent shot segment. The first target reference frame, as the frame with the highest similarity before the first frame and without subtitle obstruction, has complete and undisturbed scene background features. The second target reference frame, as the frame with consistent scene and no additional interference after the last frame, also has pure background information. The new first and last frames can provide a more reliable background benchmark for the entire shot segment, avoiding the problem of missing background information caused by the presence of subtitles in the original boundary frames.

[0067] Finally, each frame of the expanded independent shot segment is preprocessed. The feature differences between the subtitles and the background are enhanced through operations such as grayscale conversion and edge enhancement. Then, the pixels of the image are classified and filtered by combining the background feature library of the target reference frame. For example, based on multi-dimensional indicators such as brightness gradient, color threshold, and texture matching, the subtitle body, semi-transparent mask and background pixels are distinguished, and finally a binary target text mask is generated.

[0068] Step S14: Generate corresponding erased independent shot segments based on the extended independent shot segments and the target text mask using a preset erasure algorithm, and perform preset post-processing optimization operations on the erased independent shot segments to obtain target independent shot segments that meet the edge blending conditions. Integrate the target independent shot segments with the set of segments without subtitles to generate the target video.

[0069] In this embodiment, a target diffusion model is determined through a preset erasure algorithm. The target diffusion model is then used to generate corresponding erased independent shot segments based on the expanded independent shot segments and the target text mask. Specifically, the preset erasure algorithm incorporates a library of diffusion model parameters for different image features. It automatically matches and determines the optimal target diffusion model based on the background type, subtitle style, and image quality parameters of the expanded independent shot segments. The target diffusion model first accurately locates the subtitle area to be erased in each frame using the target text mask. Then, based on the target reference frame in the expanded independent shot segments, it extracts core features such as background texture, color, and lighting as the basis for generation. During the diffusion process, the model gradually reconstructs the subtitle-occluded area pixel by pixel. Through multiple rounds of denoising and feature matching, the filled pixels seamlessly blend with the surrounding background to generate the erased independent shot segments.

[0070] Furthermore, to address the potential for subtle edge traces or slight inter-frame discontinuities remaining after the diffusion model is generated, a preset edge blending operation needs to be performed on the erased independent shot segments to obtain edge-blended video segments. A preset color equalization operation is then performed on the edge-blended video segments to generate color-balanced video segments. A preset temporal filtering operation is then performed on the edge-blended video segments to generate temporally filtered video segments. Finally, a preset video quality enhancement operation is performed on the edge-blended video segments to obtain target independent shot segments that meet the edge blending conditions. Specifically, firstly, pixel gradient detection is used to locate the boundary between the filled area and the original background. Then, adaptive Poisson blending and Gaussian gradient smoothing are used to construct a transition band to eliminate the splicing effect and achieve seamless edge connection. The preset color equalization operation extracts global color features and calibrates the color of the filled area through histogram matching to ensure consistency with the original image's hue and saturation, avoiding abrupt color changes. The preset temporal filtering relies on inter-frame spatiotemporal correlation and uses adaptive median filtering to smooth inter-frame jitter in the filled area, eliminating abrupt changes and ensuring smooth dynamic images. Finally, preset video quality enhancement operations are used to improve image quality through noise reduction, sharpening, and detail restoration techniques, eliminating blur and noise. Through this series of progressive optimizations, the final result is a target independent shot segment with natural edge blending, unified colors, smooth inter-frame transitions, and clear image quality.

[0071] As described above, this embodiment first acquires the video with subtitles to be erased, determines the set of subtitled segments and the set of subtitle-free segments by subtitle detection and timestamp merging, then uses a preset segmentation algorithm to split the subtitled segments into independent shot segments and parses their first and last frames. Based on these frames, target reference frames are determined to generate extended independent shot segments. Simultaneously, the target text mask of the shot segment is segmented. Then, a preset erasure algorithm is used to generate erased independent shot segments, which are post-processed to meet edge blending conditions. Finally, the optimized target independent shot segments and the set of subtitle-free segments are integrated to generate the target video. In this way, video segments are divided into subtitled and non-subtitled segments, and only subtitled segments are erased, improving erasure performance. Simultaneously, a video segmentation algorithm is introduced to segment subtitle segments, and subtitle erasure is performed on independent shot segments. This avoids the transition effect from reducing the erasure effect of transition frames between two shot segments. Furthermore, an extended reference frame strategy is introduced. By extending reference frames that are similar to each independent shot segment and are temporally continuous, clean reference information is provided, improving the erasure effect.

[0072] As can be seen from the previous embodiment, this application discloses a video subtitle erasure method, which can accurately erase video subtitles. Next, suppose a film studio needs to erase subtitles from a 10-minute campus short drama, and will target methods such as... Figure 3 The video subtitle erasure method shown is explained in detail.

[0073] First, the original high-definition source of the short drama was acquired as the video to be erased for subtitles. A subtitle detection algorithm scanned the entire video frame by frame, accurately identifying the white bottom subtitles of dialogue, narration, and scene prompts, while simultaneously recording the timestamps corresponding to each subtitle. Since the same line, "Our class made great progress in this exam," remained unchanged for 20 consecutive frames, the consecutive timestamps corresponding to these repeated subtitles were automatically merged, ultimately dividing the video into a set of subtitled segments: such as 12 segments like 2 minutes 15 seconds to 2 minutes 18 seconds and 5 minutes 03 seconds to 5 minutes 07 seconds; and a set of subtitle-free segments: a 3-second empty shot at the beginning, a classroom scene transition from 4 minutes 20 seconds to 4 minutes 30 seconds, and a 5-second black screen at the end.

[0074] Subsequently, a pre-defined shot breakdown algorithm was used to split each segment in the set of subtitled clips. Taking the subtitled clip from 5 minutes 03 seconds to 5 minutes 07 seconds as an example, which presents a scene of a teacher announcing grades, there are two shot transitions: from 5 minutes 03 seconds to 5 minutes 05 seconds is a medium shot of the teacher, and from 5 minutes 05 seconds to 5 minutes 07 seconds it switches to a wide shot of the whole class. The shot breakdown algorithm accurately identified the shot transition at 5 minutes 05 seconds by detecting abrupt changes in the subject and background colors between frames, splitting the original clip into two independent shot segments. The first and last frames of each shot were extracted through video frame analysis, laying the foundation for subsequent targeted processing.

[0075] The target reference frame is determined based on the first and last frames of an independent shot segment: Taking the medium shot of the teacher from 5 minutes 03 seconds to 5 minutes 05 seconds as an example, the algorithm extracts the first two frames as the first video frame set and the last two frames as the second video frame set. By comparing color texture and lighting features, similarity is calculated, and the last frame at 5 minutes 02 seconds, which has the highest similarity and is not obscured by subtitles, is selected as the target reference frame. Combining this target reference frame with the original independent shot segment, the algorithm extends the segment forward by 2 frames and backward by 2 frames to generate an extended independent shot segment. Then, semantic segmentation technology is used to segment the target text mask corresponding to the subtitles and semi-transparent mask in the segment, accurately locating the pixel area to be erased.

[0076] Finally, a target diffusion model suitable for the classroom scene of the short drama was determined through a preset erasure algorithm. This target diffusion model, combined with expanded independent shot segments and target text masks, was used to reconstruct the subtitle area pixel-level, generating erased segments. Subsequently, preset post-processing optimizations were performed on the erased segments: edge blending eliminated the boundary between the filled area and the background; color balancing ensured a unified color tone; temporal filtering prevented inter-frame jitter; and quality enhancement improved image clarity, ultimately resulting in target independent shot segments that met the edge blending conditions. All optimized target independent shot segments were seamlessly integrated along the original timeline with a collection of subtitle-free segments, including opening shots and scene transitions, to generate a target video for the campus short drama with completely erased subtitles and natural, consistent image quality. This video can be directly used for subsequent multilingual subtitle addition and distribution on various platforms.

[0077] See Figure 4 As shown, an embodiment of the present invention discloses a video subtitle erasing device, comprising:

[0078] The set determination module 11 is used to acquire the video with subtitles to be erased, detect the subtitles in the video with subtitles to be erased, and merge the timestamps corresponding to the same subtitles in the video with subtitles to be erased, so as to determine the set of subtitled segments and the set of subtitle-free segments in the video with subtitles to be erased.

[0079] The video frame acquisition module 12 is used to split each subtitled segment in the set of subtitled segments into independent shot segments using a preset shot segmentation algorithm, and to perform video frame parsing on the independent shot segments to obtain the first frame and the last frame of the independent shot segment.

[0080] The mask segmentation module 13 is used to determine the target reference frame based on the first and last frames of the independent shot segment, generate an extended independent shot segment based on the target reference frame and the independent shot segment, and segment the target text mask of the extended independent shot segment.

[0081] The video generation module 14 is used to generate corresponding erased independent shot segments based on the extended independent shot segments and the target text mask using a preset erasure algorithm, and to perform preset post-processing optimization operations on the erased independent shot segments to obtain target independent shot segments that meet the edge blending conditions, and to integrate the target independent shot segments with the set of segments without subtitles to generate a target video.

[0082] As described above, this application first acquires the video with subtitles to be erased, determines the set of subtitled segments and the set of subtitle-free segments through subtitle detection and timestamp merging, then uses a preset segmentation algorithm to split the subtitled segments into independent shot segments and parses their first and last frames. Based on these frames, target reference frames are determined to generate expanded independent shot segments. Simultaneously, the target text mask of the shot segment is segmented. Then, a preset erasure algorithm is used to generate erased independent shot segments, which are post-processed to meet edge blending conditions. Finally, the optimized target independent shot segments and the set of subtitle-free segments are integrated to generate the target video. In this way, video segments are divided into subtitled and non-subtitled segments, and only subtitled segments are erased, improving erasure performance. Simultaneously, a video segmentation algorithm is introduced to segment subtitle segments, and subtitle erasure is performed on independent shot segments, thus avoiding the reduction in erasure effect of transition frames between two shot segments due to transition effects. Furthermore, an expanded reference frame strategy is introduced, providing clean reference information by expanding reference frames that are similar to each independent shot segment and sequentially continuous, improving the erasure effect.

[0083] In some specific embodiments, the video frame acquisition module 12 may specifically include:

[0084] The transition video frame recognition unit is used to identify the target transition video frame of each subtitle segment in the set of subtitle segments using a preset segmentation algorithm;

[0085] The segment splitting unit is used to split the subtitled segment into independent shot segments based on the target transition video frame.

[0086] In some specific implementations, the mask segmentation module 13 may specifically include:

[0087] The first set acquisition unit is used to acquire a first set of video frames with a preset number of frames preceding the first frame of the independent shot segment;

[0088] The second set acquisition unit is used to acquire a second set of video frames that follows the last frame of the independent shot segment by a preset number of frames.

[0089] The first similarity determination unit is used to determine the first similarity between the first frame and the video frames in the first video frame set;

[0090] The second similarity determination unit is used to determine the second similarity between the tail frame and the video frames in the second video frame set;

[0091] A reference frame determination unit is used to determine a target reference frame based on the first similarity and the second similarity.

[0092] In some specific implementations, the reference frame determination unit may specifically include:

[0093] A similarity determination subunit is used to determine whether the first similarity and / or the second similarity is greater than a preset threshold;

[0094] The first similarity determination subunit is used to determine the video frame corresponding to the first similarity as the first target reference frame corresponding to the first frame if the first similarity is greater than the preset threshold.

[0095] The second similarity determination subunit is used to determine the video frame corresponding to the second similarity as the second target reference frame corresponding to the tail frame if the second similarity is greater than the preset threshold.

[0096] In some specific implementations, the mask segmentation module 13 may specifically include:

[0097] The first frame and last frame determination unit is used to take the first target reference frame as the new first frame of the independent shot segment and the second target reference frame as the new last frame of the independent shot segment.

[0098] The segment determination unit is used to determine the extended independent shot segments based on the new first frame and the new last frame.

[0099] In some specific embodiments, the video generation module 14 may specifically include:

[0100] The model determination unit is used to determine the target diffusion model through a preset erasure algorithm;

[0101] The segment generation unit is used to generate corresponding erased independent shot segments based on the expanded independent shot segments and the target text mask using the target diffusion model.

[0102] In some specific embodiments, the video generation module 14 may specifically include:

[0103] The first segment acquisition unit is used to perform a preset edge blending operation on the erased independent shot segment to obtain an edge-blended video segment.

[0104] The second segment acquisition unit is used to perform a preset color equalization operation on the edge-blended video segment to generate a color-equalized video segment.

[0105] The third segment acquisition unit is used to perform a preset temporal filtering operation on the edge-fused video segment to generate a temporally filtered video segment.

[0106] The fourth segment acquisition unit is used to perform a preset video quality enhancement operation on the edge-blended video segment in order to obtain a target independent shot segment that meets the edge-blending conditions.

[0107] Furthermore, embodiments of this application also disclose an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0108] Figure 5 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the video subtitle erasure method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be a computer.

[0109] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0110] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0111] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the video subtitle erasure method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0112] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned video subtitle erasure method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0113] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0114] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0115] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0116] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0117] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for erasing video subtitles, characterized in that, include: The process involves acquiring a video with subtitles to be erased, detecting the subtitles in the video, and merging the timestamps corresponding to the same subtitles in the video to be erased to determine the set of segments with subtitles and the set of segments without subtitles in the video. Using a preset segmentation algorithm, each subtitled segment in the set of subtitled segments is split into independent shot segments, and video frame parsing is performed on the independent shot segments to obtain the first and last frames of the independent shot segments; A target reference frame is determined based on the first and last frames of the independent shot segment. An extended independent shot segment is generated based on the target reference frame and the independent shot segment. The target text mask of the extended independent shot segment is then segmented. A preset erasure algorithm is used to generate corresponding erased independent shot segments based on the extended independent shot segments and the target text mask. The erased independent shot segments are then subjected to preset post-processing optimization operations to obtain target independent shot segments that meet the edge blending conditions. The target independent shot segments are then integrated with the set of segments without subtitles to generate the target video.

2. The video subtitle erasure method according to claim 1, characterized in that, The step of using a preset scene segmentation algorithm to split each subtitled segment in the set of subtitled segments into independent shot segments includes: The target transition video frames of each subtitled segment in the set of subtitled segments are identified using a preset segmentation algorithm; The subtitled segment is split into independent shot segments based on the target transition video frame.

3. The video subtitle erasure method according to claim 1, characterized in that, The determination of the target reference frame based on the first and last frames of the independent shot segment includes: Obtain the first set of video frames that are a preset number of frames prior to the first frame of the independent shot segment; Obtain a second set of video frames that are a preset number of frames after the last frame of the independent shot segment; Determine the first similarity between the first frame and the video frames in the first video frame set; Determine the second similarity between the tail frame and video frames in the second video frame set; The target reference frame is determined based on the first similarity and the second similarity.

4. The video subtitle erasure method according to claim 3, characterized in that, Determining the target reference frame based on the first similarity and the second similarity includes: Determine whether the first similarity and / or the second similarity are greater than a preset threshold; If the first similarity is greater than the preset threshold, then the video frame corresponding to the first similarity is taken as the first target reference frame corresponding to the first frame. If the second similarity is greater than the preset threshold, then the video frame corresponding to the second similarity is designated as the second target reference frame corresponding to the tail frame.

5. The video subtitle erasure method according to claim 4, characterized in that, The step of generating extended independent shot segments based on the target reference frame and the independent shot segments includes: The first target reference frame is used as the new first frame of the independent shot segment, and the second target reference frame is used as the new last frame of the independent shot segment. The extended independent shot segments are determined based on the new first frame and the new last frame.

6. The video subtitle erasure method according to claim 1, characterized in that, The step of generating corresponding erased independent shot segments based on the expanded independent shot segments and the target text mask using a preset erasure algorithm includes: The target diffusion model is determined by a preset erasure algorithm; The target diffusion model is used to generate corresponding erased independent shot segments based on the expanded independent shot segments and the target text mask.

7. The video subtitle erasure method according to claim 1, characterized in that, The step of performing a preset post-processing optimization operation on the erased independent shot fragments to obtain target independent shot fragments that meet edge blending conditions includes: A preset edge blending operation is performed on the erased independent shot segments to obtain edge-blended video segments; Perform a preset color equalization operation on the edge-blended video clip to generate a color-equalized video clip; Perform a preset temporal filtering operation on the edge-fused video segment to generate a temporally filtered video segment; A preset video quality enhancement operation is performed on the edge-blended video segment to obtain a target independent shot segment that meets the edge-blending conditions.

8. A video subtitle erasing device, characterized in that, include: The set determination module is used to acquire the video with subtitles to be erased, detect the subtitles in the video with subtitles to be erased, and merge the timestamps corresponding to the same subtitles in the video with subtitles to be erased, so as to determine the set of subtitled segments and the set of subtitle-free segments in the video with subtitles to be erased. The video frame acquisition module is used to split each subtitled segment in the set of subtitled segments into independent shot segments using a preset shot segmentation algorithm, and to perform video frame parsing on the independent shot segments to obtain the first and last frames of the independent shot segments. The mask segmentation module is used to determine the target reference frame based on the first and last frames of the independent shot segment, generate an extended independent shot segment based on the target reference frame and the independent shot segment, and segment the target text mask of the extended independent shot segment. The video generation module is used to generate corresponding erased independent shot segments based on the extended independent shot segments and the target text mask using a preset erasure algorithm, and to perform preset post-processing optimization operations on the erased independent shot segments to obtain target independent shot segments that meet the edge blending conditions, and to integrate the target independent shot segments with the set of segments without subtitles to generate a target video.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the video subtitle erasure method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the video subtitle erasure method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Subtitle elimination method and device, electronic equipment and storage medium

    CN121864926A