Video text clearing method, device, electronic device and storage medium

By extracting the text mask area in the video and performing text clearing processing based on this area, the problem of poor text clearing effect in the prior art is solved, and the accurate removal of irregular shape text is achieved.

CN114598923BActive Publication Date: 2025-05-09BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210228531.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-08
Publication Date
2025-05-09
Estimated Expiration
2042-03-08

AI Technical Summary

Technical Problem

Among the existing video text clearing methods, the text clearing effect is poor, especially the irregular shape text is difficult to accurately identify and clear.

Method used

By responding to the text clearing request for the target video, a text clearing area is obtained, and the target text area image is determined from multiple initial text area images based on the area, thereby extracting the text mask area, and finally performing text clearing processing on the target video based on the text mask area.

Benefits of technology

It realizes accurate identification and removal of irregular shape text in the video, improving the effect of text clearing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114598923B_ABST
    Figure CN114598923B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video text removal method, device, electronic device and storage medium. The method comprises: in response to a text removal request for a target video, obtaining a text removal area corresponding to the text removal request; based on the text removal area, determining a target text area image corresponding to the text removal area from multiple initial text area images of the target video; extracting a text mask area from the target text area image; the text mask area is used to characterize the text area information corresponding to the text contained in the target text area image; based on the text mask area, performing text removal processing on the target video. Compared with the traditional technology in which the user directly selects the text area to be removed, the present disclosure can realize the recognition of irregularly shaped text in the video, thereby further improving the effect of text removal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and in particular to a video text removal method, device, electronic device and storage medium. Background Art

[0002] With the development of video processing technology, a technology for clearing text in videos has emerged. For example, video creators can process the subtitle area in the video frame by frame by cropping or masking to clear the subtitles presented in the video, so that the video with the subtitles cleared can be used for secondary creation.

[0003] In the related art, the current method for clearing text in a video is usually to have a user select a text area to be cleared in the video, and then clear the text area. However, the text in the video often has an irregular shape, and the user cannot accurately select the text to be cleared to ensure that only the text is cleared without processing the background of the text. Therefore, the existing video text clearing method has a poor effect of clearing text. Summary of the invention

[0004] The present disclosure provides a video text removal method, device, electronic device and storage medium to at least solve the problem of poor text removal effect in the related art. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, a method for clearing text from a video is provided, comprising:

[0006] In response to a text clearing request for a target video, obtaining a text clearing area corresponding to the text clearing request;

[0007] Based on the text clearing area, determining a target text area image that is compatible with the text clearing area from a plurality of initial text area images of the target video;

[0008] Extracting a text mask area from the target text area image; the text mask area is used to represent text area information corresponding to the text contained in the target text area image;

[0009] Based on the text mask area, the target video is subjected to text removal processing.

[0010] In an exemplary embodiment, the target video includes multiple video frames; before determining the target text area image corresponding to the text clearing area from the multiple initial text area images of the target video, it also includes: extracting target video frames from the multiple video frames according to a preset video frame interval; extracting an image area carrying text from the target video frame as the initial text area image.

[0011] In an exemplary embodiment, determining a target text area image corresponding to the text clearing area from a plurality of initial text area images of the target video includes: acquiring a plurality of initial text area images corresponding to the target video frame, and an image area corresponding to each initial text area image; acquiring an area overlap between each image area and the text clearing area; and when the area overlap is greater than a preset overlap threshold, using the initial text area image as the target text area image.

[0012] In an exemplary embodiment, extracting the text mask area from the target text area image includes: obtaining a first current video frame from the multiple video frames; when the first current video frame belongs to the target video frame, obtaining a target text area image corresponding to the first current video frame; inputting the target text area image corresponding to the first current video frame into a pre-trained text mask detection network, and obtaining the text mask area corresponding to the first current video frame through the text mask detection network; the text mask detection network is trained based on a sample text mask area of ​​a sample text area image and a sample background area of ​​the sample text area image.

[0013] In an exemplary embodiment, after obtaining the first current video frame from the multiple video frames, it also includes: when the first current video frame does not belong to the target video frame, obtaining an adjacent target video frame adjacent to the first current video frame; obtaining a target text area image corresponding to the adjacent target video frame, and a text mask area corresponding to the adjacent target video frame; obtaining a first image in the adjacent target video frame that matches the text mask area corresponding to the adjacent target video frame, and a second image in the first current video frame that matches the text mask area corresponding to the adjacent target video frame, and obtaining a difference value between the first image and the second image; when the difference value is less than a preset difference value threshold, using the text mask area corresponding to the adjacent target video frame as the text mask area corresponding to the first current video frame.

[0014] In an exemplary embodiment, the target video includes multiple video frames; the text clearing processing is performed on the target video based on the text mask area, including: obtaining a second current video frame from the multiple video frames, and a text mask area corresponding to the second current video frame; when the second current video frame is not the first frame of the multiple video frames, obtaining a previous video frame of the second current video frame from the multiple video frames; based on the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame, the text clearing processing is performed on the second current video frame to obtain a target cleared video frame corresponding to the second current video frame.

[0015] In an exemplary embodiment, the second current video frame is subjected to text clearing processing according to the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame to obtain a target cleared video frame corresponding to the second current video frame, including: obtaining an initial cleared video frame corresponding to the second current video frame according to the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame; obtaining the target cleared video frame corresponding to the previous video frame, and the text mask area corresponding to the previous video frame; inputting the target cleared video frame corresponding to the previous video frame, the text mask area corresponding to the previous video frame, the initial cleared video frame corresponding to the second current video frame, and the text mask area corresponding to the second current video frame into a pre-trained anti-flicker suppression network, and obtaining the target cleared video frame corresponding to the second current video frame through the anti-flicker suppression network.

[0016] According to a second aspect of an embodiment of the present disclosure, a video text clearing device is provided, comprising:

[0017] a clearing area acquisition unit, configured to execute, in response to a text clearing request for a target video, acquiring a text clearing area corresponding to the text clearing request;

[0018] A target image acquisition unit is configured to determine a target text area image adapted to the text clearing area from a plurality of initial text area images of the target video based on the text clearing area;

[0019] The text mask extraction unit is configured to extract a text mask area from the target text area image; the text mask area is used to represent text area information corresponding to the text contained in the target text area image;

[0020] The text removal processing unit is configured to perform text removal processing on the target video based on the text mask area.

[0021] In an exemplary embodiment, the target video includes multiple video frames; the target image acquisition unit is further configured to extract the target video frame from the multiple video frames according to a preset video frame interval; and extract the image area carrying text from the target video frame as an initial text area image.

[0022] In an exemplary embodiment, the target image acquisition unit is further configured to execute the acquisition of multiple initial text area images corresponding to the target video frame, and the image area corresponding to each initial text area image; obtain the area overlap between each image area and the text clearing area; and when the area overlap is greater than a preset overlap threshold, use the initial text area image as the target text area image.

[0023] In an exemplary embodiment, the text mask extraction unit is further configured to execute obtaining a first current video frame from the multiple video frames; in the case where the first current video frame belongs to the target video frame, obtaining a target text area image corresponding to the first current video frame; inputting the target text area image corresponding to the first current video frame into a pre-trained text mask detection network, and obtaining the text mask area corresponding to the first current video frame through the text mask detection network; the text mask detection network is trained based on a sample text mask area of ​​a sample text area image and a sample background area of ​​the sample text area image.

[0024] In an exemplary embodiment, the text mask extraction unit is further configured to execute, when the first current video frame does not belong to the target video frame, obtaining an adjacent target video frame adjacent to the first current video frame; obtaining a target text area image corresponding to the adjacent target video frame, and a text mask area corresponding to the adjacent target video frame; obtaining a first image in the adjacent target video frame that matches the text mask area corresponding to the adjacent target video frame, and a second image in the first current video frame that matches the text mask area corresponding to the adjacent target video frame, and obtaining a difference value between the first image and the second image; and when the difference value is less than a preset difference value threshold, using the text mask area corresponding to the adjacent target video frame as the text mask area corresponding to the first current video frame.

[0025] In an exemplary embodiment, the target video includes multiple video frames; the text clearing processing unit is further configured to execute obtaining a second current video frame from the multiple video frames, and a text mask area corresponding to the second current video frame; when the second current video frame is not the first frame among the multiple video frames, obtaining a previous video frame of the second current video frame from the multiple video frames; based on the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame, performing text clearing processing on the second current video frame to obtain a target cleared video frame corresponding to the second current video frame.

[0026] In an exemplary embodiment, the text clearing processing unit is further configured to execute, based on the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame, to obtain an initial clearing video frame corresponding to the second current video frame; to obtain a target clearing video frame corresponding to the previous video frame, and a text mask area corresponding to the previous video frame; and to input the target clearing video frame corresponding to the previous video frame, the text mask area corresponding to the previous video frame, the initial clearing video frame corresponding to the second current video frame, and the text mask area corresponding to the second current video frame into a pre-trained anti-flicker suppression network, and to obtain the target clearing video frame corresponding to the second current video frame through the anti-flicker suppression network.

[0027] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement a video text clearing method as described in any one of the embodiments of the first aspect.

[0028] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the video text clearing method as described in any one of the embodiments in the first aspect.

[0029] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, wherein the computer program product includes instructions, and when the instructions are executed by a processor of an electronic device, the electronic device is enabled to execute the video text clearing method as described in any one of the embodiments of the first aspect.

[0030] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0031] In response to a text clearing request for a target video, a text clearing area corresponding to the text clearing request is obtained; based on the text clearing area, a target text area image corresponding to the text clearing area is determined from a plurality of initial text area images of the target video; a text mask area is extracted from the target text area image; the text mask area is used to characterize the text area information corresponding to the text contained in the target text area image; based on the text mask area, the target video is subjected to text clearing processing. When clearing text from a video, the present invention can further extract a text mask area from the target text area image after obtaining the target text area image, and implement text clearing processing on the target video based on the text mask area. Compared with the text area directly selected for clearing by the user in the traditional technology, the present invention can realize the recognition of irregularly shaped text in the video, thereby further improving the effect of text clearing.

[0032] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0034] Figure 1 The figure is a flow chart of a method for clearing text from a video according to an exemplary embodiment.

[0035] Figure 2 The figure is a flowchart of determining a target text area image according to an exemplary embodiment.

[0036] Figure 3 The figure is a flow chart of extracting a text mask area according to an exemplary embodiment.

[0037] Figure 4 The figure is a flowchart of extracting a text mask area according to another exemplary embodiment.

[0038] Figure 5 The figure is a flowchart of performing text clearing processing on a target video according to an exemplary embodiment.

[0039] Figure 6 The figure is a flowchart of performing text clearing processing on the second current video frame according to an exemplary embodiment.

[0040] Figure 7 The figure is a flow chart of a subtitle clearing method according to an exemplary embodiment.

[0041] Figure 8 The figure is a block diagram of a device for clearing text from a video according to an exemplary embodiment.

[0042] Fig. 9 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0043] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0044] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0045] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0046] Figure 1 is a flow chart of a method for clearing text from a video according to an exemplary embodiment. Figure 1 As shown, the video text clearing method can be used in a terminal, including the following steps.

[0047] In step S101 , in response to a text clearing request for a target video, a text clearing area corresponding to the text clearing request is acquired.

[0048] The target video refers to a video that needs to be cleared of text, and the text clearing request refers to a request triggered by a user through a terminal to clear part of the text presented in the target video, and the text clearing area is a video image area selected by a user for implementing the text clearing process. In this embodiment, when a user needs to clear part of the text in a video, a text clearing request for the video can be triggered through the terminal, and the text area to be cleared can be selected. The terminal can respond to the request, and the video that needs to be cleared of text is used as the target video, and the area selected by the user is used as the text clearing area.

[0049] For example, when a user needs to clear the subtitles in a video, he can initiate a text clearing request for the video through his terminal, and select the area in the video used to present the subtitles as the text area that needs to be cleared. At this time, the terminal can respond to the request, use the video whose subtitles need to be cleared as the target video, and use the subtitle area selected by the user as the text clearing area.

[0050] In step S102, based on the text clearing area, a target text area image that is compatible with the text clearing area is determined from a plurality of initial text area images of the target video.

[0051] The initial text area image refers to the area image carrying text in the target video. In this embodiment, there may be multiple area images carrying text for the target video. For example, a target video can simultaneously present video subtitles or video title text, etc. Therefore, the target video can carry multiple initial text area images at the same time, and the target text area image refers to the initial text area image that is compatible with the text clearing area selected by the user among the multiple initial text area images. Specifically, after determining the target video, the terminal can use text recognition technology to identify the area image carrying text from the target video as the initial text area image, and can further filter out the initial text area image that is compatible with the text clearing area selected by the user from the identified initial text area image as the target text area image.

[0052] For example, the terminal may obtain multiple initial text area images for the target video through text recognition technology, which may include a subtitle area image presented as a video subtitle and a title area image presented as a video title. If the user needs to clear the video subtitles, the subtitle area of ​​the video can be selected as a text clearing area, and the terminal can then use the presented subtitle area image as the target text area image.

[0053] In step S103, a text mask area is extracted from the target text area image; the text mask area is used to represent text area information corresponding to the text contained in the target text area image;

[0054] In step S104, text removal processing is performed on the target video based on the text mask area.

[0055] The text mask area refers to the text area corresponding to the text contained in the target text area image. In this embodiment, the extracted target text area image includes not only the text image itself but also part of the background area image, and the text mask area refers to the image area corresponding to the text image itself, and the image corresponding to this area only contains the text itself. In this embodiment, after the terminal determines the target text area image, it can further extract the text mask area contained in the target text area image from the target text area image, and finally, after obtaining the text mask area, the terminal can further use the above text mask area to clear the text in the target video, for example, it can fill the text mask area with an image near the text mask area, thereby achieving the clearing of the text represented by the text mask area.

[0056] In the above-mentioned video text clearing method, by responding to a text clearing request for a target video, a text clearing area corresponding to the text clearing request is obtained; based on the text clearing area, a target text area image corresponding to the text clearing area is determined from multiple initial text area images of the target video; a text mask area is extracted from the target text area image; the text mask area is used to characterize the text area information corresponding to the text contained in the target text area image; based on the text mask area, the target video is subjected to text clearing processing. When clearing video text, the present disclosure can further extract a text mask area from the target text area image after obtaining the target text area image, and implement text clearing processing on the target video based on the text mask area. Compared with the text area directly selected for clearing by the user in the traditional technology, the present disclosure can realize the recognition of irregularly shaped text in the video, thereby further improving the effect of text clearing.

[0057] In an exemplary embodiment, the target video includes multiple video frames; before step S102, it can also include: extracting a target video frame from the multiple video frames according to a preset video frame interval; extracting an image area carrying text from the target video frame as an initial text area image.

[0058] In this embodiment, the target video may be composed of multiple video frames, and the target video frame refers to a video frame extracted from multiple video frames according to a preset video frame interval. The video frame interval may be set by the user in advance. For example, the user may set a target video frame every 3 frames. Then, after the terminal obtains the video frames of multiple target videos, it may sample from the video frames of the target video every 3 frames to obtain the corresponding target video frame. Afterwards, the terminal may use the obtained target video frame and input it into a pre-trained text recognition network. Thus, the image area containing text in the target video frame is obtained through the output of the text recognition network as the initial text area image. Compared with the need to input all the video frames in the target video into the text recognition network to obtain the initial text area image, this embodiment can extract the initial text area image of the target video frame by extracting the target video frame and only inputting the target video frame into the initial text area image, which can improve the operation speed of the model and thus improve the efficiency of text removal.

[0059] For example, the target video may include video frame 0, video frame 1, video frame 2, video frame 3, video frame 4, video frame 5, video frame 6 and video frame 7. If the pre-set video frame interval is to extract the target video frame every 3 frames, the terminal can input video frame 0, video frame 3 and video frame 6 as target video frames into the text recognition network to obtain the corresponding initial text area image, without having to input all the video frames into the text recognition network to obtain the corresponding initial text area image, thereby improving the calculation speed of the model and improving the efficiency of text removal.

[0060] In this embodiment, target video frames can be extracted from multiple video frames according to a preset video frame interval, and the corresponding initial text area image is extracted only from the target video frame. Compared with the need to extract the initial text area image from all video frames of the target video, this embodiment can improve the calculation speed of the model and improve the efficiency of text removal.

[0061] Furthermore, if Figure 2 As shown, step S102 may further include:

[0062] In step S201, a plurality of initial text region images corresponding to a target video frame and an image region corresponding to each initial text region image are obtained.

[0063] Among them, the image area corresponding to the initial text area image refers to the area corresponding to each initial text area image in the target video frame. In this embodiment, after obtaining the initial text area image of each target video frame, the terminal can further determine the areas where the above-mentioned initial text area images are located in the corresponding target video frames as the image areas corresponding to each initial text area image.

[0064] In step S202, the area overlap between each image area and the text removal area is obtained;

[0065] In step S203, when the region overlap is greater than a preset overlap threshold, the initial text region image is used as the target text region image.

[0066] The area coincidence refers to the coincidence between the image area corresponding to the initial text area image and the text clearing area. The area coincidence can be obtained by the ratio of the coincidence area between the image area corresponding to each initial text area image and the text clearing area to the area of ​​the image area corresponding to each initial text area image. In this embodiment, after obtaining the image area corresponding to each initial text area image, the terminal can compare the degree of coincidence between the image area corresponding to each initial text area image and the text clearing area, so as to obtain the area coincidence between the image area corresponding to each initial text area image and the text clearing area. After that, the terminal can also compare the obtained area coincidence with the preset coincidence threshold. Only when the area coincidence is greater than the preset coincidence threshold, the initial text area image corresponding to the image area will be used as the target text area image.

[0067] For example, the initial text area image contained in the target video frame may include: text area image A, text area image B, and text area image C, and each initial text area image corresponds to image area A, image area B, and image area C, respectively. Then, the terminal can respectively calculate the area overlap between image area A and the text clearing area selected by the user, the area overlap between image area B and the text clearing area selected by the user, and the area overlap between image area C and the text clearing area selected by the user, and compare whether the area overlaps are greater than a preset overlap threshold. Only when the area overlap is greater than the set overlap threshold, that is, the area overlap between image area B and the text clearing area selected by the user is greater than the overlap threshold, then the terminal can use text area image B as the target text area image, and if the area overlap between image area C and the text clearing area selected by the user is greater than the overlap threshold, then the terminal can use text area image C as the target text area image.

[0068] In this embodiment, after obtaining multiple initial text area images of the target video frame, the terminal can also filter out the target text area image based on the image area corresponding to each initial text area image and the area overlap between the text clearing area, thereby improving the accuracy of the target text area image screening and further improving the accuracy of the text clearing process.

[0069] In an exemplary embodiment, if Figure 3 As shown, step S103 may further include:

[0070] In step S301, a first current video frame is obtained from multiple video frames.

[0071] The first current video frame may be any one of the multiple video frames. In this embodiment, the terminal may select any one of the multiple video frames of the target video as the first current video frame.

[0072] In step S302, when the first current video frame belongs to the target video frame, a target text area image corresponding to the first current video frame is obtained.

[0073] If the first current video frame selected by the terminal in step S301 is a target video frame extracted from multiple video frames according to the video frame interval, since the terminal can extract the corresponding initial text area image from the target video frame and further find the target text area image corresponding to each target video frame, therefore, when the first current video frame belongs to the target video frame, the terminal can also further obtain the target text area image corresponding to the first current video frame.

[0074] In step S303, the target text area image corresponding to the first current video frame is input into a pre-trained text mask detection network, and the text mask area corresponding to the first current video frame is obtained through the text mask detection network; the text mask detection network is trained based on the sample text mask area of ​​the sample text area image and the sample background area of ​​the sample text area image.

[0075] The text mask detection network is a neural network used to detect text masks carried in an image. The text mask detection network can be trained using sample images carrying text, i.e., sample text area images. For example, a user can pre-collect sample text area images for training the text mask detection network, and annotate the sample text area images, and respectively determine the text mask area contained in the sample text area image, i.e., the sample text mask area, and the background area contained in the sample text area image, i.e., the sample background area. The terminal can train a neural network model of text mask areas and background areas for classification using the sample text area image, and the sample text mask areas and sample background areas corresponding to the sample text area image, thereby obtaining a text mask detection network for detecting text mask areas.

[0076] After the text mask detection network training is completed, the target text area image corresponding to the first current video frame can be further input into the trained text mask detection network, so that the text mask area corresponding to the first current video frame can be output by the text mask detection network.

[0077] In this embodiment, if the first current video frame belongs to the target video frame, the target text area image corresponding to the first current video frame can be input into a pre-trained text mask detection network, and the text mask detection network outputs the corresponding text mask area. Since the text mask detection network is based on the sample text mask area of ​​the sample text area image and is trained with the sample background area, the text mask area can output an accurate text mask area, thereby further improving the accuracy of the text mask area corresponding to the first current video frame.

[0078] In addition, if Figure 4 As shown, after step S301, the following steps may also be included:

[0079] In step S401, when the first current video frame does not belong to the target video frame, an adjacent target video frame adjacent to the first current video frame is obtained.

[0080] The adjacent target video frame refers to the target video frame adjacent to the first current video frame. If the first current video frame does not belong to the above target video frame, the terminal can further obtain the target video frame adjacent to the first current video frame as the adjacent target video frame. For example, the target video may include video frame 0, video frame 1, video frame 2, video frame 3, video frame 4, video frame 5, video frame 6 and video frame 7, and video frame 0, video frame 3 and video frame 6 are used as target video frames. If the first current video frame is video frame 1 or video frame 2, since it does not belong to the target video frame, the terminal can use the target video frame adjacent to the above first current video frame in the target video frame as its corresponding adjacent target video frame, that is, video frame 0 and video frame 3 are used as adjacent target video frames of video frame 1 or video frame 2, and if the first current video frame is video frame 4 or video frame 5, the terminal can use video frame 3 and video frame 6 as adjacent target video frames.

[0081] In step S402, a target text region image corresponding to an adjacent target video frame and a text mask region corresponding to an adjacent target video frame are obtained.

[0082] After obtaining the adjacent target video frames corresponding to each first current video frame, the terminal can obtain the initial text area images corresponding to each adjacent target video frame, and can further filter out the target text area images corresponding to each adjacent target video frame. After that, the terminal can also obtain the text mask area corresponding to each target text area image. For example, each target text area image can be input into a pre-trained text mask detection network, and the text mask area corresponding to each target text area image is obtained through the output of the text mask detection network, thereby obtaining the text mask area corresponding to each adjacent target video frame.

[0083] In step S403, a first image in an adjacent target video frame that matches a text mask region corresponding to the adjacent target video frame and a second image in a first current video frame that matches a text mask region corresponding to the adjacent target video frame are obtained, and a difference value between the first image and the second image is obtained;

[0084] In step S404, when the difference value is less than a preset difference value threshold, the text mask area corresponding to the adjacent target video frame is used as the text mask area corresponding to the first current video frame.

[0085] In this embodiment, the first image refers to the image displayed in the text mask area in the adjacent target video frame, that is, it can be the text image itself that needs to be cleared in the target video frame, and the second image is the image displayed in the text mask area of ​​the adjacent target video frame in the first current video frame. If the difference between the first image and the second image is small, that is, the difference value between the first image and the second image is less than the preset difference value threshold, then it can be indicated that the similarity between the first image and the second image is large, that is, the corresponding text image in the adjacent target video frame is also displayed in the first current video frame. Therefore, the text mask area corresponding to the adjacent target video frame can be used as the text mask area corresponding to the first current video frame.

[0086] Taking video frame 1 as the first current video frame as an example, its corresponding adjacent target video frames are video frame 0 and video frame 3, wherein the first image corresponding to video frame 0 is image A, the first image corresponding to video frame 3 is image B, and the second image corresponding to video frame 2 is image C, then the terminal can calculate the difference value between image A and image C, and the difference value between image B and image C respectively. If the difference value between image A and image C is less than the preset difference value threshold, it means that the image displayed by image C in the text mask area corresponding to image A is the same as the image displayed by image A, that is, image C also displays the text corresponding to image A. Therefore, the terminal can use the text mask area corresponding to video frame 0 as the text mask area corresponding to video frame 1, and if the difference value between image B and image C is less than the preset difference value threshold, then the terminal can use the text mask area corresponding to video frame 3 as the text mask area corresponding to video frame 1.

[0087] In this embodiment, if the first current video frame does not belong to the target video frame, this embodiment can perform a difference value comparison between the image displayed in the text mask area of ​​the adjacent target video frame and the image displayed in the text mask area of ​​the first current video frame. If the difference value is less than the set difference value threshold, the text mask area of ​​the adjacent target video frame can be used as the text mask area of ​​the first current video frame. Thereby, there is no need to detect the target text area image of the first current video frame, and its corresponding text mask area can be obtained, thereby further improving the efficiency of video text removal.

[0088] In an exemplary embodiment, the target video includes a plurality of video frames; Figure 5 As shown, step S104 may further include:

[0089] In step S501, a second current video frame and a text mask area corresponding to the second current video frame are obtained from a plurality of video frames.

[0090] The second current video frame can also be any one of multiple video frames. In this embodiment, after obtaining the text mask area corresponding to each video frame in the target video, the terminal can also select any one video frame as the second current video frame, and use the text mask area corresponding to the video frame as the text mask area corresponding to the second current video frame.

[0091] In step S502, when the second current video frame is not the first frame among the multiple video frames, a previous video frame of the second current video frame is obtained from the multiple video frames;

[0092] In step S503, text clearing processing is performed on the second current video frame according to the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame to obtain a target cleared video frame corresponding to the second current video frame.

[0093] The previous video frame refers to the previous video frame of the second current video frame, and the target cleared video frame refers to the video frame displayed after the text clearing process is performed on the video frame of the target video. If the second current video frame is not the first frame among multiple video frames, in addition to using the second current video frame and the text mask area corresponding to the second current video frame to realize text clearing of the second current video frame, in this embodiment, the previous video frame of the second current video frame is further introduced to realize text clearing. The text clearing method can be realized by a neural network model for realizing image clearing, by using the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame as model inputs, and inputting them into the neural network model for image clearing, and the neural network model realizes the text clearing process for the second current video frame. By introducing the previous video frame, the inconsistency of the results in time caused by single-frame text clearing can be alleviated, thereby further improving the effect of text clearing.

[0094] In this embodiment, in the process of performing text clearing processing on the second current video frame to obtain the target cleared video frame, the previous video frame of the second current video frame is further introduced as input, thereby alleviating the temporal inconsistency of the results caused by single-frame text clearing and further improving the effect of text clearing.

[0095] Furthermore, if Figure 6 As shown, step S503 may further include:

[0096] In step S601, an initial cleared video frame corresponding to the second current video frame is obtained according to the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame.

[0097] In this embodiment, obtaining the target cleared video frame corresponding to the second current video frame can be achieved by using a neural network model for achieving text clearing, and the neural network model can be composed of two parts, including a text clearing network for clearing text, and a flicker suppression network for further suppressing flicker problems caused by the lack of temporal continuity, wherein the initial cleared video frame refers to the output result corresponding to the second current video frame obtained by the output of the text clearing network. Specifically, the terminal can input the previous video frame of the second current video frame, the second current video frame, and the text mask area corresponding to the second current video frame into a pre-trained text clearing network, and the text clearing network outputs the initial cleared video frame corresponding to the second current video frame.

[0098] In step S602, a target cleared video frame corresponding to the previous video frame and a text mask area corresponding to the previous video frame are obtained;

[0099] In step S603, the target cleared video frame corresponding to the previous video frame, the text mask area corresponding to the previous video frame, the initial cleared video frame corresponding to the second current video frame, and the text mask area corresponding to the second current video frame are input into a pre-trained anti-flicker suppression network, and the target cleared video frame corresponding to the second current video frame is obtained through the anti-flicker suppression network.

[0100] The target cleared video frame is the final text clearing processing result of the video frame obtained by outputting the flicker suppression network. In this embodiment, after obtaining the initial cleared video frame corresponding to the second current video frame, the terminal can further obtain the target cleared video frame corresponding to the previous video frame of the second current video frame, and the text mask area corresponding to the previous video frame, and the target cleared video frame of the previous video frame, the text mask area of ​​the previous video frame, the initial cleared video frame of the second current video frame, and the text mask area of ​​the second current video frame can be used as model inputs and input into a pre-trained anti-flicker suppression network, which is composed of an encoder (Encoder)-gated recurrent unit (GRU)-decoder (Decoder), wherein the gated recurrent unit can carry implicit values, and the target cleared video frame corresponding to the second current video frame can be obtained through the anti-flicker suppression network.

[0101] Taking video frame 2 as the second current video frame as an example, video frame 1 can be used as the previous video frame of the second current video frame. The terminal can first input video frame 1, video frame 2 and the text mask area corresponding to video frame 2 into the text clearing network for clearing text. The initial cleared video frame corresponding to video frame 2 can be obtained through the text clearing network, and the target cleared video frame corresponding to video frame 1, the text mask area corresponding to video frame 1, the initial cleared video frame corresponding to video frame 2, and the text mask area corresponding to video frame 2 can be further input into the pre-trained anti-flicker suppression network, so that the anti-flicker suppression network can be used to obtain the target cleared video frame corresponding to video frame 2.

[0102] In this embodiment, when performing text clearing processing on the second current video frame, an anti-flicker suppression network is further introduced, so as to further suppress the flicker problem caused by the lack of temporal continuity and further improve the effect of text clearing.

[0103] In an exemplary embodiment, a subtitle clearing method for clearing video subtitles is also provided, such as Figure 7 As shown, the method may specifically include the following steps:

[0104] (1) Subtitle area detection

[0105] First, if you need to clear subtitles, you must determine the specific range of the subtitle area. This method has strict speed requirements, and for a video, the time required to detect subtitles is linearly related to the number of frames that call the subtitle detection algorithm. Therefore, for subtitle area detection, a frame skipping detection method is designed. That is, for every n video frames of cache length, only the subtitle area in the nth video frame is detected.

[0106] When detecting subtitles, a general algorithm in the field of text recognition is used to detect the rectangular subtitle area, namely:

[0107] B i =D(I i ),

[0108] Among them, D() represents the subtitle detection network, I i is the i-th video frame, B i is the set of subtitle rectangular areas detected in the i-th frame.

[0109] (2) Subtitle Mask Detection Network

[0110] For the detected subtitle rectangular area, it is also necessary to accurately obtain the specific mask of the subtitle as the input of the subsequent clearing module. This can be done by using a classification network to classify the subtitles and background as two different scenes, namely:

[0111]

[0112] in, is the subtitle mask corresponding to the k-th subtitle rectangular area in the i-th frame. E() is the designed subtitle mask recognition network. G() is the subtitle mask recognition network from I i Frame and corresponding rectangular area The region used to input into the subtitle mask recognition network is obtained.

[0113] (3) Subtitle tracking

[0114] At the same time, since subtitle detection is not frame-by-frame, a letter tracking model is also designed. According to the unified subtitle scene, the detected subtitle mask area value remains unchanged, but the background area may change. Given a threshold σ, for I i The adjacent frame I i±1 , if you have:

[0115]

[0116] That is, the adjacent frame and Some subtitles are continuous, where Diff() is the difference between frames.

[0117] Therefore, the starting and ending positions of the unified subtitles in the video frames of the n-th buffer length can be propagated from the n-th frame to the head and tail frames in the video frames of the n-th buffer length, and the starting and ending positions of the unified subtitles in the video frames of the n-th buffer length can be obtained.

[0118] (4) Subtitle clearing

[0119] In order to make the network's clearing ability fast and stable, the designed clearing algorithm is based on a two-stage network structure. In order to achieve lightweight, the convolutional layers, attention mechanisms and other structures in the network are greatly reduced, but similar effects are maintained. In addition, unlike traditional image clearing algorithms, in order to reduce the inconsistency of the results in time due to single-frame clearing, the input of the previous frame is added to ensure the consistency of the output results:

[0120]

[0121] in is the output result after erasing the i frame, and F() is the designed clearing network.

[0122] After that, in order to further suppress the flicker problem caused by the lack of temporal continuity, a lightweight and high-speed anti-flicker suppression network is added before obtaining the final result, which consists of an encoder (Encoder)-gated recurrent unit (GRU)-decoder (Decoder), that is,

[0123]

[0124] For the final stable output result, H is the implicit value required in the gated recurrent unit, and C() is the designed anti-flicker suppression network.

[0125] Through the above embodiment, compared with the existing subtitle clearing algorithm, the subtitle range can be accurately and quickly detected, and subtitles in different languages ​​and fonts can be accurately identified, so that subtitles can be cleared with high quality and stability.

[0126] It should be understood that although Figure 1-Figure 7 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1-Figure 7 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0127] It can be understood that the same / similar parts between the various embodiments of the above method in this specification can refer to each other, and each embodiment focuses on the differences from other embodiments. For related points, please refer to the description of other method embodiments.

[0128] Figure 8 is a block diagram of a video text removal device according to an exemplary embodiment. Figure 8 The device includes a clearing area acquisition unit 801, a target image acquisition unit 802, a text mask extraction unit 803 and a text clearing processing unit 804.

[0129] The clearing area acquisition unit 801 is configured to execute, in response to a text clearing request for a target video, to acquire a text clearing area corresponding to the text clearing request;

[0130] The target image acquisition unit 802 is configured to determine a target text region image that is compatible with the text removal region from a plurality of initial text region images of the target video based on the text removal region;

[0131] The text mask extraction unit 803 is configured to extract a text mask area from the target text area image; the text mask area is used to represent text area information corresponding to the text contained in the target text area image;

[0132] The text removal processing unit 804 is configured to perform text removal processing on the target video based on the text mask area.

[0133] In an exemplary embodiment, the target video includes multiple video frames; the target image acquisition unit 802 is further configured to extract the target video frame from the multiple video frames according to a preset video frame interval; and extract the image area carrying text from the target video frame as an initial text area image.

[0134] In an exemplary embodiment, the target image acquisition unit 802 is further configured to execute the acquisition of multiple initial text area images corresponding to the target video frame, and the image area corresponding to each initial text area image; to obtain the area overlap between each image area and the text clearing area; and when the area overlap is greater than a preset overlap threshold, to use the initial text area image as the target text area image.

[0135] In an exemplary embodiment, the text mask extraction unit 803 is further configured to execute obtaining a first current video frame from multiple video frames; when the first current video frame belongs to a target video frame, obtaining a target text area image corresponding to the first current video frame; inputting the target text area image corresponding to the first current video frame into a pre-trained text mask detection network, and obtaining a text mask area corresponding to the first current video frame through the text mask detection network; the text mask detection network is trained based on a sample text mask area of ​​a sample text area image and a sample background area of ​​the sample text area image.

[0136] In an exemplary embodiment, the text mask extraction unit 803 is also configured to execute, when the first current video frame does not belong to the target video frame, obtaining an adjacent target video frame adjacent to the first current video frame; obtaining a target text area image corresponding to the adjacent target video frame, and a text mask area corresponding to the adjacent target video frame; obtaining a first image in the adjacent target video frame that matches the text mask area corresponding to the adjacent target video frame, and a second image in the first current video frame that matches the text mask area corresponding to the adjacent target video frame, and obtaining a difference value between the first image and the second image; when the difference value is less than a preset difference value threshold, using the text mask area corresponding to the adjacent target video frame as the text mask area corresponding to the first current video frame.

[0137] In an exemplary embodiment, the target video includes multiple video frames; the text clearing processing unit 804 is further configured to execute obtaining a second current video frame and a text mask area corresponding to the second current video frame from the multiple video frames; when the second current video frame is not the first frame among the multiple video frames, obtaining the previous video frame of the second current video frame from the multiple video frames; performing text clearing processing on the second current video frame according to the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame to obtain a target cleared video frame corresponding to the second current video frame.

[0138] In an exemplary embodiment, the text clearing processing unit 804 is further configured to execute, based on the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame, to obtain an initial clearing video frame corresponding to the second current video frame; obtain a target clearing video frame corresponding to the previous video frame, and a text mask area corresponding to the previous video frame; input the target clearing video frame corresponding to the previous video frame, the text mask area corresponding to the previous video frame, the initial clearing video frame corresponding to the second current video frame, and the text mask area corresponding to the second current video frame into a pre-trained anti-flicker suppression network, and obtain the target clearing video frame corresponding to the second current video frame through the anti-flicker suppression network.

[0139] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0140] Fig. 9 1 is a block diagram of an electronic device 900 for clearing video text according to an exemplary embodiment. For example, the electronic device 900 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0141] Reference Fig. 9 , the electronic device 900 may include one or more of the following components: a processing component 902 , a memory 904 , a power component 906 , a multimedia component 908 , an audio component 910 , an input / output (I / O) interface 912 , a sensor component 914 , and a communication component 916 .

[0142] The processing component 902 generally controls the overall operation of the electronic device 900, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 902 may include one or more processors 920 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 902 may include one or more modules to facilitate the interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate the interaction between the multimedia component 908 and the processing component 902.

[0143] The memory 904 is configured to store various types of data to support operations on the electronic device 900. Examples of such data include instructions for any application or method operating on the electronic device 900, contact data, phone book data, messages, pictures, videos, etc. The memory 904 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, an optical disk, or a graphene memory.

[0144] The power supply component 906 provides power to the various components of the electronic device 900. The power supply component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 900.

[0145] The multimedia component 908 includes a screen that provides an output interface between the electronic device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the electronic device 900 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0146] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC), and when the electronic device 900 is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 904 or sent via the communication component 916. In some embodiments, the audio component 910 also includes a speaker for outputting an audio signal.

[0147] I / O interface 912 provides an interface between processing component 902 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0148] The sensor assembly 914 includes one or more sensors for providing various aspects of status assessment for the electronic device 900. For example, the sensor assembly 914 can detect the open / closed state of the electronic device 900, the relative positioning of the components, such as the display and keypad of the electronic device 900, and the sensor assembly 914 can also detect the position change of the electronic device 900 or the electronic device 900 components, the presence or absence of contact between the user and the electronic device 900, the orientation or acceleration / deceleration of the device 900 and the temperature change of the electronic device 900. The sensor assembly 914 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 914 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 914 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0149] The communication component 916 is configured to facilitate wired or wireless communication between the electronic device 900 and other devices. The electronic device 900 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0150] In an exemplary embodiment, the electronic device 900 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0151] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 904 including instructions, and the above instructions can be executed by the processor 920 of the electronic device 900 to perform the above method. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0152] In an exemplary embodiment, a computer program product is further provided. The computer program product includes instructions. The instructions can be executed by the processor 920 of the electronic device 900 to complete the above method.

[0153] It should be noted that the above-mentioned devices, electronic devices, computer-readable storage media, computer program products, etc. may also include other implementation methods according to the description of the method embodiments. The specific implementation methods can refer to the description of the relevant method embodiments, which will not be described one by one here.

[0154] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.

[0155] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A video text removal method, characterized in that: include: In response to a text clearing request for a target video, obtaining a text clearing area corresponding to the text clearing request; Extracting a target video frame from a plurality of video frames of the target video according to a preset video frame interval; extracting an image region carrying text from the target video frame as a plurality of initial text region images of the target video; Based on the text clearing area, determining a target text area image that is compatible with the text clearing area from a plurality of initial text area images of the target video; Extracting a text mask area from the target text area image; The text mask area is used to characterize the text area information corresponding to the text contained in the target text area image; including: obtaining a first current video frame from the multiple video frames; when the first current video frame belongs to the target video frame, obtaining the target text area image corresponding to the first current video frame; inputting the target text area image corresponding to the first current video frame into a pre-trained text mask detection network, and obtaining the text mask area corresponding to the first current video frame through the text mask detection network; the text mask detection network is trained based on the sample text mask area of ​​the sample text area image and the sample background area of ​​the sample text area image; when the first current video frame is not In the case of the target video frame, an adjacent target video frame adjacent to the first current video frame is obtained; a target text area image corresponding to the adjacent target video frame and a text mask area corresponding to the adjacent target video frame are obtained; a first image matching the text mask area corresponding to the adjacent target video frame in the adjacent target video frame and a second image matching the text mask area corresponding to the adjacent target video frame in the first current video frame are obtained, and a difference value between the first image and the second image is obtained; in the case where the difference value is less than a preset difference value threshold, the text mask area corresponding to the adjacent target video frame is used as the text mask area corresponding to the first current video frame; Based on the text mask area, the target video is subjected to text removal processing.

2. The method according to claim 1, characterized in that The step of determining a target text area image that is compatible with the text clearing area from a plurality of initial text area images of the target video includes: Acquire a plurality of initial text region images corresponding to the target video frame, and an image region corresponding to each initial text region image; Obtaining the area overlap between each image area and the text clearing area; When the area overlap is greater than a preset overlap threshold, the initial text area image is used as the target text area image.

3. The method according to claim 1 or 2, characterized in that: The target video includes a plurality of video frames; and performing text removal processing on the target video based on the text mask area includes: Acquire a second current video frame and a text mask area corresponding to the second current video frame from the multiple video frames; When the second current video frame is not the first frame of the multiple video frames, obtaining a previous video frame of the second current video frame from the multiple video frames; According to the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame, the second current video frame is subjected to text clearing processing to obtain a target cleared video frame corresponding to the second current video frame.

4. The method according to claim 3, characterized in that The step of performing text clearing processing on the second current video frame according to the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame to obtain a target cleared video frame corresponding to the second current video frame includes: Obtaining an initial cleared video frame corresponding to the second current video frame according to the previous video frame, the second current video frame, and a text mask area corresponding to the second current video frame; Obtaining a target cleared video frame corresponding to the previous video frame and a text mask area corresponding to the previous video frame; The target cleared video frame corresponding to the previous video frame, the text mask area corresponding to the previous video frame, the initial cleared video frame corresponding to the second current video frame, and the text mask area corresponding to the second current video frame are input into a pre-trained anti-flicker suppression network, and the target cleared video frame corresponding to the second current video frame is obtained through the anti-flicker suppression network.

5. A video text removal device, characterized in that: include: a clearing area acquisition unit, configured to execute, in response to a text clearing request for a target video, acquiring a text clearing area corresponding to the text clearing request; The target image acquisition unit is configured to extract a target video frame from a plurality of video frames of the target video according to a preset video frame interval; extract an image region carrying text from the target video frame as a plurality of initial text region images of the target video; A target image acquisition unit is configured to determine a target text area image adapted to the text clearing area from a plurality of initial text area images of the target video based on the text clearing area; A text mask extraction unit is configured to extract a text mask area from the target text area image; The text mask area is used to characterize the text area information corresponding to the text contained in the target text area image; it is further configured to execute obtaining a first current video frame from the multiple video frames; when the first current video frame belongs to the target video frame, obtaining the target text area image corresponding to the first current video frame; inputting the target text area image corresponding to the first current video frame into a pre-trained text mask detection network, and obtaining the text mask area corresponding to the first current video frame through the text mask detection network; the text mask detection network is obtained by training based on a sample text mask area of ​​a sample text area image and a sample background area of ​​the sample text area image; in the first current video frame In the case where the first current video frame does not belong to the target video frame, an adjacent target video frame adjacent to the first current video frame is obtained; a target text area image corresponding to the adjacent target video frame and a text mask area corresponding to the adjacent target video frame are obtained; a first image matching the text mask area corresponding to the adjacent target video frame in the adjacent target video frame and a second image matching the text mask area corresponding to the adjacent target video frame in the first current video frame are obtained, and a difference value between the first image and the second image is obtained; in the case where the difference value is less than a preset difference value threshold, the text mask area corresponding to the adjacent target video frame is used as the text mask area corresponding to the first current video frame; The text removal processing unit is configured to perform text removal processing on the target video based on the text mask area.

6. The device according to claim 5, characterized in that The target image acquisition unit is further configured to acquire a plurality of initial text area images corresponding to the target video frame, and an image area corresponding to each initial text area image; and acquire a degree of area overlap between each image area and the text clearing area; When the region overlap is greater than a preset overlap threshold, the initial text region image is used as the target text region image.

7. The device according to claim 5 or 6, characterized in that The target video includes multiple video frames; the text clearing processing unit is further configured to execute obtaining a second current video frame and a text mask area corresponding to the second current video frame from the multiple video frames; when the second current video frame is not the first frame among the multiple video frames, obtaining a previous video frame of the second current video frame from the multiple video frames; and performing text clearing processing on the second current video frame according to the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame to obtain a target cleared video frame corresponding to the second current video frame.

8. The device according to claim 7, characterized in that The text clearing processing unit is further configured to execute, based on the previous video frame, the second current video frame, and the text mask area corresponding to the second current video frame, to obtain an initial clearing video frame corresponding to the second current video frame; obtain a target clearing video frame corresponding to the previous video frame, and a text mask area corresponding to the previous video frame; input the target clearing video frame corresponding to the previous video frame, the text mask area corresponding to the previous video frame, the initial clearing video frame corresponding to the second current video frame, and the text mask area corresponding to the second current video frame into a pre-trained anti-flicker suppression network, and obtain the target clearing video frame corresponding to the second current video frame through the anti-flicker suppression network.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the video text clearing method as described in any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the video text clearing method as claimed in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method and device for eliminating target image in video, electronic equipment and storage medium

    CN111179159A

  • Data collection method and device, storage medium and electronic equipment

    CN111445902A

  • Text elimination method and device, electronic device and storage medium

    CN112183294A

  • Video trace removing method and video trace removing device

    CN112233055A

  • Image processing method and related equipment thereof

    CN114004751A