Video data processing method and device, readable storage medium and program product
By extracting optical flow data of video frame sequence and splicing local optical flow data, the text removal model is trained, and the problem of flickering between video frames is solved, automated and flicker-free text removal is achieved, and the visual effect and user experience of video processing is improved.
Patent Information
- Application Number
- CN202510295078.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-04
AI Technical Summary
In the existing video data processing methods, obvious flickering problems are prone to occur between frames after removing text, resulting in poor visual effects, cumbersome operation, time-consuming and labor-consuming.
By extracting the optical flow data of the video frame sequence, the local optical flow data of the text area is determined, and splicing it with the video frame sequence to obtain the sample video frame sequence. The text removal model is trained based on the sample video frame sequence to achieve the effect of automatically removing text.
It effectively alleviates the flickering problem after removing text between frames, improves the visual effect of video data processing, improves user experience, and simplifies the operation process.
Smart Images

Figure CN120260049A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a video data processing method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] With the development of computer technology and Internet technology, users' demands for personalized images and videos are increasing day by day, which are widely used in scenarios such as job hunting, exam registration, social networking, and entertainment. In daily life, text is usually added to video or picture content for information annotation such as subtitles and copyright. However, in some cases, this text may interfere with the appreciation or analysis of visual content. Users may hope to remove the text or subtitles in the video or picture for different considerations such as aesthetics, copyright regulations, and privacy protection.
[0003] However, in the current video data processing methods, the methods for removing text in videos or pictures usually rely on image editing software. For example, Adobe Photoshop (abbreviated as PS, an image processing software developed and distributed by Adobe Systems). This usually requires users to manually select the text area in the video and process it frame by frame. It is inconvenient, time-consuming, and laborious for users. In addition, there will be many defects after frame-by-frame smearing. For example, there is an obvious flickering problem between frames, which leads to a poor visual effect of the smeared video. Therefore, how to alleviate the flickering problem after text removal between frames in the video and effectively improve the visual effect of video data processing has become an urgent problem to be solved. Summary of the Invention
[0004] Based on this, this application provides a video data processing method, apparatus, computer device, computer-readable storage medium, and computer program product, which can effectively alleviate the flickering problem after text removal between frames in the video, thereby effectively improving the visual effect of video data processing, and at the same time enhancing the user experience and bringing convenience to users.
[0005] On the one hand, this application provides a video data processing method, including: extracting optical flow data corresponding to a video frame sequence; determining local optical flow data for a text area in the video frame sequence based on the optical flow data; splicing the local optical flow data with the video frame sequence to obtain a sample video frame sequence; training an initial text removal model based on the sample video frame sequence to obtain a text removal model; and processing a target video frame sequence through the text removal model to obtain the target video frame sequence without the text area.
[0006] In one embodiment, extracting the optical flow data corresponding to the video frame sequence includes: processing the video frame sequence through an optical flow estimation model to obtain the optical flow data between every two adjacent frames in the video frame sequence; using the optical flow data between every two adjacent frames in the video frame sequence as the optical flow data corresponding to the video frame sequence.
[0007] In one embodiment, determining the local optical flow data for the text region in the video frame sequence based on the optical flow data includes: determining the binary mask data for the text region in the video frame sequence; determining the local optical flow data for the text region in the video frame sequence based on the optical flow data and the binary mask data.
[0008] In one embodiment, determining the binary mask data for the text region in the video frame sequence includes: detecting the video frame sequence through a pre-trained text detection model to obtain the text detection boxes for each frame image in the video frame sequence; converting the text detection boxes into the binary mask data for the text regions of each frame image, and using the binary mask data for the text regions of each frame image as the binary mask data for the text region in the video frame sequence.
[0009] In one embodiment, determining the local optical flow data for the text region in the video frame sequence based on the optical flow data and the binary mask data includes: determining the product between the optical flow data and the binary mask data; using the product as the local optical flow data for the text region in the video frame sequence.
[0010] In one embodiment, the color channel of the local optical flow data includes a target color channel, and the color channels of each frame image in the video frame sequence include a first color channel, a second color channel, and a third color channel; splicing the local optical flow data with the video frame sequence to obtain a sample video frame sequence includes: splicing the values of the target color channel, the values of the first color channel, the values of the second color channel, and the values of the third color channel according to the color channels to obtain the video frame sequence with the values of the target color channel superimposed; using the video frame sequence with the values of the target color channel superimposed as the sample video frame sequence.
[0011] In one embodiment, training the initial text removal model based on the sample video frame sequence to obtain a text removal model includes: using the sample video frame sequence as an input parameter to train the initial text removal model until the training stops when the target loss value determined based on the sample video frame sequence and the video frame sequence with removed text output by the initial text removal model meets a preset loss condition, thereby obtaining the trained text removal model; wherein, the target loss value is determined by the sum of a first loss value and a second loss value, the first loss value is determined based on the difference between the non-text regions of each frame image in the sample video frame sequence and the non-text regions of each frame image in the video frame sequence with removed text, and the second loss value is determined based on the difference between the text regions of every two adjacent frames in the video frame sequence with removed text.
[0012] On the one hand, the present application also provides a video data processing device, including: an extraction module configured to extract optical flow data corresponding to a video frame sequence; a determination module configured to determine local optical flow data for the text regions in the video frame sequence based on the optical flow data; a splicing module configured to splice the local optical flow data with the video frame sequence to obtain a sample video frame sequence; a training module configured to train an initial text removal model based on the sample video frame sequence to obtain a text removal model; and a processing module configured to process a target video frame sequence through the text removal model to obtain the target video frame sequence without the text regions.
[0013] On the one hand, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: extracting optical flow data corresponding to a video frame sequence; determining local optical flow data for the text regions in the video frame sequence based on the optical flow data; splicing the local optical flow data with the video frame sequence to obtain a sample video frame sequence; training an initial text removal model based on the sample video frame sequence to obtain a text removal model; and processing a target video frame sequence through the text removal model to obtain the target video frame sequence without the text regions.
[0014] On the one hand, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: extracting optical flow data corresponding to a video frame sequence; determining local optical flow data for a text region in the video frame sequence based on the optical flow data; splicing the local optical flow data with the video frame sequence to obtain a sample video frame sequence; training an initial text removal model based on the sample video frame sequence to obtain a text removal model; and processing a target video frame sequence through the text removal model to obtain the target video frame sequence without the text region.
[0015] On the one hand, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented: extracting optical flow data corresponding to a video frame sequence; determining local optical flow data for a text region in the video frame sequence based on the optical flow data; splicing the local optical flow data with the video frame sequence to obtain a sample video frame sequence; training an initial text removal model based on the sample video frame sequence to obtain a text removal model; and processing a target video frame sequence through the text removal model to obtain the target video frame sequence without the text region.
[0016] For the above video data processing method, apparatus, computer device, computer-readable storage medium and computer program product, by extracting the optical flow data corresponding to the video frame sequence, determining the local optical flow data for the text region in the video frame sequence based on the optical flow data, and splicing the local optical flow data with the video frame sequence to obtain a sample video frame sequence; further, training an initial text removal model based on the sample video frame sequence to obtain a text removal model, and processing the target video frame sequence through the text removal model to obtain the target video frame sequence without the text region. Since the optical flow data can be used to accurately estimate the position change of pixel points in each frame of the video frame sequence, the local optical flow data determined for the text region in the video frame sequence based on the optical flow data corresponding to the video frame sequence can be used to accurately estimate the position change of pixel points in the text region of each frame of the video frame sequence. Furthermore, the sample video frame sequence obtained by splicing the local optical flow data with the original video frame sequence is more suitable as sample data for training the initial text removal model. When the trained text removal model processes the target video frame sequence, it can achieve the effect of clean and non-flashing text removal, effectively alleviating the flashing problem between frames after text removal in the video, thereby effectively improving the visual effect of video data processing and also enhancing the user experience, bringing convenience to users. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0018] Figure 1 It is an application environment diagram of the video data processing method in an embodiment;
[0019] Figure 2 It is a schematic flowchart of the video data processing method in an embodiment;
[0020] Figure 3 It is a schematic diagram of the display interface of the video data processing method on the product side in an embodiment;
[0021] Figure 4 It is a schematic diagram of the overall flowchart of the video data processing method in an embodiment;
[0022] Figure 5 It is a structural block diagram of the video data processing device in an embodiment;
[0023] Figure 6 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0024] In order to make the purpose, technical solutions and beneficial effects of the present application clearer and more understandable, the following further details the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0025] The video data processing method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed on the cloud or other network servers. The server 104 can be the background server of an image application. After the terminal 102 obtains a video frame sequence, the terminal can send the obtained video frame sequence to the background server of the image application, that is, the server 104, so that the server 104 extracts the optical flow data corresponding to the video frame sequence, and determines the local optical flow data for the text area in the video frame sequence based on the optical flow data, splices the local optical flow data with the video frame sequence to obtain a sample video frame sequence; further, the server 104 can train an initial text removal model based on the sample video frame sequence to obtain a text removal model, and process the target video frame sequence through the text removal model to obtain a target video frame sequence without a text area, and return the target video frame sequence without a text area to the terminal 102, so that the terminal 102 displays the target video frame sequence without a text area.
[0026] Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0027] In an exemplary embodiment, as Figure 2 shown, a video data processing method is provided. Taking the terminal in Figure 1 as an example for illustration, it includes the following steps 202 to step 210. Among them:
[0028] Step 202, extract the optical flow data corresponding to the video frame sequence.
[0029] Among them, the video frame sequence refers to a sequence composed of multiple consecutive pictures. For example, in this application, the video frame sequence can be obtained by frame-dividing different videos. It can be understood that the video frame sequence in this application can be obtained by frame-dividing videos such as local videos and short videos intercepted from the network.
[0030] Optical flow data refers to the data obtained after performing optical flow calculation (estimation) on a specified video frame sequence. The optical flow data in this application can be an optical flow map. For example, the optical flow data in this application can be the optical flow map between every two adjacent frames (pictures) in the video frame sequence obtained after processing the video frame sequence using the RAFT model (a model for optical flow estimation). It can be understood that the optical flow data in this application can be used to reflect the position change of pixel points between every two adjacent frames (pictures) in the video frame sequence.
[0031] Step 204: Determine local optical flow data for the text regions in the video frame sequence based on the optical flow data.
[0032] Among them, the text region refers to the text regions in each frame (picture) of the video frame sequence. For example, the text region in this application can be a text detection box. For example, the terminal can use a pre-trained text detection model to detect the video frame sequence, obtain the text detection box in each frame (picture) of the video frame sequence, and use the text detection box in each frame (picture) as the text region in the video frame sequence.
[0033] The local optical flow data refers to the optical flow data corresponding to the text regions in the video frame sequence. For example, the local optical flow data in this application can be obtained by converting the text detection box into binary mask data and then calculating the binary mask data and the optical flow data.
[0034] Step 206: Stitch the local optical flow data with the video frame sequence to obtain a sample video frame sequence.
[0035] Among them, stitching refers to the operation of superimposing the local optical flow data and each frame (picture) in the video frame sequence according to the color channels. For example, assuming that the local optical flow data is a picture with one color channel, and each frame (picture) in the video frame sequence is an RGB picture, that is, a picture with three color channels, then stitching the local optical flow data with the video frame sequence means superimposing the picture with one color channel and the RGB pictures with three color channels (each frame in the video frame sequence) according to the color channels to obtain continuous frame pictures with four color channels (i.e., each frame in the stitched video frame sequence). The continuous frame pictures with four color channels are the sample video frame sequence.
[0036] The sample video frame sequence refers to the video frame sequence obtained by stitching the local optical flow data of the text regions to the original video frame sequence. For example, each frame (picture) in the original video frame sequence in this application is an RGB picture with three color channels, and each frame (picture) in the sample video frame sequence is a picture with four color channels (i.e., the color channel with the local optical flow data superimposed).
[0037] Step 208: Train the initial text removal model based on the sample video frame sequence to obtain the text removal model.
[0038] The initial text removal model refers to an untrained model, and the text removal model refers to a trained model. In this application, the trained text removal model can be used to automatically remove text (such as subtitles) from a video (video frame sequence). For example, the initial text removal model in this application can be a general video model 3D-Unet.
[0039] Optionally / Exemplarily, devices used by different users (operating objects) can all interact with a video application (or video generation system). When a user (operating object) wants to generate a personalized video (such as a video without subtitles), the user can trigger an operation to open the video application (Application, APP) on the terminal and enter the video generation page of the video application through a selection operation, that is, the user can log in to the video application (such as a video generation application) through the trigger operation. Further, in the generation page of the video application displayed on the terminal, the user can trigger an information input operation or an information selection operation, so that the terminal responds to the above information input operation or information selection operation triggered by the user, obtains the target video frame sequence (video) input by the user, and the target video frame sequence (video) contains text (such as pure English subtitles). Further, the terminal can respond to the text removal request triggered by the operating object in the generation page of the video application, call the pre-trained text removal model for removing text, and process the target video frame sequence through the text removal model, and then the target video frame sequence without text areas (i.e., with text removed) can be obtained.
[0040] When training the text removal model for removing text, the terminal can obtain the original video frame sequence and extract the optical flow data corresponding to the video frame sequence. Further, the terminal can determine the local optical flow data for the text area in the video frame sequence based on the extracted optical flow data, and splice the local optical flow data with the video frame sequence to obtain the sample video frame sequence; the terminal trains the initial text removal model based on the sample video frame sequence, and then the trained text removal model can be obtained.
[0041] It can be understood that the method provided in this application can be implemented through the interaction between the terminal and the background server of the video application, or through the interaction between the front end and the back end of the terminal. That is, the front end of the terminal is used to display the target video frame sequence without text areas (i.e., with text removed), and the back end of the terminal is equivalent to the background server, which is used to train the initial text removal model and call the trained text removal model to perform logical processing such as text removal on the target video frame sequence.
[0042] For example, the scenario of removing subtitles from a personalized video will be used as an example for illustration. As Figure 3 shown, it is a schematic diagram of the display interface on the product side of the video data processing method provided by this application. That is, when user A (the operation object) wants to remove the subtitles from a personalized video, user A can open video application A on the terminal through a trigger operation, and enter the main page of video application A as shown in Figure 3 through a selection operation. That is, user A can log in to video application A through a trigger operation. Further, user A can trigger an operation in the main page shown in Figure 3 to trigger the display of a page for starting model training. For example, in the main page shown in Figure 3 displayed on the terminal, user A can click on the "Train" control in the left figure shown in Figure 3 so that the terminal responds to the above click operation of user A and displays the page for starting model training shown on the right in Figure 3 . After user A completes the training input operation (such as determining the original video frame sequence for training input), user A can further click on the "Start" control in the page for starting model training shown on the right in Figure 3 to initiate a model training request. Then the terminal responds to the click "Start" operation triggered by user A in this page (i.e., the model training request), obtains the input original video frame sequence, extracts the optical flow data corresponding to the original video frame sequence through an optical flow estimation model. Further, the terminal determines the local optical flow data for the text region in the original video frame sequence based on the extracted optical flow data, and splices the local optical flow data with the video frame sequence to obtain the final sample video frame sequence for input into the initial model. That is, the terminal can use the obtained sample video frame sequence as an input parameter to train the initial text removal model until the target loss value determined based on the sample video frame sequence and the text-removed video frame sequence output by the initial text removal model satisfies the preset loss condition and then stops training, thus obtaining the trained text removal model. Among them, the target loss value is determined by the sum of the first loss value and the second loss value. The first loss value is determined based on the difference between the non-text regions of each frame image in the input sample video frame sequence and the non-text regions of each frame image in the output text-removed video frame sequence. The second loss value is determined based on the difference between every two adjacent frames in the output text-removed video frame sequence.
[0043] Step 210, process the target video frame sequence through the text removal model to obtain a target video frame sequence without text regions.
[0044] Among them, the target video frame sequence refers to the video frame sequence corresponding to the target video to be processed. It can be understood that the target video frame sequence in this application can be the video frame sequence corresponding to any video custom-selected by the user.
[0045] The target video frame sequence without text regions refers to the target video frame sequence with text removed obtained after being processed by the text removal model. For example, if the target video contains 10 consecutive frames, after processing each frame in the target video frame sequence corresponding to the target video by the text removal model, 10 consecutive frames with text removed are obtained, that is, a target video frame sequence (10 consecutive frames) without text regions is obtained.
[0046] Specifically, after the model training is completed, when the user (the operating object) selects the target video for which they want to remove text, the terminal responds to the above selection operation triggered by the user, obtains the target video frame sequence corresponding to the target video selected by the user, and the target video frame sequence (video) contains text (such as pure English subtitles). Further, the terminal can call the pre-trained text removal model for removing text and process the target video frame sequence through this text removal model, and then a target video frame sequence without text regions (i.e., with text removed) can be obtained.
[0047] For example, as Figure 4 shown, it is a schematic diagram of the overall process of the video data processing method provided by this application. When the text removal model as Figure 4 shown is trained, the terminal can directly call the trained text removal model to process any video frame sequence (or single-frame image), and then a target video frame sequence (or single-frame image) with text removed can be obtained. That is, the method provided in the embodiments of this application can automatically achieve the effect of text removal applicable to pictures or videos without other additional operations.
[0048] In this embodiment, by extracting the optical flow data corresponding to the video frame sequence, determining the local optical flow data for the text regions in the video frame sequence based on the optical flow data, and splicing the local optical flow data with the video frame sequence, a sample video frame sequence is obtained; further, based on the sample video frame sequence, an initial text removal model is trained to obtain a text removal model, and the target video frame sequence is processed by the text removal model to obtain a target video frame sequence without text regions. Since the optical flow data can be used to accurately estimate the position change of pixel points in each frame of the video frame sequence, the local optical flow data determined for the text regions in the video frame sequence based on the optical flow data corresponding to the video frame sequence can be used to accurately estimate the position change of pixel points in the text regions of each frame of the video frame sequence. Furthermore, the sample video frame sequence obtained by splicing the local optical flow data with the original video frame sequence is more suitable as the sample data for training the initial text removal model. When the trained text removal model processes the target video frame sequence, it can achieve the effect of clean and non-flashing text removal, effectively alleviating the flashing problem of the removed text between frames in the video, thereby effectively improving the visual effect of video data processing and enhancing the user experience, bringing convenience to the user.
[0049] In an exemplary embodiment, the step of extracting the optical flow data corresponding to the video frame sequence includes:
[0050] Processing the video frame sequence through an optical flow estimation model to obtain the optical flow data between every two adjacent frames in the video frame sequence;
[0051] Taking the optical flow data between every two adjacent frames in the video frame sequence as the optical flow data corresponding to the video frame sequence.
[0052] Among them, the optical flow data between every two adjacent frames is used to reflect the position change of pixel points between every two adjacent frames.
[0053] Specifically, as Figure 4As shown in [description], during the model training phase, assume that the original video frame sequence A consists of 5 consecutive frames. After the terminal obtains the original video frame sequence A input by the user, the terminal can process the video frame sequence A through an optical flow estimation model. For example, the terminal can perform optical flow calculation on the video frame sequence A through the RAFT model, and then obtain the optical flow data between every two adjacent frames in the video frame sequence A. For example, after the terminal performs optical flow calculation on the video frame sequence A through the RAFT model, the optical flow data between every two adjacent frames in the video frame sequence A obtained finally is: the optical flow map of S1 - S2 (representing the optical flow map between the first frame and the second frame obtained by calculation), the optical flow map of S2 - S3, the optical flow map of S3 - S4, the optical flow map of S4 - S5, and the optical flow map of the automatically supplemented frame 0 (S0 - S1). That is, 5 optical flow maps are finally obtained, and the optical flow data (optical flow maps) between every two frames in the calculated video frame sequence A is used as the optical flow data corresponding to the video frame sequence (i.e., 5 optical flow maps). This enables, by using optical flow to track the text area in the video, to automatically and accurately estimate the movement of the text and achieve the effect of clean and non - flickering text removal, that is, it can effectively alleviate the flickering problem of the removed text between frames.
[0054] In an exemplary embodiment, the step of determining local optical flow data for the text area in the video frame sequence based on the optical flow data includes:
[0055] Determine the binary mask data of the text area in the video frame sequence;
[0056] Determine the local optical flow data for the text area in the video frame sequence based on the optical flow data and the binary mask data.
[0057] Among them, the binary mask data in this application may include that the mask value of the target area (text area) is 1 (white), and the mask value of other areas is 0 (black).
[0058] Specifically, during the model training phase, assume that the original video frame sequence A consists of 5 consecutive frames. After the terminal extracts the optical flow data corresponding to the original video frame sequence (i.e., 5 optical flow maps), the terminal can also detect the original video frame sequence through a pre - trained text detection model, and then obtain the text detection boxes of each frame image in the video frame sequence (i.e., the text detection boxes in 5 frame images), and convert the text detection boxes into the binary mask data of the text area of each frame image (i.e., the binary mask data of the text area in 5 frame images); further, as Figure 4As shown, the terminal can determine the local optical flow data for the text regions of each frame in the video frame sequence (i.e., the local optical flow data for the text regions in 5 frames of images) based on the optical flow data (i.e., 5 optical flow maps) and the binary mask data (i.e., the binary mask data for the text regions in 5 frames of images). This enables, by using optical flow to track the text regions in the video, to automatically and accurately estimate the movement of the text and achieve a clean and flicker-free text removal effect, that is, it can effectively alleviate the flicker problem after text removal between frames.
[0059] In one exemplary embodiment, the step of determining the binary mask data for the text regions in the video frame sequence includes:
[0060] Detect the video frame sequence through a pre-trained text detection model to obtain the text detection boxes for each frame image in the video frame sequence;
[0061] Convert the text detection boxes into the binary mask data for the text regions of each frame image, and use the binary mask data for the text regions of each frame image as the binary mask data for the text regions in the video frame sequence.
[0062] Specifically, as Figure 4 shown, in the model training stage, assuming the original video frame sequence A consists of 5 consecutive frames, after the terminal extracts the optical flow data corresponding to the original video frame sequence (i.e., 5 optical flow maps), the terminal can also detect the original video frame sequence through a pre-trained text detection model to obtain the text detection boxes for each frame image in the video frame sequence (i.e., the text detection boxes in 5 frames of images), and convert the text detection boxes into the binary mask data for the text regions of each frame image (i.e., the binary mask data for the text regions in 5 frames of images). For example, the terminal can obtain a preset fill rectangle function and, based on the fill rectangle function, convert the text detection boxes into the binary mask data for the text regions of each frame image (i.e., the binary mask data for the text regions in 5 frames of images). Among them, the preset fill rectangle function includes but is not limited to the fill rectangle function in the opencv library. This enables, by using optical flow to track the text regions in the video, to automatically and accurately estimate the movement of the text and achieve a clean and flicker-free text removal effect, that is, it can effectively alleviate the flicker problem after text removal between frames.
[0063] In one exemplary embodiment, the step of determining the local optical flow data for the text regions in the video frame sequence based on the optical flow data and the binary mask data includes:
[0064] Determine the product between the optical flow data and the binary mask data;
[0065] Use the product as the local optical flow data for the text regions in the video frame sequence.
[0066] Specifically, in the model training stage, assuming that the original video frame sequence A consists of 5 consecutive frames, after the terminal extracts the optical flow data corresponding to the original video frame sequence (i.e., 5 optical flow maps), the terminal can also detect the original video frame sequence through a pre-trained text detection model, and then obtain the text detection boxes of each frame image in the video frame sequence (i.e., the text detection boxes in 5 frames of images), and convert the text detection boxes into binary mask data of the text regions of each frame image (i.e., the binary mask data of the text regions in 5 frames of images); further, as Figure 4 shown, the terminal can calculate the product between the optical flow data (i.e., 5 optical flow maps) and the binary mask data (i.e., the binary mask data of the text regions in 5 frames of images) respectively (i.e., obtain 5 products), and use the calculated products (i.e., the obtained 5 products) as the local optical flow data for the text regions in the video frame sequence. For example, the terminal can calculate the product between the optical flow data (i.e., 5 optical flow maps) and the binary mask data (1) (i.e., the binary mask data of the text regions in 5 frames of images) respectively (i.e., the obtained 5 optical flow maps), and use the calculated products (i.e., the obtained 5 optical flow maps) as the local optical flow data for the text regions in the video frame sequence (5 optical flow maps). This enables, by using optical flow to track the text regions in the video, to automatically and accurately estimate the movement of the text and achieve the effect of clean and non-flashing text removal, that is, it can effectively alleviate the flashing problem after text removal between frames.
[0067] In an exemplary embodiment, the color channels of the local optical flow data include a target color channel, and the color channels of each frame image in the video frame sequence include a first color channel, a second color channel, and a third color channel; the step of splicing the local optical flow data with the video frame sequence to obtain a sample video frame sequence includes:
[0068] Splice the values of the target color channel, the first color channel, the second color channel, and the third color channel according to the color channels to obtain a video frame sequence with the value of the target color channel superimposed;
[0069] Use the video frame sequence with the value of the target color channel superimposed as the sample video frame sequence.
[0070] Specifically, in the model training stage, assume that the original video frame sequence A consists of 5 consecutive frames. After the terminal determines the local optical flow data (i.e., 5 optical flow maps) for the text region in the original video frame sequence based on the optical flow data (i.e., 5 optical flow maps), since the local optical flow data, i.e., each optical flow map, is a single-channel image, that is, each optical flow map only includes one color channel (the target color channel). Assume that each frame image in the original video frame sequence A is an RGB image, that is, the color channels of each frame image in the original video frame sequence A include the first color channel (R), the second color channel (G), and the third color channel (B). Therefore, when the terminal splices the local optical flow data with the video frame sequence, the terminal needs to stack the values (S values) of the target color channel corresponding to each optical flow map with the values (R values), (G values), and (B values) of the first, second, and third color channels of each frame image in the original video frame sequence A according to the color channels (for example: L = S value * a + R value * b + G value * c + B value * d), so as to obtain each frame image (four color channels) with the values (S values) of the target color channel stacked, and use the obtained each frame image (four color channels) with the values (S values) of the target color channel stacked as the sample video frame sequence. Wherein, a, b, c, and d are coefficients, that is, the values of a, b, c, and d can be pre-configured. Thus, by using optical flow to track the text region in the video, the movement of the text can be automatically and accurately estimated, and the effect of clean and non-flashing text removal can be achieved, that is, the flashing problem after text removal between frames can be effectively alleviated.
[0071] In an exemplary embodiment, the step of training the initial text removal model based on the sample video frame sequence to obtain the text removal model includes:
[0072] Using the sample video frame sequence as an input parameter to train the initial text removal model until the training stops when the target loss value determined based on the sample video frame sequence and the text-removed video frame sequence output by the initial text removal model satisfies the preset loss condition, and obtaining the trained text removal model;
[0073] Wherein, the target loss value is determined by the sum of the first loss value and the second loss value. The first loss value is determined based on the difference between the non-text regions of each frame image in the sample video frame sequence and the non-text regions of each frame image in the text-removed video frame sequence, and the second loss value is determined based on the difference between the text regions of every two adjacent frames in the text-removed video frame sequence.
[0074] Specifically, as Figure 4As shown in [description], during the model training phase, the terminal can use the sample video frame sequence (each frame image includes four color channels) as an input parameter to train the initial text removal model until the training stops when the target loss value determined based on the input sample video frame sequence and the video frame sequence with removed text output by the initial text removal model meets the preset loss condition, and then the trained text removal model can be obtained. That is, in the model training phase of this application, a joint training method is adopted. The loss function used during training includes: in order to enable the model to learn the ability to remove text regions without changing non-text regions, a reconstruction loss of non-text regions (i.e., the first loss function) is designed to keep the non-text regions in the input video frame sequence and the output video frame sequence consistent. In addition, in order to reduce the jump between frames in the video frame sequence output by the model, by minimizing the optical flow loss of text regions (i.e., the second loss function), the optical flow of text regions between frames is nearly zero, that is, the text regions can also ensure temporal stability after text removal. Thus, through the injection of optical flow, the flicker and jump generated between frames after text removal are effectively reduced. During the user usage phase, the finally trained model can receive a single frame of picture or a batch of sequential frames of video, and without other additional operations, it can automatically be applied to text removal processing of pictures and videos, thereby effectively improving the automation level and efficiency of video processing.
[0075] In one exemplary embodiment, the steps of determining the target loss value based on the sample video frame sequence and the video frame sequence with removed text output by the initial text removal model include:
[0076] Determine the first loss value based on the non-text regions of each frame image in the sample video frame sequence and the non-text regions of each frame image in the video frame sequence with removed text;
[0077] Determine the second loss value based on the optical flow change between every two adjacent frames in the video frame sequence with removed text;
[0078] Determine the sum value between the first loss value and the second loss value, and use the sum value as the target loss value.
[0079] Among them, the first loss value and the second loss value in this application are only used to distinguish different loss values. For example, the first loss value in this application can also be called the reconstruction loss value of non-text regions, and the second loss value can also be called the optical flow loss value of text regions. That is, the first loss value is used to reflect the difference between the non-text regions in the input video frame sequence and the output video frame sequence, and the second loss value is used to reflect the difference between the text regions of every two adjacent frames in the output video frame sequence, that is, the optical flow difference between the text regions of every two adjacent frames.
[0080] Specifically, as Figure 4As shown in , in the model training stage, the terminal can use the sample video frame sequence (each frame image includes four color channels) as an input parameter to train the initial text removal model until the training stops when the target loss value determined based on the input sample video frame sequence and the video frame sequence with text removed output by the initial text removal model meets the preset loss condition, and then the trained text removal model can be obtained. That is, in the model training stage of this application, a joint training method is adopted. The loss function used in training includes: in order to enable the model to learn the ability to remove text regions without changing non-text regions, a reconstruction loss of non-text regions (i.e., the first loss function) is designed, that is, the first loss value determined based on the non-text regions of each frame image in the sample video frame sequence and the non-text regions of each frame image in the video frame sequence with text removed. Among them, the calculation method of the first loss value L1 (or the first loss function) can be using the following formula (1):
[0081] (1)
[0082] where y i is the non-text region in the i-th input video frame, and f(x i ) is the non-text region in the i-th video frame with text removed output by the model. n is the number of pixels in the picture (video frame). What this formula calculates is the mean absolute error between the input video frame and the output video frame. The smaller the error, the better, indicating that the non-text region is reconstructed better. The value range is between 0 and 1.
[0083] In addition, in order to reduce the jump between frames of the video frame sequence output by the model, by minimizing the optical flow loss of the text region (i.e., the second loss function), that is, determining the second loss value based on the optical flow change between every two adjacent frames in the video frame sequence with text removed, so that the optical flow of the text region between frames is almost zero, that is, the text region can also ensure temporal stability after text removal. Among them, the calculation method of the second loss value L2 (or the second loss function) can be using the following formula (2):
[0084] L2 = sum (calculate 1 optical flow map between every two adjacent frames output by the model) / (width of the optical flow map * height of the optical flow map) (2)
[0085] where sum represents summing up each pixel point of the optical flow map. The closer to 0, the smaller the position change of the pixel points between two adjacent frames, that is, the smaller the jump. Dividing by the denominator in formula (2) means normalizing, so that the value range of the optical flow loss is also between 0 and 1.
[0086] After the terminal determines the first loss value according to the above formula (1) and the second loss value according to the above formula (2), the terminal can determine the target loss value based on the first loss value (the reconstruction loss of the non-text area) and the second loss value (the optical flow loss between every two adjacent frames). Among them, the calculation method of the target loss value L can be to use the following formula (3):
[0087] L = a * L1 + b * L2 (3)
[0088] Among them, L1 represents the first loss value, L2 represents the second loss value, and a and b represent weight coefficients, which can be pre-set weight coefficients. For example, if a = 1 and b = 0.5 are pre-set, then L = 1 * L1 + 0.5 * L2.
[0089] This application also provides an application scenario that applies the above video data processing method. The method provided in the embodiments of this application can be applied to various scenarios of removing text from personalized videos or images. The following takes the scenario where a user interacts with a video processing system as an example to illustrate the video data processing method provided in the embodiments of this application.
[0090] In daily life, text is usually added to video or picture content for information annotation such as subtitles and copyright. However, in some cases, these texts will interfere with the appreciation or analysis of visual content. Users may hope to remove the text or subtitles in the video or picture for different considerations such as aesthetics, copyright regulations, and privacy protection.
[0091] Traditional methods for removing text from videos or images mainly rely on image editing software such as Adobe Photoshop. This usually requires users to manually select the text area in the video and process each frame. The removal operation is very inconvenient, time-consuming, and laborious for users. In addition, there will be more defects after frame-by-frame smearing. For example, obvious flickering is likely to occur between frames after smearing. In addition, there are also some methods that attempt to identify and remove static text in videos through template matching technology. However, these methods work well when the text style and position are relatively fixed, but for the dynamically changing text or background in the video, their effects are limited, that is, the effect of removing text is still not ideal, and obvious flickering problems are still likely to occur between frames. That is, there are few open-source methods for removing text from videos and images currently, especially in dealing with videos. To improve the user experience, this application proposes an automatic text removal method based on optical flow. That is, for the specific Chinese-English subtitle scenario, the text detection model is first fine-tuned. At the same time, considering that optical flow technology can be used to estimate the movement of pixel points in an image sequence, optical flow is used to track the text area in the video to alleviate the removal flickering problem between frames. This enables the user side to only input an image or a video (video frame sequence), and can automatically and accurately estimate the movement of the text, and achieve the effect of clean and flicker-free text removal in the video. Thus, not only can the automation level and efficiency of video processing be improved, but also diverse user needs can be met, providing more flexible and personalized video processing options for users, thereby enhancing the user experience and bringing convenience to users.
[0092] In the technical solution provided by this application, text detection and text removal are connected in series to implement an automated video or image text removal technology, as Figure 4 shown below, and the specific process is as follows:
[0093] (1) Text detection
[0094] Since the current open-source text detection models have limited detection effects on smaller text regions. In addition, the project scenarios of video tasks are mostly in Chinese and English. Although the open-source text detection models support many languages, they do not have particular advantages in the vertical scenario of pure Chinese and English detection. Therefore, in the technical solution provided in this application, the text detection model is fine-tuned. First, 500 real Chinese and English images, most of which are smaller texts, are collected and the corresponding text detection frames are manually labeled. That is, these 500 Chinese and English images are used as sample images, and the labeled text detection frames are used as the labels of each sample image. The labels are used to identify the text regions in each sample image. That is, the terminal can obtain these 500 Chinese and English images (sample image set) and the corresponding manually labeled text detection frames (labels of each sample image), and use these 500 Chinese and English images (sample image set) and the corresponding manually labeled text detection frames (labels of each sample image) as training data to train the initial text detection model to obtain a trained text detection model. That is, the labeled text detection frames in the training data in this application are all smaller texts. Therefore, during the training process, the detection effect of smaller text regions can be improved by increasing the prediction size of the image. The finally obtained text detection model has more accurate detection results in the scenarios of small texts, Chinese and English. That is, when an input image or video is processed by the trained text detection model, the corresponding text detection frame can be obtained. Since the subsequent text removal model requires binary mask data of the text region, the text detection frame needs to be converted into a binary mask. For example, the terminal can use the fill rectangle function in the opencv library to convert the text detection frame into a binary mask.
[0095] (2)Text removal
[0096] In the model training stage of the embodiments of this application, before the data is input into the text removal model, the terminal can perform preprocessing operations of text detection and optical flow calculation (such as using the RAFT model for optical flow calculation) on the data, and then the binary mask of the text and the optical flow can be obtained. Then, the terminal can multiply the two, that is, multiply the binary mask of the text and the optical flow (graph), and the local optical flow for the text region can be obtained. Further, the terminal then splices the local optical flow for the text region and the original input video frame sequence by channels to obtain a spliced video frame sequence, and inputs the spliced video frame sequence into the (initial) text removal model for training. The text removal model here can be based on the general video model 3D-Unet.
[0097] The specific loss functions for training include: To enable the model to learn the ability to remove text regions without changing non-text regions, the reconstruction loss of non-text regions is minimized, and the calculation method is as shown in the aforementioned formula (1), so that the non-text regions in the input and output are kept consistent. In addition, to reduce the jump between frames in the model output, the optical flow loss of text regions is minimized, and the calculation method is as shown in the aforementioned formula (2), so that the optical flow of text regions between frames is almost zero, that is, the temporal stability can be ensured for text regions after removing text.
[0098] Through the injection of optical flow, the final training effectively reduces the flicker and jump generated between frames after removing text. In the user usage stage, the finally trained model can receive a single frame of image or a batch of sequential frames of video, and without other additional operations, it can automatically be applied to text removal processing for images and videos.
[0099] The beneficial effects produced by the technical solution of this application include:
[0100] The technical solution provided by this application can effectively alleviate the problem of video frame flickering after removing text and concatenate text detection. During user testing, just input an image or a video, and it can automatically and accurately estimate the movement of text and achieve the effect of clean and flicker-free text removal. It can not only improve the automation level and efficiency of video processing, but also meet diverse user needs and provide more flexible and personalized video processing options for users.
[0101] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0102] Based on the same inventive concept, the embodiments of this application also provide a video data processing device for implementing the above-mentioned video data processing method. The implementation solutions for solving problems provided by this device are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the video data processing device provided below can refer to the limitations on the video data processing method in the above text, and will not be repeated here.
[0103] In an exemplary embodiment, as Figure 5 shown, a video data processing device is provided, including: an extraction module 502, a determination module 504, a splicing module 506, a training module 508, and a processing module 510, where:
[0104] The extraction module 502 is configured to extract optical flow data corresponding to a video frame sequence.
[0105] The determination module 504 is configured to determine local optical flow data for a text region in the video frame sequence based on the optical flow data.
[0106] The splicing module 506 is configured to splice the local optical flow data with the video frame sequence to obtain a sample video frame sequence.
[0107] The training module 508 is configured to train an initial text removal model based on the sample video frame sequence to obtain a text removal model.
[0108] The processing module 510 is configured to process a target video frame sequence through the text removal model to obtain the target video frame sequence without the text region.
[0109] In an embodiment, the processing module is further configured to process the video frame sequence through an optical flow estimation model to obtain optical flow data between every two adjacent frames in the video frame sequence; and use the optical flow data between every two adjacent frames in the video frame sequence as the optical flow data corresponding to the video frame sequence.
[0110] In an embodiment, the determination module is further configured to determine binary mask data for a text region in the video frame sequence; and determine local optical flow data for the text region in the video frame sequence based on the optical flow data and the binary mask data.
[0111] In an embodiment, the device further includes: a detection module configured to detect the video frame sequence through a pre-trained text detection model to obtain text detection frames for each frame image in the video frame sequence; and a conversion module configured to convert the text detection frames into binary mask data for text regions of each frame image, and use the binary mask data for text regions of each frame image as the binary mask data for the text region in the video frame sequence.
[0112] In an embodiment, the determination module is further configured to determine the product between the optical flow data and the binary mask data; and use the product as the local optical flow data for the text region in the video frame sequence.
[0113] In one embodiment, the color channel of the local optical flow data includes a target color channel, and the color channels of each frame image in the video frame sequence include a first color channel, a second color channel, and a third color channel; the splicing module is further configured to splice the values of the target color channel, the values of the first color channel, the values of the second color channel, and the values of the third color channel according to the color channels, so as to obtain the video frame sequence with the values of the target color channel superimposed; and use the video frame sequence with the values of the target color channel superimposed as the sample video frame sequence.
[0114] In one embodiment, the training module is further configured to use the sample video frame sequence as an input parameter to train the initial text removal model until the training stops when the target loss value determined based on the sample video frame sequence and the video frame sequence with text removed output by the initial text removal model meets a preset loss condition, and obtain the trained text removal model; wherein, the target loss value is determined by the sum of a first loss value and a second loss value, the first loss value is determined based on the difference between the non-text regions of each frame image in the sample video frame sequence and the non-text regions of each frame image in the video frame sequence with text removed, and the second loss value is determined based on the difference between the text regions of every two adjacent frames in the video frame sequence with text removed.
[0115] Each module in the above video data processing device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0116] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a video data processing method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0117] Those skilled in the art can understand that Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0118] In an exemplary embodiment, in an embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0119] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0120] In an embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0121] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0122] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0123] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0124] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. A video data processing method, characterized in that, The method includes: extracting the optical flow data corresponding to the video frame sequence; determining local optical flow data for the text regions in the video frame sequence based on the optical flow data; concatenating the local optical flow data with the video frame sequence to obtain a sample video frame sequence; training an initial text removal model based on the sample video frame sequence to obtain a text removal model; processing a target video frame sequence through the text removal model to obtain the target video frame sequence without the text regions.
2. The method according to claim 1, characterized in that The extracting the optical flow data corresponding to the video frame sequence includes: processing the video frame sequence through an optical flow estimation model to obtain the optical flow data between every two adjacent frames in the video frame sequence; using the optical flow data between every two adjacent frames in the video frame sequence as the optical flow data corresponding to the video frame sequence.
3. The method according to claim 1, wherein The determining local optical flow data for the text regions in the video frame sequence based on the optical flow data includes: determining binary mask data for the text regions in the video frame sequence; determining local optical flow data for the text regions in the video frame sequence based on the optical flow data and the binary mask data.
4. The method according to claim 3, wherein The determining binary mask data for the text regions in the video frame sequence includes: detecting the video frame sequence through a pre-trained text detection model to obtain text detection boxes for each frame image in the video frame sequence; converting the text detection boxes into binary mask data for the text regions of each frame image, and using the binary mask data for the text regions of each frame image as the binary mask data for the text regions in the video frame sequence.
5. The method according to claim 3, characterized in that, The determining local optical flow data for the text regions in the video frame sequence based on the optical flow data and the binary mask data includes: determining the product between the optical flow data and the binary mask data; using the product as the local optical flow data for the text regions in the video frame sequence.
6. The method according to claim 1, characterized in that The color channels of the local optical flow data include a target color channel, and the color channels of each frame image in the video frame sequence include a first color channel, a second color channel, and a third color channel; the concatenating the local optical flow data with the video frame sequence to obtain a sample video frame sequence includes: concatenating the values of the target color channel, the values of the first color channel, the values of the second color channel, and the values of the third color channel according to the color channels to obtain the video frame sequence with the values of the target color channel superimposed; using the video frame sequence with the values of the target color channel superimposed as the sample video frame sequence.
7. The method according to claim 1, wherein The training an initial text removal model based on the sample video frame sequence to obtain a text removal model includes: using the sample video frame sequence as an input parameter to train the initial text removal model until the target loss value determined based on the sample video frame sequence and the video frame sequence with text removed output by the initial text removal model satisfies a preset loss condition, and then stopping the training to obtain the trained text removal model; Among them, the target loss value is determined by the sum of a first loss value and a second loss value. The first loss value is determined based on the difference between the non-text regions of each frame image in the sample video frame sequence and the non-text regions of each frame image in the video frame sequence with text removed. The second loss value is determined based on the difference between the text regions of every two adjacent frames in the video frame sequence with text removed.
8. A video data processing device, characterized in that, The device includes: an extraction module, configured to extract optical flow data corresponding to the video frame sequence; a determination module, configured to determine local optical flow data for the text regions in the video frame sequence based on the optical flow data; a splicing module, configured to splice the local optical flow data with the video frame sequence to obtain a sample video frame sequence; a training module, configured to train an initial text removal model based on the sample video frame sequence to obtain a text removal model; a processing module, configured to process a target video frame sequence through the text removal model to obtain the target video frame sequence without the text regions.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.