Video processing methods, equipment, storage media and software products
By combining optical flow information with adjacent video frames, and integrating optical flow information with a preset video restoration model, the problem of low video erasure quality in existing technologies is solved, achieving efficient target object erasure and restoration.
Patent Information
- Application Number
- CN202411749904.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Existing video processing methods are not high-quality when erasing subtitles, advertisements, people, icons and other content from videos. They are prone to omissions, blurring, or other issues, requiring manual frame-by-frame repair, which is costly, or may result in the video being played directly, affecting the user's viewing experience.
By acquiring the optical flow information of the video to be processed, the location of the target object is determined and segmented. The optical flow information and adjacent video frames are combined for preliminary repair, preserving the original texture and reducing the repair area. A preset video repair model is then used for further repair to improve the erasure quality.
It improves the quality and efficiency of erasing video target objects, reduces the area to be repaired, lowers the cost of manual repair, and enhances the user experience.
Smart Images

Figure CN119603477B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer and network communication technology, and in particular to a video processing method, device, storage medium, and program product. Background Technology
[0002] During video editing and processing, users may need to remove subtitles, advertisements, people, icons, and other content from the video.
[0003] However, some existing video processing methods do not produce high-quality erasure, especially in complex scenes, where they are prone to omissions, scratches, or blurring. These issues require manual frame-by-frame repair, which is costly. Alternatively, playing the video directly without manual repair can negatively impact the user experience. Summary of the Invention
[0004] This disclosure provides a video processing method, apparatus, storage medium, and program product to improve the erasure quality of target objects in a video.
[0005] In a first aspect, embodiments of this disclosure provide a video processing method, including:
[0006] Obtain the optical flow information of the video to be processed;
[0007] The location of the target object in any initial video frame of the video to be processed is determined, and the initial video frame is segmented according to the location of the target object to obtain a first intermediate video frame with the target object removed and an initial mask image; wherein the target object is located in a non-mask area in the initial mask image, which is used to identify the area to be repaired in the first intermediate video frame;
[0008] Based on the optical flow information and the adjacent video frames of the initial video frame, the pixel value of at least one target pixel in the repair area of the first intermediate video frame is determined. Based on the pixel value of the target pixel, the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image are filled to obtain the second intermediate video frame and the intermediate mask image.
[0009] Based on the optical flow information, the second intermediate video frame, and the intermediate mask image, a preset video restoration model is invoked to restore the second intermediate video frame, thereby obtaining the target video frame.
[0010] In a second aspect, embodiments of this disclosure provide a video processing apparatus, including:
[0011] Optical flow acquisition unit, used to acquire optical flow information of the video to be processed;
[0012] The target object recognition unit is used to determine the position of the target object in any initial video frame of the video to be processed, and to segment the initial video frame according to the target object position to obtain a first intermediate video frame with the target object removed and an initial mask image; wherein the target object is located in a non-mask area in the initial mask image, which is used to identify the area to be repaired in the first intermediate video frame;
[0013] The first repair unit is configured to determine the pixel value of at least one target pixel in the repair area of the first intermediate video frame based on the optical flow information and the adjacent video frames of the initial video frame, and fill the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image based on the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image.
[0014] The second repair unit is used to repair the second intermediate video frame by calling a preset video repair model based on the optical flow information, the second intermediate video frame and the intermediate mask image, so as to obtain the target video frame.
[0015] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor and a memory;
[0016] The memory stores computer-executed instructions;
[0017] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the video processing method as described in the first aspect and various possible designs of the first aspect.
[0018] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the video processing method described in the first aspect and various possible designs of the first aspect.
[0019] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the video processing method described in the first aspect and various possible designs of the first aspect.
[0020] The video processing method, device, storage medium, and program product provided in this disclosure acquire optical flow information of the video to be processed; determine the position of a target object in any initial video frame of the video to be processed; segment the initial video frame according to the target object position to obtain a first intermediate video frame with the target object removed and an initial mask image; wherein the target object is located in a non-mask area in the initial mask image, which is used to identify the area to be repaired in the first intermediate video frame; determine the pixel value of at least one target pixel in the area to be repaired in the first intermediate video frame according to the optical flow information and adjacent video frames of the initial video frame; fill the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image according to the pixel value of the target pixel to obtain a second intermediate video frame and an intermediate mask image; and call a preset video repair model to repair the second intermediate video frame according to the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame. In this disclosure, preliminary repair by combining optical flow information and adjacent video frames preserves richer original textures and reduces the area to be repaired. Further repair using a video repair model improves repair efficiency and the quality of target object erasure. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A schematic diagram of a scene for a video processing method provided in an embodiment of this disclosure;
[0023] Figure 2 This is a schematic flowchart of a video processing method provided in an embodiment of the present disclosure;
[0024] Figure 3 A schematic flowchart illustrating a video processing method provided in another embodiment of this disclosure;
[0025] Figure 4 A schematic flowchart illustrating a video processing method provided in another embodiment of this disclosure;
[0026] Figure 5 A schematic flowchart illustrating a video processing method provided in another embodiment of this disclosure;
[0027] Figure 6 This is a structural block diagram of a video processing device provided in one embodiment of the present disclosure;
[0028] Figure 7This is a schematic diagram of the hardware structure of a video processing device provided in an embodiment of the present disclosure. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0030] During video editing and processing, users often need to erase subtitles, advertisements, people, icons, and other content from videos. However, some existing video processing methods do not produce high-quality erasures, especially in complex scenes, where they are prone to omissions, blurring, or other issues. These require manual frame-by-frame repair, which is costly. Alternatively, playing videos directly without manual repair can negatively impact the user experience.
[0031] To address the aforementioned technical problems, this disclosure provides a video processing method. The method involves acquiring optical flow information of the video to be processed; determining the location of a target object in any initial video frame of the video to be processed; segmenting the initial video frame according to the target object location to obtain a first intermediate video frame with the target object removed and an initial mask image; wherein the target object is located in a non-masked area in the initial mask image, used to identify the area to be repaired in the first intermediate video frame; determining the pixel value of at least one target pixel within the area to be repaired in the first intermediate video frame based on the optical flow information and adjacent video frames of the initial video frame; filling the target pixel in the first intermediate video frame and the corresponding pixel in the initial mask image according to the pixel value of the target pixel to obtain a second intermediate video frame and an intermediate mask image; and using a preset video repair model to repair the second intermediate video frame based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame. In this embodiment, preliminary repair using optical flow information and adjacent video frames preserves richer original textures and reduces the area requiring repair. Further repair using a video repair model improves repair efficiency and the quality of target object erasure.
[0032] The application scenarios of the video processing method in this disclosure embodiment are as follows: Figure 1As shown, the method of this embodiment can be applied to electronic devices such as terminal devices or servers to obtain optical flow information of the video to be processed; determine the position of the target object in any initial video frame of the video to be processed; segment the initial video frame according to the target object position to obtain a first intermediate video frame with the target object removed and an initial mask image; wherein the target object is located in a non-mask area in the initial mask image, which is used to identify the repair area of the first intermediate video frame; determine the pixel value of at least one target pixel in the repair area of the first intermediate video frame according to the optical flow information and the adjacent video frames of the initial video frame; fill the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image according to the pixel value of the target pixel to obtain a second intermediate video frame and an intermediate mask image; call a preset video repair model to repair the second intermediate video frame according to the optical flow information, the second intermediate video frame and the intermediate mask image to obtain the target video frame, and finally output the target video with the target object erased.
[0033] The video processing method of this disclosure will be described in detail below with reference to specific embodiments.
[0034] refer to Figure 2 , Figure 2 This is a schematic flowchart of a video processing method according to an embodiment of the present disclosure. The method of this embodiment can be applied to electronic devices such as terminal devices or servers. The video processing method includes:
[0035] S201. Obtain the optical flow information of the video to be processed.
[0036] In this embodiment, to perform fine-grained erasure of target objects in the video to be processed, considering that the part of the video frame occluded by the target object may not be occluded in adjacent video frames, the original texture in the adjacent video frames can be used to complete the part of the video frame occluded by the target object. However, since objects in the video may move, the part of the video frame occluded by the target object may not be in the same position in adjacent video frames. Therefore, it is also necessary to determine the corresponding position of the part of the video frame occluded by the target object in adjacent video frames. Considering that optical flow is a two-dimensional vector field describing the trajectory and velocity of pixels in a continuous frame image over time, used to capture the motion of objects and displacement information between adjacent frames, this embodiment can use optical flow information to determine the corresponding position of the part of the video frame occluded by the target object in adjacent video frames. In this embodiment, optical flow information can be obtained from the video to be processed, and any known method can be used to obtain the optical flow information, such as sparse optical flow algorithms, dense optical flow algorithms, deep learning-based optical flow algorithms, etc.
[0037] Optionally, in this embodiment, a preset optical flow model can be used to obtain optical flow information from the video to be processed, wherein the preset optical flow model can be any known optical flow algorithm model.
[0038] Furthermore, considering that the acquired optical flow information may be inaccurate or missing, optical flow completion can be performed on the initial optical flow information based on the acquired initial optical flow information. The optical flow completion method can be any known method, such as optical flow completion methods based on forward flow, reverse flow, and non-adjacent frames, or optical flow completion methods based on deep learning, etc.
[0039] Optionally, in this embodiment, a preset optical flow completion model can be called to complete the initial optical flow information based on the video to be processed and the initial optical flow information to obtain the final optical flow information. The preset optical flow completion model can be any known optical flow completion model.
[0040] S202. Determine the position of the target object in any initial video frame of the video to be processed, and perform target object segmentation on the initial video frame according to the target object position to obtain a first intermediate video frame with the target object removed and an initial mask image; wherein the target object is located in a non-masked area in the initial mask image, which is used to identify the area to be repaired in the first intermediate video frame.
[0041] In this embodiment, in order to perform fine-grained erasure of target objects in the video to be processed, it is also necessary to determine the position of the target object in each initial video frame of the video to be processed. Any recognition method can be used to identify the target object from the initial video frames to determine the position of the target object. Optionally, a recognition model can be used to identify the target object from the initial video frames. The recognition model can be any machine learning model.
[0042] Furthermore, the initial video frame can be segmented based on the target object's location, and the target object can be extracted from the initial video frame to obtain a video frame with the target object removed, referred to here as the first intermediate video frame. In addition, a mask image of the target object can be obtained, referred to as the initial mask image. The mask image is a binary image with the same size as the original image, in which the selected area (non-masked area or unmasked area) is marked as 1 (or True), while the remaining areas (masked area or masked area) are marked as 0 (or False), used to indicate the specific area to be processed when applying certain image processing operations. In this embodiment, the target object's area is located in the non-masked area (selected area or unmasked area) of the initial mask image, indicating that this area is the area to be repaired. That is, the non-target object areas in the initial mask image do not need to be repaired and are masked, while the target object area needs to be repaired and is not masked.
[0043] Optionally, when determining the location of a target object in any initial video frame, the location of the target object to be erased can also be determined according to user needs. Users can pre-set preset filtering conditions for target objects, such as the type, color, and location of the target object. That is, only target objects that meet the preset filtering conditions need to be erased. Therefore, when determining the location of a target object in any initial video frame, all target objects in any initial video frame of the video to be processed can be identified first. Then, based on the preset filtering conditions, target objects that meet the preset filtering conditions can be filtered from all target objects in the initial video frame, and the location of the target objects that meet the preset filtering conditions can be determined. Subsequent steps are then performed based on the location of the target object, enabling the erasure of only target objects that meet the preset filtering conditions.
[0044] It should be noted that in this embodiment, the execution order of S201 and S202 is not limited, and they can be executed in parallel.
[0045] S203. Based on the optical flow information and the adjacent video frames of the initial video frame, determine the pixel value of at least one target pixel in the repair area of the first intermediate video frame, and fill the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image according to the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image.
[0046] In this embodiment, based on optical flow information and adjacent video frames of the initial video frame, the corresponding pixel of at least one target pixel in the repair area of the first intermediate video frame can be determined in the adjacent video frames. Then, the pixel value of the target pixel is determined based on the corresponding pixel in the adjacent video frames. Based on the pixel value of the target pixel, the corresponding pixel of the target pixel in the first intermediate video frame is filled in the second intermediate video frame (the non-masked area is changed into a masked area, in other words, the corresponding pixel is changed from 1 to 0). This indicates that the position of the corresponding pixel in the initial mask image has been repaired and no further repair is needed, thus obtaining the intermediate mask image.
[0047] In practice, the pixel value of the target pixel can be filled into the target pixel of the first intermediate video frame to obtain the second intermediate video frame. In addition, the corresponding pixel of the target pixel in the corresponding initial mask image is determined, and the corresponding pixel in the initial mask image is masked to change the non-masked area into a masked area, and finally the intermediate mask image is obtained.
[0048] Optionally, in this embodiment, when determining the pixel value of at least one target pixel in the repair area of the first intermediate video frame, the corresponding pixel of at least one target pixel in the repair area of the first intermediate video frame in the adjacent video frame can be determined according to the optical flow information. Then, the pixel value of the target pixel can be determined according to the corresponding pixel in the adjacent video frame, so that the original texture of the target pixel can be preserved as much as possible based on the corresponding pixel in the adjacent video frame.
[0049] For pixels in the area to be repaired that do not have a corresponding pixel in an adjacent video frame, no action needs to be taken in this step.
[0050] Of course, the above method for determining the pixel value of the target pixel can be achieved by any other feasible method, such as any feasible algorithm or machine learning model.
[0051] S204. Based on the optical flow information, the second intermediate video frame, and the intermediate mask image, a preset video repair model is invoked to repair the second intermediate video frame to obtain the target video frame.
[0052] In this embodiment, since the second intermediate video frame has undergone preliminary repair, but may not have repaired all pixels in the area to be repaired, a preset video repair model can be invoked based on the optical flow information, the second intermediate video frame, and the intermediate mask image to further repair the second intermediate video frame, so that the second intermediate video frame is completely repaired, and finally the target video frame is obtained, which can then be output as the final target video. The preset video repair model can be any machine learning model, such as the Transformer model, which is pre-trained to have the ability to perform video repair based on optical flow information, video frames, and mask images.
[0053] It should be noted that the target object in this embodiment can be any object in the video to be processed that needs to be erased, such as subtitles, advertising labels, people, icons, etc. For such objects as subtitles and text-based advertisements, since they are generally small and have complex strokes, the shape of the occlusion area is also relatively complex, thus requiring fine-grained erasure, and the video processing method described above can be used. For other target objects that do not require fine-grained erasure, although the video processing method described above can also be used, the cost is relatively high. Therefore, other erasure methods can also be used. For example, the repair can be performed based solely on optical flow information and the first intermediate video frame, calling the aforementioned preset video repair model or other video repair models.
[0054] Optionally, in this embodiment, the video processing method of this embodiment can be executed on all video frames of the video to be processed, or the video processing method of this embodiment can be executed by extracting frames from the video to be processed.
[0055] Furthermore, after erasing the target object, other video editing operations can be performed, such as adding new subtitles (e.g., subtitles in different colors and languages), adding new advertisements, icons, etc.
[0056] The video processing method provided in this embodiment obtains optical flow information of the video to be processed; determines the position of the target object in any initial video frame of the video to be processed; segments the initial video frame according to the target object position to obtain a first intermediate video frame with the target object removed and an initial mask image; wherein the target object is located in a non-mask area in the initial mask image, which is used to identify the area to be repaired in the first intermediate video frame; determines the pixel value of at least one target pixel in the area to be repaired in the first intermediate video frame according to the optical flow information and adjacent video frames of the initial video frame; fills the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image according to the pixel value of the target pixel to obtain a second intermediate video frame and an intermediate mask image; and calls a preset video repair model to repair the second intermediate video frame according to the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame. In this embodiment, preliminary repair by combining optical flow information and adjacent video frames preserves richer original textures and reduces the area to be repaired. Then, the video repair model is used for repair, which improves the repair efficiency and the erasure quality of the target object.
[0057] In any of the above embodiments, when determining the pixel value of the target pixel based on the corresponding pixels in the adjacent video frames, one of the following methods may be used:
[0058] Method 1: Determine the pixel value that appears most frequently among the corresponding pixel values in the adjacent video frames, and determine the pixel value that appears most frequently as the pixel value of the target pixel.
[0059] In this embodiment, since there may be multiple adjacent video frames, such as several consecutive frames, the pixel value of the target pixel in each adjacent video frame may be exactly the same or there may be some differences. In this embodiment, the pixel value that appears most frequently can be determined from the pixel value of the corresponding pixel in the adjacent video frames. The pixel value that appears most frequently is more accurate. Therefore, the pixel value that appears most frequently can be determined as the pixel value of the target pixel.
[0060] Method 2: Calculate a weighted average of the pixel values of corresponding pixels in the adjacent video frames, and determine the weighted average result as the pixel value of the target pixel.
[0061] In this embodiment, a weighted average can also be performed on the pixel values of corresponding pixels in adjacent video frames. By comprehensively considering the pixel values of corresponding pixels in multiple adjacent video frames, the weighted average result is determined as the pixel value of the target pixel. The weights of the weighted average can be set according to the actual situation, or they can be determined through pre-training.
[0062] Of course, the pixel value of the target pixel can also be determined by other methods based on the corresponding pixels in adjacent video frames, which will not be elaborated here.
[0063] Based on any of the above embodiments, since the pixel value of the target pixel is inferred, it may not be accurate. Therefore, before filling the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image according to the pixel value of the target pixel, the confidence level of the pixel value of the target pixel can also be considered. Only when the confidence level is high, for example, higher than a preset confidence threshold, is the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image filled according to the pixel value of the target pixel. Otherwise, if the confidence level is not high, the pixel value of the target pixel is ignored.
[0064] The confidence level of the target pixel's pixel value can be determined in any feasible way, such as through a pre-trained machine learning model. The video to be processed, the location of the target pixel, and the pixel value can be input into the model, and the model outputs the confidence level of the target pixel's pixel value. The specific model is not limited in this embodiment.
[0065] Based on any of the above embodiments, S204 calls a preset video restoration model to restore the second intermediate video frame according to the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame, which may specifically include:
[0066] Based on the optical flow information, the second intermediate video frame, and the intermediate mask image, the preset video repair model is invoked to repair the area to be repaired in the second intermediate video frame to obtain the target video frame, wherein the area to be repaired in the second intermediate video frame is the area in the second intermediate video frame that corresponds to the non-masked area of the intermediate mask image.
[0067] In this embodiment, in order to reduce the computational load of the preset video restoration model and improve the restoration efficiency, the preset video restoration model can be made to only repair the area to be repaired in the second intermediate video frame. The area to be repaired in the second intermediate video frame is the area in the second intermediate video frame that corresponds to the non-masked area of the intermediate mask image, that is, the remaining non-masked area after removing the pixels filled by the initial restoration from the non-masked area of the initial mask image. This greatly reduces the area of the area to be repaired, thereby reducing the computational load of the preset video restoration model.
[0068] Based on any of the above embodiments, feature extraction can be performed on the area to be repaired in the current second intermediate video frame within a certain spatiotemporal range when repairing it. Local features can be extracted by combining images of the same window area of adjacent second intermediate video frames. Of course, global features can also be extracted by combining the overall image of the current second intermediate video frame to improve the accuracy and completeness of the features. Then, the area to be repaired in the current second intermediate video frame can be repaired based on the features.
[0069] In one optional embodiment, the preset video restoration model is a Transformer model, a neural network model based on a self-attention mechanism. In this embodiment, the self-attention mechanism can simultaneously process images of target regions from multiple second intermediate video frames, enabling better spatiotemporal prediction learning, capturing long-distance dependencies, and restoring the region to be restored in the current second intermediate video frame. Of course, other models can also be used. The preset video restoration model includes an encoder and a decoder, where the encoder extracts features and the decoder generates the target video frame. The following description uses the Transformer model; the processing of other models can be adaptively adjusted according to the specific model. Specifically, for example... Figure 3 As shown, the step of invoking the preset video insulation model to repair the area to be repaired in the second intermediate video frame based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame includes:
[0070] S301. Obtain an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image;
[0071] S302, Input the optical flow information, the image of the target region in the current second intermediate video frame, and the image of the target region in the adjacent second intermediate video frames into the encoder of the preset video restoration model to extract feature maps;
[0072] S303. Input the feature map into the decoder of the preset video restoration model to generate the target video frame.
[0073] In this embodiment, to improve the processing efficiency of the Transformer model and reduce the amount of data, this embodiment can focus only on the image of the neighborhood range of the region to be repaired (a preset range around the region to be repaired) in the second intermediate video frame, that is, the image of the region corresponding to the preset neighborhood of the non-masked region of the intermediate mask image in the second intermediate video frame, without using the entire second intermediate video frame as the input of the Transformer model. At the same time, it is necessary to refer to the images of the same region in adjacent second intermediate video frames to form a set of temporal images. Then, the optical flow information, the image of the target region of the current second intermediate video frame, and the images of the target regions of adjacent second intermediate video frames are input into the encoder of the Transformer model to extract feature maps. The encoder of the Transformer model is used to extract features from the input data and output feature maps (vectors), which contain the dependencies and contextual information of the input sequence. Specifically, it may include a multi-head self-attention layer to capture the global dependencies between image patches, and a feed-forward neural network. The network is used to perform further nonlinear transformations on the features of each image patch to enhance the expressive power of the model. The residual connection and layer normalization (Add & Norm) layers are used to alleviate the gradient vanishing problem through residual connections and to make the input distribution of each layer more stable through layer normalization. The specific processing process is not described in detail here.
[0074] More specifically, before inputting into the encoder of the Transformer model, the optical flow information, the image of the target region of the current second intermediate video frame, and the image of the target region of the adjacent second intermediate video frames can be embedded to generate corresponding vector representations, and then the corresponding vector representations are input into the encoder of the Transformer model.
[0075] Furthermore, the feature map output by the encoder is input into the decoder of the Transformer model for prediction, thereby generating the repaired target video frame. The decoder of the Transformer model is used to predict and generate sequences based on the encoder output sequence. Specifically, it may include a masked multi-head self-attention layer, which is used to calculate the attention score between each token and other tokens in the target sequence; a multi-head self-attention mechanism layer, which is used to align the current state of the decoder with the output of the encoder, allowing the decoder to refer to the information of the input sequence when generating the target sequence; a feed-forward network, which is used to further process the output of the attention layer through a series of nonlinear transformations; and residual connections and layer normalization (Add & Norm) layers, which are used to alleviate the gradient vanishing problem through residual connections and to make the input distribution of each layer more stable through layer normalization, etc. The specific processing procedures are not elaborated here.
[0076] In this embodiment, only the optical flow information, the image of the target region of the current second intermediate video frame, and the images of the target regions of adjacent second intermediate video frames are used as input. Global features are not considered. However, when global features are considered, there are two cases: one is to reduce the size of the current second intermediate video frame to the size of the target region and use it as input to the encoder to achieve image-level fusion; the other is to extract features from the current second intermediate video frame to obtain a global feature map, and then perform feature-level fusion with the feature map output by the encoder. These are described in detail below:
[0077] In one alternative embodiment, such as Figure 4 As shown, the step involves using the optical flow information, the second intermediate video frame, and the intermediate mask image to call the preset video restoration model to restore the region to be restored in the second intermediate video frame, thereby obtaining the target video frame, including...
[0078] S401. Obtain an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image.
[0079] S402. Reduce the current second intermediate video frame to the size of the target area to obtain the reduced current second intermediate video frame;
[0080] S403. Input the optical flow information, the scaled-down current second intermediate video frame, the image of the target region of the current second intermediate video frame, and the image of the target region of the adjacent second intermediate video frames into the encoder of the preset video restoration model to extract feature maps.
[0081] S404. Input the feature map into the decoder of the preset video restoration model to generate the target video frame.
[0082] In this embodiment, the image of the target region within the current second intermediate video frame and the images of the target regions within a preset number of adjacent second intermediate video frames are obtained as in embodiment S301 above. In addition, the current second intermediate video frame is reduced (resized) to the size of the target region. Then, the optical flow information, the reduced current second intermediate video frame, the image of the target region of the current second intermediate video frame, and the images of the target regions of adjacent second intermediate video frames are input into the encoder of the Transformer model to extract feature maps. The obtained feature maps are feature maps that fuse global features. Then, the feature maps are input into the decoder of the Transformer model to generate the repaired target video frame, as in embodiment S303 above.
[0083] In another alternative embodiment, such as Figure 5 As shown, the step involves using the optical flow information, the second intermediate video frame, and the intermediate mask image to call the preset video restoration model to restore the region to be restored in the second intermediate video frame, thereby obtaining the target video frame, including...
[0084] S501. Obtain an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image.
[0085] S502, Input the optical flow information, the image of the target region in the current second intermediate video frame and the image of the target region in the adjacent second intermediate video frames into the encoder of the preset video restoration model to extract the first feature map;
[0086] S503. Extract features from the current second intermediate video frame to obtain the second feature map;
[0087] S504. The first feature map and the second feature map are fused to obtain a fused feature map;
[0088] S505. Input the fused feature map into the decoder of the preset video restoration model to generate the target video frame.
[0089] In this embodiment, steps S501-S502 are the same as steps S301-S302 in the above embodiment, that is, the optical flow information, the image of the target region of the current second intermediate video frame, and the images of the target regions of adjacent second intermediate video frames are input into the encoder of the Transformer model to extract a first feature map, which is a local feature map; while step S503 extracts features from the current second intermediate video frame separately to obtain a second feature map, which is a global feature map. The feature extraction of the current second intermediate video frame can be performed using a separate feature extraction layer, or the encoder of the Transformer model can be used. Furthermore, the first feature map and the second feature map can be fused to achieve the fusion of local and global features to obtain a fused feature map, which is then input into the decoder of the Transformer model to generate the repaired target video frame, as in step S303 of the above embodiment.
[0090] The video processing methods described in the above embodiments can achieve fine-grained erasure of the target object in the video to be processed. In the initial repair process, optical flow information is combined to preserve as much of the original texture as possible, which also reduces the area that needs to be repaired. Then, a video repair model is used for repair, which improves the repair efficiency and the erasure quality of the target object.
[0091] It should be noted that the above Figures 3-5 The method shown is based on a video restoration model that is not limited to the Transformer model. The processing principle of other models is similar and can be adjusted according to the specific model.
[0092] Corresponding to the video processing method in the above embodiments, Figure 6 This is a structural block diagram of a video processing apparatus provided in an embodiment of this disclosure. For ease of explanation, only the parts relevant to the embodiments of this disclosure are shown. (Refer to...) Figure 6 The video processing device 600 includes: an optical flow acquisition unit 601, a target object recognition unit 602, a first repair unit 603, and a second repair unit 604.
[0093] Among them, the optical flow acquisition unit 601 is used to acquire the optical flow information of the video to be processed;
[0094] The target object recognition unit 602 is used to determine the position of the target object in any initial video frame of the video to be processed, and to segment the initial video frame according to the target object position to obtain a first intermediate video frame with the target object removed and an initial mask image; wherein the target object is located in a non-mask area in the initial mask image, which is used to identify the area to be repaired in the first intermediate video frame.
[0095] The first repair unit 603 is used to determine the pixel value of at least one target pixel in the repair area of the first intermediate video frame according to the optical flow information and the adjacent video frames of the initial video frame, and fill the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image according to the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image.
[0096] The second repair unit 604 is used to repair the second intermediate video frame by calling a preset video repair model based on the optical flow information, the second intermediate video frame and the intermediate mask image, so as to obtain the target video frame.
[0097] In one or more embodiments of this disclosure, when the first repair unit 603 fills the corresponding pixels of the target pixel in the first intermediate video frame and the corresponding pixels of the initial mask image according to the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image, it is used to:
[0098] The pixel value of the target pixel is filled into the target pixel of the first intermediate video frame to obtain the second intermediate video frame;
[0099] The corresponding pixel of the target pixel in the initial mask image is determined, and the corresponding pixel in the initial mask image is masked to obtain the intermediate mask image.
[0100] In one or more embodiments of this disclosure, when the first repair unit 603 determines the pixel value of at least one target pixel within the repair area of the first intermediate video frame based on the optical flow information and adjacent video frames of the initial video frame, it is configured to:
[0101] Based on the optical flow information, determine the corresponding pixel of at least one target pixel in the repair area of the first intermediate video frame in the adjacent video frame, and determine the pixel value of the target pixel based on the corresponding pixel in the adjacent video frame.
[0102] In one or more embodiments of this disclosure, when the first repair unit 603 determines the pixel value of the target pixel based on corresponding pixels in the adjacent video frames, it is configured to:
[0103] Determine the pixel value that appears most frequently among the corresponding pixel values in the adjacent video frames, and set the pixel value that appears most frequently as the pixel value of the target pixel; or
[0104] The pixel values of corresponding pixels in the adjacent video frames are weighted and averaged, and the weighted average result is determined as the pixel value of the target pixel.
[0105] In one or more embodiments of this disclosure, when the first repair unit 603 fills the corresponding pixels of the target pixel in the first intermediate video frame and the corresponding pixels of the initial mask image according to the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image, it is used to:
[0106] Determine the confidence level of the pixel value of the target pixel. If the confidence level is higher than a preset confidence threshold, then fill the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image according to the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image.
[0107] In one or more embodiments of this disclosure, when the second repair unit 604 repairs the second intermediate video frame by calling a preset video repair model based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame, it is used to:
[0108] Based on the optical flow information, the second intermediate video frame, and the intermediate mask image, the preset video repair model is invoked to repair the area to be repaired in the second intermediate video frame to obtain the target video frame, wherein the area to be repaired in the second intermediate video frame is the area in the second intermediate video frame that corresponds to the non-masked area of the intermediate mask image.
[0109] In one or more embodiments of this disclosure, when the second repair unit 604 repairs the area to be repaired in the second intermediate video frame by calling the preset video repair model based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame, it is used to:
[0110] Acquire an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image;
[0111] The optical flow information, the image of the target region in the current second intermediate video frame, and the image of the target region in the adjacent second intermediate video frames are input into the encoder of the preset video restoration model to extract feature maps.
[0112] The feature map is input into the decoder of the preset video restoration model to generate the target video frame.
[0113] In one or more embodiments of this disclosure, when the second repair unit 604 repairs the area to be repaired in the second intermediate video frame by calling the preset video repair model based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame, it is used to:
[0114] Acquire an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image;
[0115] The current second intermediate video frame is reduced to the size of the target area to obtain the reduced current second intermediate video frame;
[0116] The optical flow information, the scaled-down current second intermediate video frame, the image of the target region of the current second intermediate video frame, and the image of the target region of the adjacent second intermediate video frames are input into the encoder of the preset video restoration model to extract feature maps.
[0117] The feature map is input into the decoder of the preset video restoration model to generate the target video frame.
[0118] In one or more embodiments of this disclosure, when the second repair unit 604 repairs the area to be repaired in the second intermediate video frame by calling the preset video repair model based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame, it is used to:
[0119] Acquire an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image;
[0120] The optical flow information, the image of the target region in the current second intermediate video frame, and the image of the target region in the adjacent second intermediate video frames are input into the encoder of the preset video restoration model to extract the first feature map;
[0121] Extract features from the current second intermediate video frame to obtain the second feature map;
[0122] The first feature map and the second feature map are fused to obtain a fused feature map;
[0123] The fused feature map is input into the decoder of the preset video restoration model to generate the target video frame.
[0124] In one or more embodiments of this disclosure, the optical flow acquisition unit 601, when acquiring optical flow information of the video to be processed, is used to:
[0125] Based on the video to be processed, the initial optical flow information is obtained by calling the preset optical flow model;
[0126] Based on the video to be processed and the initial optical flow information, a preset optical flow completion model is invoked to complete the initial optical flow information, thereby obtaining the optical flow information of the video to be processed.
[0127] In one or more embodiments of this disclosure, the target object recognition unit 602, when determining the position of a target object in any initial video frame of the video to be processed, is configured to:
[0128] Identify all target objects in any initial video frame of the video to be processed;
[0129] Select target objects that meet preset filtering conditions from all target objects in the initial video frame, and determine the location of the target objects that meet the preset filtering conditions.
[0130] In one or more embodiments of this disclosure, the target object is a caption.
[0131] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.
[0132] To implement the above embodiments, this disclosure also provides an electronic device.
[0133] refer to Figure 7 The diagram illustrates a structural schematic of an electronic device 700 suitable for implementing embodiments of the present disclosure. The electronic device 700 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0134] like Figure 7As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0135] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0136] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.
[0137] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0138] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0139] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0140] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0142] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0143] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0144] In a first aspect, according to one or more embodiments of this disclosure, a video processing method is provided, comprising:
[0145] Obtain the optical flow information of the video to be processed;
[0146] The location of the target object in any initial video frame of the video to be processed is determined, and the initial video frame is segmented according to the location of the target object to obtain a first intermediate video frame with the target object removed and an initial mask image; wherein the target object is located in a non-mask area in the initial mask image, which is used to identify the area to be repaired in the first intermediate video frame;
[0147] Based on the optical flow information and the adjacent video frames of the initial video frame, the pixel value of at least one target pixel in the repair area of the first intermediate video frame is determined. Based on the pixel value of the target pixel, the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image are filled to obtain the second intermediate video frame and the intermediate mask image.
[0148] Based on the optical flow information, the second intermediate video frame, and the intermediate mask image, a preset video restoration model is invoked to restore the second intermediate video frame, thereby obtaining the target video frame.
[0149] According to one or more embodiments of this disclosure, the step of filling the target pixel of the first intermediate video frame and the corresponding pixels of the initial mask image with the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image includes:
[0150] The pixel value of the target pixel is filled into the target pixel of the first intermediate video frame to obtain the second intermediate video frame;
[0151] The corresponding pixel of the target pixel in the initial mask image is determined, and the corresponding pixel in the initial mask image is masked to obtain the intermediate mask image.
[0152] According to one or more embodiments of this disclosure, determining the pixel value of at least one target pixel within the repair area of the first intermediate video frame based on the optical flow information and adjacent video frames of the initial video frame includes:
[0153] Based on the optical flow information, determine the corresponding pixel of at least one target pixel in the repair area of the first intermediate video frame in the adjacent video frame, and determine the pixel value of the target pixel based on the corresponding pixel in the adjacent video frame.
[0154] According to one or more embodiments of this disclosure, determining the pixel value of the target pixel based on corresponding pixels in the adjacent video frames includes:
[0155] Determine the pixel value that appears most frequently among the corresponding pixel values in the adjacent video frames, and set the pixel value that appears most frequently as the pixel value of the target pixel; or
[0156] The pixel values of corresponding pixels in the adjacent video frames are weighted and averaged, and the weighted average result is determined as the pixel value of the target pixel.
[0157] According to one or more embodiments of this disclosure, the step of filling the target pixel of the first intermediate video frame and the corresponding pixels of the initial mask image with the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image includes:
[0158] Determine the confidence level of the pixel value of the target pixel. If the confidence level is higher than a preset confidence threshold, then fill the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image according to the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image.
[0159] According to one or more embodiments of this disclosure, the step of invoking a preset video restoration model to restore the second intermediate video frame based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame includes:
[0160] Based on the optical flow information, the second intermediate video frame, and the intermediate mask image, the preset video repair model is invoked to repair the area to be repaired in the second intermediate video frame to obtain the target video frame, wherein the area to be repaired in the second intermediate video frame is the area in the second intermediate video frame that corresponds to the non-masked area of the intermediate mask image.
[0161] According to one or more embodiments of this disclosure, the step of invoking the preset video insulation model to repair the region to be repaired in the second intermediate video frame based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame includes:
[0162] Acquire an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image;
[0163] The optical flow information, the image of the target region in the current second intermediate video frame, and the image of the target region in the adjacent second intermediate video frames are input into the encoder of the preset video restoration model to extract feature maps.
[0164] The feature map is input into the decoder of the preset video restoration model to generate the target video frame.
[0165] According to one or more embodiments of this disclosure, the step of invoking the preset video insulation model to repair the region to be repaired in the second intermediate video frame based on the optical flow information, the second intermediate video frame, and the intermediate mask image, to obtain the target video frame, includes...
[0166] Acquire an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image;
[0167] The current second intermediate video frame is reduced to the size of the target area to obtain the reduced current second intermediate video frame;
[0168] The optical flow information, the scaled-down current second intermediate video frame, the image of the target region of the current second intermediate video frame, and the image of the target region of the adjacent second intermediate video frames are input into the encoder of the preset video restoration model to extract feature maps.
[0169] The feature map is input into the decoder of the preset video restoration model to generate the target video frame.
[0170] According to one or more embodiments of this disclosure, the step of invoking the preset video insulation model to repair the region to be repaired in the second intermediate video frame based on the optical flow information, the second intermediate video frame, and the intermediate mask image, to obtain the target video frame, includes...
[0171] Acquire an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image;
[0172] The optical flow information, the image of the target region in the current second intermediate video frame, and the image of the target region in the adjacent second intermediate video frames are input into the encoder of the preset video restoration model to extract the first feature map;
[0173] Extract features from the current second intermediate video frame to obtain the second feature map;
[0174] The first feature map and the second feature map are fused to obtain a fused feature map;
[0175] The fused feature map is input into the decoder of the preset video restoration model to generate the target video frame.
[0176] According to one or more embodiments of this disclosure, obtaining the optical flow information of the video to be processed includes:
[0177] Based on the video to be processed, the initial optical flow information is obtained by calling the preset optical flow model;
[0178] Based on the video to be processed and the initial optical flow information, a preset optical flow completion model is invoked to complete the initial optical flow information, thereby obtaining the optical flow information of the video to be processed.
[0179] According to one or more embodiments of this disclosure, determining the location of the target object in any initial video frame of the video to be processed includes:
[0180] Identify all target objects in any initial video frame of the video to be processed;
[0181] Select target objects that meet preset filtering conditions from all target objects in the initial video frame, and determine the location of the target objects that meet the preset filtering conditions.
[0182] According to one or more embodiments of this disclosure, the target object is a subtitle.
[0183] Secondly, according to one or more embodiments of this disclosure, a video processing apparatus is provided, comprising:
[0184] Optical flow acquisition unit, used to acquire optical flow information of the video to be processed;
[0185] The target object recognition unit is used to determine the position of the target object in any initial video frame of the video to be processed, and to segment the initial video frame according to the target object position to obtain a first intermediate video frame with the target object removed and an initial mask image; wherein the target object is located in a non-mask area in the initial mask image, which is used to identify the area to be repaired in the first intermediate video frame;
[0186] The first repair unit is configured to determine the pixel value of at least one target pixel in the repair area of the first intermediate video frame based on the optical flow information and the adjacent video frames of the initial video frame, and fill the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image based on the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image.
[0187] The second repair unit is used to repair the second intermediate video frame by calling a preset video repair model based on the optical flow information, the second intermediate video frame and the intermediate mask image, so as to obtain the target video frame.
[0188] According to one or more embodiments of this disclosure, when the first repair unit fills the corresponding pixels of the target pixel in the first intermediate video frame and the corresponding pixels of the initial mask image according to the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image, it is used to:
[0189] The pixel value of the target pixel is filled into the target pixel of the first intermediate video frame to obtain the second intermediate video frame;
[0190] The corresponding pixel of the target pixel in the initial mask image is determined, and the corresponding pixel in the initial mask image is masked to obtain the intermediate mask image.
[0191] According to one or more embodiments of this disclosure, when the first repair unit determines the pixel value of at least one target pixel within the repair area of the first intermediate video frame based on the optical flow information and adjacent video frames of the initial video frame, it is configured to:
[0192] Based on the optical flow information, determine the corresponding pixel of at least one target pixel in the repair area of the first intermediate video frame in the adjacent video frame, and determine the pixel value of the target pixel based on the corresponding pixel in the adjacent video frame.
[0193] According to one or more embodiments of this disclosure, when the first repair unit determines the pixel value of the target pixel based on corresponding pixels in the adjacent video frames, it is configured to:
[0194] Determine the pixel value that appears most frequently among the corresponding pixel values in the adjacent video frames, and set the pixel value that appears most frequently as the pixel value of the target pixel; or
[0195] The pixel values of corresponding pixels in the adjacent video frames are weighted and averaged, and the weighted average result is determined as the pixel value of the target pixel.
[0196] According to one or more embodiments of this disclosure, when the first repair unit fills the corresponding pixels of the target pixel in the first intermediate video frame and the corresponding pixels of the initial mask image according to the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image, it is used to:
[0197] Determine the confidence level of the pixel value of the target pixel. If the confidence level is higher than a preset confidence threshold, then fill the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image according to the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image.
[0198] According to one or more embodiments of this disclosure, when the second repair unit repairs the second intermediate video frame by calling a preset video repair model based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame, it is configured to:
[0199] Based on the optical flow information, the second intermediate video frame, and the intermediate mask image, the preset video repair model is invoked to repair the area to be repaired in the second intermediate video frame to obtain the target video frame, wherein the area to be repaired in the second intermediate video frame is the area in the second intermediate video frame that corresponds to the non-masked area of the intermediate mask image.
[0200] According to one or more embodiments of this disclosure, when the second repair unit calls the preset video repair model to repair the area to be repaired in the second intermediate video frame based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame, it is used to:
[0201] Acquire an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image;
[0202] The optical flow information, the image of the target region in the current second intermediate video frame, and the image of the target region in the adjacent second intermediate video frames are input into the encoder of the preset video restoration model to extract feature maps.
[0203] The feature map is input into the decoder of the preset video restoration model to generate the target video frame.
[0204] According to one or more embodiments of this disclosure, when the second repair unit calls the preset video repair model to repair the area to be repaired in the second intermediate video frame based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame, it is used to:
[0205] Acquire an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image;
[0206] The current second intermediate video frame is reduced to the size of the target area to obtain the reduced current second intermediate video frame;
[0207] The optical flow information, the scaled-down current second intermediate video frame, the image of the target region of the current second intermediate video frame, and the image of the target region of the adjacent second intermediate video frames are input into the encoder of the preset video restoration model to extract feature maps.
[0208] The feature map is input into the decoder of the preset video restoration model to generate the target video frame.
[0209] According to one or more embodiments of this disclosure, when the second repair unit calls the preset video repair model to repair the area to be repaired in the second intermediate video frame based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame, it is used to:
[0210] Acquire an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image;
[0211] The optical flow information, the image of the target region in the current second intermediate video frame, and the image of the target region in the adjacent second intermediate video frames are input into the encoder of the preset video restoration model to extract the first feature map;
[0212] Extract features from the current second intermediate video frame to obtain the second feature map;
[0213] The first feature map and the second feature map are fused to obtain a fused feature map;
[0214] The fused feature map is input into the decoder of the preset video restoration model to generate the target video frame.
[0215] According to one or more embodiments of this disclosure, when acquiring optical flow information of the video to be processed, the optical flow acquisition unit is used to:
[0216] Based on the video to be processed, the initial optical flow information is obtained by calling the preset optical flow model;
[0217] Based on the video to be processed and the initial optical flow information, a preset optical flow completion model is invoked to complete the initial optical flow information, thereby obtaining the optical flow information of the video to be processed.
[0218] According to one or more embodiments of this disclosure, when determining the location of a target object in any initial video frame of the video to be processed, the target object recognition unit is configured to:
[0219] Identify all target objects in any initial video frame of the video to be processed;
[0220] Select target objects that meet preset filtering conditions from all target objects in the initial video frame, and determine the location of the target objects that meet the preset filtering conditions.
[0221] According to one or more embodiments of this disclosure, the target object is a subtitle.
[0222] Thirdly, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;
[0223] The memory stores computer-executed instructions;
[0224] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the video processing method as described in the first aspect and various possible designs of the first aspect.
[0225] Fourthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, and when a processor executes the computer-executable instructions, the video processing method described in the first aspect and various possible designs of the first aspect is implemented.
[0226] Fifthly, according to one or more embodiments of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the video processing method described in the first aspect and various possible designs of the first aspect.
[0227] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0228] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0229] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A video processing method, characterized in that, include: Obtain the optical flow information of the video to be processed; The location of the target object in any initial video frame of the video to be processed is determined, and the initial video frame is segmented according to the location of the target object to obtain a first intermediate video frame with the target object removed and an initial mask image; wherein the target object is located in a non-mask area in the initial mask image, which is used to identify the area to be repaired in the first intermediate video frame; Based on the optical flow information and the adjacent video frames of the initial video frame, the pixel value of at least one target pixel in the repair area of the first intermediate video frame is determined. Based on the pixel value of the target pixel, the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image are filled to obtain the second intermediate video frame and the intermediate mask image. Based on the optical flow information, the second intermediate video frame, and the intermediate mask image, a preset video restoration model is invoked to restore the second intermediate video frame, thereby obtaining the target video frame.
2. The method according to claim 1, characterized in that, The step of filling the target pixel of the first intermediate video frame and the corresponding pixels of the initial mask image according to the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image includes: The pixel value of the target pixel is filled into the target pixel of the first intermediate video frame to obtain the second intermediate video frame; The corresponding pixel of the target pixel in the initial mask image is determined, and the corresponding pixel in the initial mask image is masked to obtain the intermediate mask image.
3. The method according to claim 1, characterized in that, The step of determining the pixel value of at least one target pixel within the repair area of the first intermediate video frame based on the optical flow information and adjacent video frames of the initial video frame includes: Based on the optical flow information, determine the corresponding pixel of at least one target pixel in the repair area of the first intermediate video frame in the adjacent video frame, and determine the pixel value of the target pixel based on the corresponding pixel in the adjacent video frame.
4. The method according to claim 3, characterized in that Determining the pixel value of the target pixel based on corresponding pixels in the adjacent video frames includes: Determine the pixel value that appears most frequently among the corresponding pixel values in the adjacent video frames, and set the pixel value that appears most frequently as the pixel value of the target pixel; or The pixel values of corresponding pixels in the adjacent video frames are weighted and averaged, and the weighted average result is determined as the pixel value of the target pixel.
5. The method according to any one of claims 1-4, characterized in that, The step of filling the target pixel of the first intermediate video frame and the corresponding pixels of the initial mask image according to the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image includes: Determine the confidence level of the pixel value of the target pixel. If the confidence level is higher than a preset confidence threshold, then fill the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image according to the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image.
6. The method according to any one of claims 1-4, characterized in that, The step of repairing the second intermediate video frame by calling a preset video restoration model based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame includes: Based on the optical flow information, the second intermediate video frame, and the intermediate mask image, the preset video repair model is invoked to repair the area to be repaired in the second intermediate video frame to obtain the target video frame, wherein the area to be repaired in the second intermediate video frame is the area in the second intermediate video frame that corresponds to the non-masked area of the intermediate mask image.
7. The method according to claim 6, characterized in that, The step of invoking the preset video restoration model to restore the region to be restored in the second intermediate video frame based on the optical flow information, the second intermediate video frame, and the intermediate mask image to obtain the target video frame includes: Acquire an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image; The optical flow information, the image of the target region in the current second intermediate video frame, and the image of the target region in the adjacent second intermediate video frames are input into the encoder of the preset video restoration model to extract feature maps. The feature map is input into the decoder of the preset video restoration model to generate the target video frame.
8. The method according to claim 6, characterized in that, The process involves using the optical flow information, the second intermediate video frame, and the intermediate mask image to call the preset video restoration model to restore the region to be restored in the second intermediate video frame, thereby obtaining the target video frame. Acquire an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image; The current second intermediate video frame is reduced to the size of the target area to obtain the reduced current second intermediate video frame; The optical flow information, the scaled-down current second intermediate video frame, the image of the target region of the current second intermediate video frame, and the image of the target region of the adjacent second intermediate video frames are input into the encoder of the preset video restoration model to extract feature maps. The feature map is input into the decoder of the preset video restoration model to generate the target video frame.
9. The method according to claim 6, characterized in that, The process involves using the optical flow information, the second intermediate video frame, and the intermediate mask image to call the preset video restoration model to restore the region to be restored in the second intermediate video frame, thereby obtaining the target video frame. Acquire an image of the target region within the current second intermediate video frame and images of the target region within a preset number of adjacent second intermediate video frames, wherein the target region is a region corresponding to a preset neighborhood of the non-masked region of the intermediate mask image; The optical flow information, the image of the target region in the current second intermediate video frame, and the image of the target region in the adjacent second intermediate video frames are input into the encoder of the preset video restoration model to extract the first feature map; Extract features from the current second intermediate video frame to obtain the second feature map; The first feature map and the second feature map are fused to obtain a fused feature map; The fused feature map is input into the decoder of the preset video restoration model to generate the target video frame.
10. The method according to claim 1, characterized in that, The acquisition of optical flow information of the video to be processed includes: Based on the video to be processed, the initial optical flow information is obtained by calling the preset optical flow model; Based on the video to be processed and the initial optical flow information, a preset optical flow completion model is invoked to complete the initial optical flow information, thereby obtaining the optical flow information of the video to be processed.
11. The method according to claim 1, characterized in that, Determining the location of the target object in any initial video frame of the video to be processed includes: Identify all target objects in any initial video frame of the video to be processed; Select target objects that meet preset filtering conditions from all target objects in the initial video frame, and determine the location of the target objects that meet the preset filtering conditions.
12. The method according to claim 1, characterized in that, The target object is the subtitle.
13. A video processing device, characterized in that, include: Optical flow acquisition unit, used to acquire optical flow information of the video to be processed; The target object recognition unit is used to determine the position of the target object in any initial video frame of the video to be processed, and to perform target object segmentation on the initial video frame according to the target object position to obtain a first intermediate video frame with the target object removed and an initial mask image. In the initial mask image, the target object is located in the non-masked area, which is used to identify the area to be repaired in the first intermediate video frame; The first repair unit is configured to determine the pixel value of at least one target pixel in the repair area of the first intermediate video frame based on the optical flow information and the adjacent video frames of the initial video frame, and fill the target pixel of the first intermediate video frame and the corresponding pixel of the initial mask image based on the pixel value of the target pixel to obtain the second intermediate video frame and the intermediate mask image. The second repair unit is used to repair the second intermediate video frame by calling a preset video repair model based on the optical flow information, the second intermediate video frame and the intermediate mask image, so as to obtain the target video frame.
14. An electronic device, characterized in that, include: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-11.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-11.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-11.
Citation Information
Patent Citations
Video restoration method, related device, equipment and storage medium
CN115170400A
Video target segmentation method and system for removing background interference through semi-supervised learning
CN115830505A