Video processing method, device, equipment and medium
The video processing method allows users to specify target areas through smearing operations, using segmentation algorithms on a user terminal for flexible and efficient object extraction, addressing the limitations of conventional multimedia clip software.
Patent Information
- Application Number
- JP2025522583
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-26
- Filing Date
- 2023-12-11
- Publication Date
- 2025-10-22
AI Technical Summary
Conventional multimedia clip software lacks flexibility in identifying and extracting objects other than those predetermined by its algorithms, such as humans or faces, and requires image upload to servers for processing, leading to inefficiencies and security concerns.
A video processing method that allows users to specify target areas in a video through smearing operations, using segmentation algorithms to determine and track or reverse-track these areas locally on a user terminal, enabling flexible and efficient object extraction without server interaction.
Enables flexible and efficient object extraction in videos, improving user experience and reducing processing time and costs while ensuring real-time capabilities and information security.
Smart Images

Figure 2025535169000001_ABST
Abstract
Description
[Technical Field]
[0001] [Cross-Citation of Related Applications] This application claims priority from Chinese Patent Application No. 202211677133.3, entitled "Video Processing Method, Apparatus, Device and Medium," filed on December 26, 2022, the entire contents of which are incorporated herein by reference.
[0002] [Technical field] The present disclosure relates to the technical field of video processing, and more particularly to a video processing method, apparatus, device and medium. [Background technology]
[0003] Conventional multimedia clip software can provide users with features such as smart matting or animation special effects, which can automatically identify system-identifiable objects in an image, such as people, cats, and dogs, and then perform operations such as matting or adding special effects based on the identified object regions. However, the effects that these features can achieve are entirely determined by the algorithms used by the software itself. For example, if the software uses a human matting algorithm, it can only identify and extract all human images in an image, while if the software uses a face special effects algorithm, it can only identify all faces in an image and add special effects to them. These methods lack flexibility and are difficult to fully meet user needs. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure provides an image processing method, apparatus, device, and medium. [Means for solving the problem]
[0005] An embodiment of the present disclosure provides a video processing method, the method including: acquiring a smear area of a target image in a target video, the smear area being an area determined based on a smearing operation on the target image by a user; inputting area information of the target image and the smear area corresponding to the target image into a segmentation algorithm to obtain a target area in the target image output by the segmentation algorithm, the target area including an object area specified by a user, the object area specified by the user being an area of a main object to which the smear area belongs or the smear area; tracking a first image in the target video based on the target area in the target image to determine the target area in the first image, and / or reverse-tracking a second image in the target video based on the target area in the target image to determine the target area in the second image, the first image being a frame image after the target image in the target video, and the second image being a frame image before the target image in the target video.
[0006] An embodiment of the present disclosure further provides a video processing device, the video processing device including: a smear area acquisition module for acquiring a smear area of a target image in a target video, the smear area being an area determined based on a smearing operation on the target image by a user; and a first target area determination module for inputting area information of the target image and the smear area corresponding to the target image into a segmentation algorithm to obtain a target area in the target image output by the segmentation algorithm, the target area including an object area specified by a user, the object area specified by the user being an area of a main object to which the smear area belongs. a first target area determination module for determining a target area in the first image by tracking a first image in the target video based on the target area in the target image, and / or a second target area determination module for determining a target area in the second image by reverse tracking a second image in the target video based on the target area in the target image, wherein the first image is a frame image in the target video that follows the target image, and the second image is a frame image in the target video that precedes the target image.
[0007] An embodiment of the present disclosure further provides an electronic device, the electronic device including a processor and a memory for storing instructions executable by the processor, the processor being used to read the executable instructions from the memory and execute the instructions to realize a video processing method according to an embodiment of the present disclosure.
[0008] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium having a computer program stored thereon, the computer program being used to perform a video processing method according to an embodiment of the present disclosure.
[0009] It should be understood that the contents described in this section are not intended to represent key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become more readily apparent from the following description. [Brief explanation of the drawings]
[0010] The drawings herein, which are incorporated in and constitute a part of this specification, illustrate preferred embodiments of the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.
[0011] In order to more clearly explain the technical solutions in the embodiments or related technologies of the present disclosure, the following briefly introduces drawings that need to be used in the description of the embodiments or related technologies, and it is obvious that those skilled in the art can derive other drawings based on these drawings without exerting any creative effort. [Figure 1] 1 is a flowchart of a video processing method according to an embodiment of the present disclosure. [Figure 2] FIG. 1 is a principle schematic diagram of a preset matting algorithm according to an embodiment of the present disclosure. [Figure 3] 1 is a flowchart of a matting process according to an embodiment of the present disclosure. [Figure 4] FIG. 1 is a schematic diagram of task execution according to an embodiment of the present disclosure. [Figure 5] 10 is a flowchart of an execution of an interaction splitting task according to an embodiment of the present disclosure. [Figure 6] 10 is a flowchart illustrating converting brush data into a matting result according to an embodiment of the present disclosure. [Figure 7] 10 is a flowchart of the execution of a tracking task according to an embodiment of the present disclosure. [Figure 8] 1 is a structural schematic diagram of a video processing device according to an embodiment of the present disclosure; [Figure 9] 1 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0012] In order to make the above-mentioned objects, features and advantages of the present disclosure more clearly understood, the present disclosure will be further described below. It should be noted that, if there is no conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0013] In order to fully understand the present disclosure, numerous details are set forth in the following description; however, the present disclosure may be embodied in forms different from those described herein, and it is to be understood that the embodiments in the specification are merely some embodiments of the present disclosure, and not all embodiments.
[0014] Through research, the inventors have found that conventional smart matting functions or functions such as animation special effects lack flexibility. Taking the smart matting function as an example, if the matting algorithm used by a system can only identify human figures, the smart matting function can only identify human figures in the image to be processed, but cannot identify and extract other object types. For example, a user cannot extract a cat from the image using the smart matting function. Furthermore, if multiple human figures are present in the image to be processed, the conventional matting function will simultaneously extract multiple human figures, preventing the user from extracting only one of the human figures required. To address these issues, the inventors have provided a video processing method, device, equipment, and media, which will be described in detail below.
[0015] In the technical solution according to the embodiments of the present disclosure, a user can flexibly specify the required target area (which may include the smear area or the area of the main object to which the smear area belongs) in the target image in the video by using the smearing method, and then easily and quickly determine the target areas of other frame images in the video by using the tracking and / or reverse tracking method, and the target areas determined by the above methods can flexibly and easily meet the user's individual needs.
[0016] FIG. 1 is a flowchart of a video processing method according to an embodiment of the present disclosure, which may be performed by a video processing device, which may be implemented by adopting software and / or hardware, and may generally be integrated into an electronic device, which may be a user terminal such as a mobile phone, a computer, etc. As shown in FIG. 1, the method mainly includes the following steps S102 to S106.
[0017] Step S102: Obtain a smear area of the target image in the target video, where the smear area is an area determined based on a smear operation on the target image by a user.
[0018] The target video is a video to be processed. The embodiments of the present disclosure are not limited to the video content of the target video. The target image may be any frame image selected by the user from the target video, and is not particularly limited thereto. In some embodiments, the smearing operation is an operation in which the user uses a virtual brush to smear on the target image. For example, an interaction interface provides the user with a virtual brush. The user clicks on the virtual brush and then controls the virtual brush to form a smearing path on the target image, for example, by sliding a finger or a mouse. The path of the finger sliding on the target image becomes the smearing path of the virtual brush, and the thickness of the virtual brush can be set by the user. Accordingly, the step of obtaining a smeared area on the target image includes obtaining a brush thickness of the virtual brush selected by the user and information about the smearing path the user uses the virtual brush to smear on the target frame image, and determining a smeared area on the target image based on the brush thickness and the smearing path information. In some specific embodiments, the smear area may be determined, for example, by a feature point sampling method, for example, by determining the brush stroke outline based on the brush thickness and the smear trajectory, and then determining the smear area based on the stroke outline.
[0019] In step S104, the target image and the area information of the smeared area corresponding to the target image are input into a segmentation algorithm to obtain a target area in the target image output by the segmentation algorithm. The target area includes an object area designated by a user, which is the area of the main object to which the smeared area belongs or the smeared area. The embodiments of the present disclosure do not limit the function of the target area. For example, the target area may be used for matting processing, i.e., the target area may be the image area to be extracted. For example, the target area may be used for special effects processing. For example, if the target area is a human image smeared by a user, designated special effects such as a headpiece or facial feature modification may be added to the human image. The above is merely illustrative and should not be considered limiting. The segmentation algorithm can determine the target area to be extracted in the target image based on the area information of the smeared area corresponding to the target image. The area information of the smear area is used to identify the location of the smear area. For example, the area information of the smear area may be expressed in the form of a mask of the smear area (i.e., a mask image). The size of the smear area mask and the size of the target image are the same, but the pixels in the smear area mask where the smear area is located are not 0 (for example, all are 1), and the pixels in the background area other than the smear area are all 0. In this manner, the location of the smear area in the image can be clearly identified.
[0020] In practical applications, the object region may be the smeared region alone, or the region of the main object to which the smeared region belongs. Specifically, the object region may be determined by user settings. For example, the method for determining the object region differs depending on the brush type selected by the user. In some embodiments, the object region specified by the user is the region of the main object to which the smeared region belongs. For example, if the user smears a cat's body with a single stroke, i.e., if the smeared region is located in the area of the cat's body, the main object to which the smeared region belongs will be the cat. It should also be noted that the main object classification method may be configured as needed. For example, a complete individual may be defined as a main object, such as a cat or a tree. Complete localized parts may also be defined as main objects, such as a cat's head or a tree trunk, or a cat's ears or leaves. The classification granularity of the main object may be flexibly configured as needed and is not particularly limited herein. It can be understood that the area of the main object to which the smear area belongs is larger than the smear area, and the user does not need to completely smear or draw an outline of the entire main object to be extracted. Instead, the user can simply smear, for example, draw a line on the body of a cat, and the segmentation algorithm can quickly determine that the area to be extracted is the entire cat, thereby improving matting efficiency and user experience. In some other embodiments, when the object area specified by the user is a smear area, it means that only the part smeared by the user will be extracted. The method of determining the target area to be extracted based on the smear area can be flexibly selected by the user as needed, and is not particularly limited here.
[0021] Step S106: based on the target area in the target image, track a first image in the target video to determine a target area in the first image, and / or based on the target area in the target image, reverse track a second image in the target video to determine a target area in the second image, where the first image is a frame image that follows the target image in the target video, and the second image is a frame image that follows the target image in the target video. That is, an image before the target image may be reverse tracked, or an image after the target image may be tracked. After determining the target area of the target image, target areas of other frame images in the target video can be directly determined easily and quickly by tracking or reverse tracking, and the target areas of other frame images may also be used for matting processing, etc., and are not particularly limited here.
[0022] In the above technical solution according to the embodiments of the present disclosure, a user can flexibly specify the required target area (which may include the smear area or the area of the main object to which the smear area belongs) in the target image in the video by using the smearing method, and then can easily and quickly determine the target areas of other frame images in the video by using the tracking and / or reverse tracking method, and the target areas determined by the above methods can well meet the user's individual needs.
[0023] It should also be noted that in the related technology, the matting function needs to be realized via a server, and users need to upload images to the server to realize the matting function. This not only raises certain concerns about information security, but also requires interaction flows such as image transmission between the server and the system, which consumes time and communication traffic, has poor real-time capabilities, low matting efficiency, and requires high matting costs. The above method according to the embodiments of the present disclosure may be executed by a user terminal, and can realize a method in which the user selects a target area by the user terminal. Based on the smear area determined by the user's smearing operation, a segmentation algorithm is used to determine the target area that the user needs to extract (including the object area specified by the user). That is, the user terminal can combine front-end interactions (e.g., the user's smearing operation) with background algorithms (e.g., the segmentation algorithm), thereby reliably realizing a matting function customized by the user locally on the user terminal. This not only fully meets the user's image matting needs, but also eliminates the need to upload images to a server for processing, thereby avoiding problems such as time consumption and communication traffic consumption caused by interactions between the server and the mobile terminal, thereby ensuring information security and effectively ensuring real-time processing, which is helpful in improving matting efficiency and reducing matting costs.
[0024] The method further includes acquiring a region to be superimposed corresponding to the target image, where the region to be superimposed corresponding to the target image is a region that needs to be acquired from the target image before the smearing operation. For example, assuming that a user has previously enabled a preset matting function (which may also be referred to as a smart matting function), the smart matting function is a function for extracting a region of a preset type object in an image using a preset matting algorithm. The preset type object is an object type included in the training sample of the preset matting algorithm. For example, the preset matting algorithm is a human image matting algorithm trained with training samples including a human image. In this case, the smart matting function can automatically extract the human image, and the region to be superimposed needs to include the human image in the target image. The region to be superimposed may also include a custom object specified by the user for another image preceding the target image. For example, if the target image belongs to an image in a target video and the user sets the need to extract a dog image from another image in the target video, the region to be superimposed corresponding to the target image may include the dog image. In addition, the area to be superimposed may be an empty area, and for example, if there is no need to extract other areas from the target image before the smearing operation, the area to be superimposed may be considered to be empty, or the area information of the area to be superimposed may be considered to be empty data.
[0025] Furthermore, the step of inputting the target image and the area information of the smear area corresponding to the target image into the segmentation algorithm includes inputting the target image, the area information of the smear area corresponding to the target image, and the area information of the area to be superimposed into the segmentation algorithm, where the target area further includes the area to be superimposed. That is, the segmentation algorithm can comprehensively determine the target area to be extracted using the area information of the smear area and the area information of the area to be superimposed. For example, the target area may not only include the object area determined based on the smear area, but may also include the area to be superimposed.
[0026] In some embodiments, the embodiment of obtaining the region to be superimposed corresponding to the target image may be implemented with reference to (1) and (2) below.
[0027] (1) When the preset matting function is enabled, the area of a preset type object in the target image is obtained based on the preset matting algorithm, and the area of the preset type object is used as the area to be superimposed. Here, the preset type object is an object type included in the training sample of the preset matting algorithm. The embodiments of the present disclosure are not limited to the preset type object, and may be, for example, a human figure. In this case, the smart matting function automatically extracts the human figure. It should be noted that the user cannot control which specific preset type object the smart matting function extracts. The extractable preset type objects of the smart matting function are determined by the background algorithm. The user can only select whether to enable the smart matting function.
[0028] (2) When the preset matting function is off or there is no preset type object in the target image, it is determined that the area to be superimposed corresponding to the target image is an empty area. An empty area is an area that does not exist in the target image. Specifically, meaningless data may be adopted to represent the area to be superimposed.
[0029] In some embodiments, the segmentation algorithm determines a target region to be extracted in the target image according to the following steps a to b.
[0030] Step a: Determine an object area designated by a user based on area information of the smear area.
[0031] If the smearing operation is an operation in which a user uses a virtual brush to smear on a target frame image, then step a above may be implemented by referring to the following steps: Obtain the brush type of the virtual brush selected by the user. If the brush type is a first type, determine that the object area designated by the user is a main object area belonging to the smear area; if the brush type is a second type, determine that the object area designated by the user is a smear area. In practical applications, different types of virtual brushes can be provided for the user. For example, the first type of brush is a fast brush, and the user can determine that the main object belonging to the smear area by simply smearing on the object requiring extraction. For example, the user can determine that the object area designated by the user is a cat image by simply drawing a stroke on the cat image requiring extraction. The second type of brush is a normal brush, which means that only the user's smeared area is extracted. This smeared area is the area of the handwriting outline drawn by the user on the image, and the object area designated by the user is only the area of the handwriting outline.
[0032] Step b: obtaining a target area to be extracted in the target image based on area information of the object area corresponding to the target image and area information of the area to be superimposed corresponding to the target image.
[0033] In some specific embodiments, the region information of the object region includes a mask of the object region, and the region information of the region to be superimposed includes a mask of the region to be superimposed. The masks can be identified by adopting different pixel values for different regions. For example, the object region mask, the region to be superimposed mask, and the target image all have the same size, but in the object region mask, the pixel value of the object region is 1 and the pixel values of the remaining background region are all 0; and in the region to be superimposed mask, the pixel value of the region to be superimposed is 1 and the pixel values of the remaining background region are all 0. The masks can clearly mark the regions that need to be extracted, and can then be merged with the target image, i.e., the object region and the region to be superimposed can be extracted.
[0034] Furthermore, the above step b may be implemented by referring to the following steps: merging the mask of the object region corresponding to the target image with the mask of the region to be superimposed corresponding to the target image, and determining the target region to be extracted in the target image based on the merged mask. That is, the masks of all the regions to be extracted are merged into one mask, and the merged mask can clearly indicate the target region to be extracted according to the above pixel marking method. For example, all regions with a pixel value of 1 on the merged mask belong to the target region. By merging the masks, the merged mask can be directly used to merge with the target image, which helps to easily and quickly extract the target region from the target image.
[0035] In some embodiments, the target image belongs to a target video, i.e., the target image is a frame image in the target video. The embodiments of the present disclosure can not only determine a target area to be extracted in the target image to facilitate matting processing for the target image, but also extract a target that requires user extraction from the entire video. For example, the above method according to the embodiments of the present disclosure can further include tracking and / or reverse tracking, and may be specifically performed with reference to the following (1) and / or (2).
[0036] (1) A specific embodiment of tracking a first image in a target video (a frame image that follows the target image in the target video) based on a target area in a target image to determine a target area in the first image may be performed with reference to the following steps A1 to A3.
[0037] Step A1: The frame images in the target video after the target image are selected as the first images in order from front to back until the end frame is reached, where the end frame may be the last frame of the target video or a frame set by the user.
[0038] Step A2: Obtain the area to be superimposed corresponding to the first image. The area to be superimposed corresponding to the first image is the remaining area that needs to be extracted other than the area of the main object to which the smear area of the target image belongs. For example, if the user has previously turned on a preset matting function (for example, to intelligently extract a person's image), the area to be superimposed includes the person's image area in the first image. The area to be superimposed corresponding to the first image is similar to the method of determining the area to be superimposed corresponding to the target image, and will not be further described here.
[0039] Step A3: Input the region information of the first image and the region to be superimposed corresponding to the first image into a tracking algorithm, and the tracking algorithm determines a target region in the first image based on the region information of the object region of the target image and the region information of the region to be superimposed corresponding to the first image. The tracking algorithm can determine the region of the main object in the first image that belongs to the smear region based on the region information of the object region of the target image, and further combine it with the region to be superimposed to obtain the target region to be extracted in the first image. For example, if the region to be superimposed is a human figure, and the user instructs the target image to be smeared to extract a cat figure (i.e., the user customizes the extraction of a cat figure), the tracking algorithm can determine the cat figure in the first image using a tracking method, and in combination with the region to be superimposed, determine that the region to be extracted in the first image is a human figure and a cat figure.
[0040] (2) A specific embodiment of reverse tracking a second image in the target video based on a target area in the target image to determine a target area in the second image may be performed with reference to the following steps B1 to B3.
[0041] Step B1: Frame images in front of the target image in the target video are sequentially taken as second images in the order from rear to front.
[0042] Step B2: Obtain the area to be superimposed corresponding to the second image.
[0043] In step B3, the second image and the area information of the area to be superimposed corresponding to the second image are input into a reverse tracking algorithm, and the reverse tracking algorithm determines the target area to be extracted in the second image based on the area information of the object area of the target image and the area information of the area to be superimposed corresponding to the second image. The reverse tracking algorithm is similar to the tracking algorithm, except that the tracking algorithm processes frame images from front to back, while the reverse tracking algorithm processes frame images from back to front. For example, assuming that the target video has a total of 200 frames and the user performs a smear operation on the 100th frame (target image), the tracking algorithm processes the images from the 100th frame to the 200th frame from front to back, and the reverse tracking algorithm processes the images from the 100th frame to the 1st frame from back to front.
[0044] The specific implementation manner of the step of obtaining the target area to be extracted in the second image by the reverse tracking algorithm may refer to the relevant content of the step of obtaining the target area to be extracted in the second image by the tracking algorithm above, and will not be further described here.
[0045] In practical applications, the above tracking algorithm and / or reverse tracking algorithm can be flexibly selected as needed, for example, only the target area corresponding to the frame image after the target image in the target video can be extracted, or only the target area corresponding to the frame image before the target image in the target video can be extracted, or all the target areas in the frame images of the entire target video can be extracted, thereby meeting various needs of users.
[0046] In some embodiments, a target region corresponding to an arbitrary frame image in the target video includes a region to be superimposed corresponding to the frame image and an object region corresponding to the frame image. The object region corresponding to the frame image is the region in the frame image of the main object to which the smear region belongs. This arbitrary frame image may be, for example, the first image, the second image, or the target image. In practical applications, the frame image may simultaneously segment an object region (e.g., a cat image) specified by a user and segment other regions, such as a person image that requires extraction by a preset matting algorithm. In specific implementation, when performing custom segmentation, the region that previously required segmentation may be superimposed, and then a mask may be output in which the cat image and the person image are simultaneously labeled, thereby allowing the cat image and the person image to be simultaneously extracted based on this mask. Furthermore, region information of the region to be superimposed included in the target region corresponding to an arbitrary frame image in the target video is obtained according to the following steps 1 to 4.
[0047] Step 1: When the preset matting function is on for any frame image in the target video, it is determined whether the area information of the area to be superimposed corresponding to this frame image has been successfully read.
[0048] Step 2: if not read, input this frame image into a preset matting algorithm to obtain region information of a preset type object corresponding to this frame image through the preset matting algorithm, where the preset type object is an object type included in the training sample of the preset matting algorithm, and the preset matting function is a function for extracting the region of the preset type object in the image through the preset matting algorithm.
[0049] For ease of understanding, the following example will be used: In actual application, a user performs custom matting after turning on the preset matting function, and the turned-on preset matting function employs a preset matting algorithm (e.g., a human image matting algorithm) to automatically extract human images in video frame images, that is, it is assumed that the area to be superimposed contains a human image. For a certain frame image, if the preset matting algorithm has already been applied to this frame image, the area information (e.g., human image information) of the area to be superimposed corresponding to this frame image can be directly read. If the preset matting algorithm has not yet been applied to this frame image, for example, if the user suddenly drags from frame 110 to frame 180, the preset matting algorithm will not be able to be applied to frame 180 in time. In this case, the area to be superimposed corresponding to frame 180 has not yet been read. Therefore, the preset matting algorithm can be used again to process the 180th frame image to obtain the human image information corresponding to the 180th frame image.
[0050] Furthermore, the embodiment of the present disclosure further provides a principle schematic diagram of the preset matting algorithm, as shown in Figure 2, and takes the preset matting algorithm as an example, showing that the original frame image is proportionally reduced to obtain a sub-image. For example, the sub-image can be obtained by proportionally reducing the short side length of the original frame image, and then inputting the sub-image into a human body segmentation algorithm model, performing an edge smoothing operation to obtain a human image mask corresponding to the sub-image, and then performing an enlargement process on the human image mask to obtain a human image mask corresponding to the original frame image, that is, the size of the human image mask is the same as the size of the original frame image, the pixel values of the human image region are not 0 (e.g., 1), and the pixel values of the non-human image region are 0. Using the above method, the human image region can be clearly identified from the frame image.
[0051] Step 3: The area information of the area to be superimposed corresponding to this read frame image, or the area information of the preset type object corresponding to this frame image obtained by the preset matting algorithm, is used as the area information of the area to be superimposed corresponding to this frame image.
[0052] Step 4: If the preset matting function is off or there is no preset-type object in this frame image, construct empty data and use the empty data as the region information of the region to be superimposed corresponding to this frame image. If the user does not enable the preset matting function, a completely custom matting is used, or there is no human figure in the frame image that requires extraction by the preset matting function, empty data of the same size as the frame image can be created. Specifically, the empty data can be a mask of the region to be superimposed, but all pixel values in the mask are unified and are the same special numerical value, which has no practical meaning and can be understood as meaning that there is no region to be superimposed in the frame image. The purpose of using the empty data as the region information of the region to be superimposed is to preserve the interface for subsequent algorithm upgrades and to support the use of various settings other than the preset matting function to superimpose extracted regions.
[0053] In some other embodiments, the target region corresponding to a given frame image in the target video includes only the object region corresponding to this frame image, where the object region corresponding to this frame image is the region in this frame image of the main object to which the smear region belongs. That is, custom segmentation may be performed only on the frame image to segment the object region specified by the user. When using a segmentation algorithm to segment the user-specified object region, segmentation of other regions, such as human images, that require extraction by a preset matting algorithm is not performed; the output result of the segmentation algorithm is a standalone result. For example, a custom object region (e.g., the aforementioned cat image) may be segmented using a custom matting algorithm, and a preset type object region (e.g., the aforementioned human image) may be segmented using a preset matting algorithm. Then, if image matting is required, the custom object region and the preset type object region may be merged. For example, a cat image mask and a human image mask may be obtained, and then the cat image mask and the human image mask may be merged. The merged mask may then be used to process the image to extract the cat image and the human image. In addition, the above method further includes the following steps 1) and 2).
[0054] Step 1), if the preset matting function is on, input each frame image in the target video into the preset matting algorithm, and determine the area of the preset type object in each frame image through the preset matting algorithm.
[0055] Step 2) When performing matting processing, the mask of the target area corresponding to each frame image is integrated with the mask of the area of the preset type object, and matting processing is performed based on the integrated mask corresponding to each frame image.
[0056] Custom matting and smart matting (i.e., the preset matting function described above) can be performed separately and then integrated again when necessary. In actual applications, the required processing method can be flexibly selected according to needs, and custom matting and smart matting can be performed simultaneously or separately, and there is no particular limitation here.
[0057] To summarize the above, in practical applications, during the matting process for any frame image in the target video, all regions requiring extraction may be output simultaneously, and regions previously requiring extraction may be included in addition to the object region customized by the user. Only the object region customized by the user may be output, without extracting the region previously requiring extraction, or the region previously requiring extraction may be extracted separately. For example, assume that the target video has a total of 200 frames, and the user enables the smart matting function in the first frame and uses this preset matting function to automatically extract a human figure. The user then indicates that a cat figure needs to be extracted by smearing in the 100th frame. In this case, the human figure and the cat figure may be extracted simultaneously from the frame images, or they may be extracted separately and then merged. Note that superimposition may not be performed; for example, the cat figure may be extracted from the images from the 100th to 200th frames without extracting the human figure.
[0058] For ease of understanding, the following provides a specific example of applying the above-described aspects of the embodiments of the present disclosure. In this example, matting is performed on an entire video. For the overall flow, refer to the matting flowchart shown in FIG. 3. A user performs a smearing operation on the xth frame image of a target video, where the xth frame image is the target image. A segmentation task may then be performed based on the xth frame image, a tracking task may be performed on images after the xth frame image, and a reverse tracking task may be performed on images before the xth frame image. FIG. 3 is merely an illustrative example and should not be considered limiting. For example, in a practical application, only one of the tracking task or the reverse tracking task may be performed. For example, end frames may be set for the tracking task and the reverse tracking task, respectively. The end frames may not necessarily be the last frame or the first frame of the video. For ease of understanding, reference may be made to the task execution schematic diagram shown in Figure 4, which simply illustrates the execution node of the interaction task (i.e., executing a segmentation task in response to the user's smear operation, which may also be called an interaction segmentation task), and the processing direction of the frame images of the reverse tracking task and the tracking task. As shown in Figure 4, the reverse tracking task is processed from back to front, and the tracking task is processed from front to back. Specifically, reference may be made to the related content mentioned above, and no further description will be given here.
[0059] The interaction division task and the tracking task / reverse tracking task will be specifically described below.
[0060] The execution mode of the interaction segmentation task may be seen in Figure 5. The SDK (Software Development Kit) can obtain the brush type, brush thickness, and brush trajectory data points through the UI layer (which may also be referred to as the APP layer), and obtain the brush mask (corresponding to the aforementioned smear area). The obtained brush mask, the mask currently displayed on the screen (corresponding to the aforementioned area to be superimposed), and the current video frame (e.g., the aforementioned target image) are input to the algorithm module, and the segmentation algorithm in the algorithm module obtains the result (existing mask + newly added mask). Here, the above masks both represent masks, for example, by using pixel values 0 or 1 to mark corresponding areas. The existing mask corresponds to the area that needs to be extracted before the smearing operation, and the newly added mask corresponds to the object area specified by the user through the smearing operation. That is, the matting result of this frame image may simultaneously include the user's custom matting area and the area to be superimposed that needs to be extracted before the user customizes the matting. It should be noted that the above UI layer, SDK and algorithm module are all configured on a user terminal, and specifically, the above UI layer, SDK and algorithm module may all belong to a client configured on a user terminal.
[0061] In practical applications, the brush type, brush thickness, and brush trajectory data points (corresponding to the above-mentioned smear trajectory information) can be converted into a brush mask through an interface provided by the preset rendering module. Here, the rendering module can blend the video frame with the acquired mask, removing the background and leaving the main area corresponding to the mask (e.g., a person, a cat, etc.), thereby achieving a matting effect.
[0062] It can be understood that in an interaction segmentation task, a user's custom mask needs to be determined based on the brush mask, where the brush mask corresponds to brush data, and the user's custom mask corresponds to the object area specified by the user. The embodiments of the present disclosure provide a specific example of converting brush data into a matting result. As shown in FIG. 6 , feature sampling can be performed based on the user's brush input. Specifically, the brush outline can be calculated based on brush information such as brush thickness and brush trajectory data points. Then, outline point sampling can be performed, specifically using grid sampling, to describe the area to be extracted based on the sampled points. The sampled points and an existing mask (e.g., the area that needs to be extracted from the image before the current brush trajectory, corresponding to the area to be superimposed) can then be fed into a custom matting algorithm model to obtain the final mask. In practical application, the custom matting algorithm model can analyze the main object covering all the sampled points based on the sampled points and determine the final mask based on the existing mask. For example, the final mask can be obtained by performing a union operation between the custom mask corresponding to the brush and the existing mask. In some specific embodiments, if the area that the user wants to smear in the image with the current brush (corresponding to the custom mask) and the existing area to be extracted (corresponding to the existing mask) have a common body and the two areas are close, the union area corresponding to the custom mask and the existing mask may be determined as the common body area. For example, if a mask of a cat's head already exists and the user smears the cat's body with one stroke, the union is considered to be the entire cat. If the area that the user wants to smear in the image with the current brush and the existing area to be extracted do not have a common body, the union area corresponding to the custom mask and the existing mask is simply a combination of the two areas.
[0063] If the brush type is the first type (i.e., the first brush), the custom mask may be determined based on the brush mask, and the custom mask may be determined based on the thickness and stroke length of the user's brush. That is, the area of the main object to which the smeared area obtained by the user's smearing operation with the brush belongs is determined. For example, if the user uses a thin brush and the stroke length is short, the sampling points are relatively concentrated, and the main object to be extracted is expected to be small, for example, extracting only the head of a cat. If the user uses a thick brush and the stroke length is long, the sampling points are relatively sparse, and the main object to be extracted is expected to be large, for example, extracting the entire cat. The above is merely an illustrative description and should not be considered limiting.
[0064] By executing the interaction splitting task, an interaction splitting result can be obtained. The result may be represented by a mask only, for example, the interaction splitting result includes a user's custom mask. Then, a tracking task and / or a reverse tracking task may be executed. A specific description will be given below using the tracking task as an example.
[0065] The execution flow of the tracking task may refer to FIG. 7. First, preprocessing may be performed on the current frame. Preprocessing operations include, but are not limited to, decoding, resizing, format conversion, etc. For example, the size of the frame image in the video may be adjusted to a specified size, and the image format may be converted to a format processable by the algorithm module. After preprocessing, it can be determined whether the current frame is based on smart matting, i.e., whether the current frame requires smart matting. In other words, it can be determined whether the user has enabled the smart matting function and will perform custom matting. If the current frame is not based on smart matting (i.e., the user has completely custom matting), an empty smart matting mask may be constructed. This empty smart matting mask corresponds to the aforementioned empty area or empty data. For example, the size of the empty smart matting mask matches the size of the frame image, but all pixels are unified and all have a special value, such as all 1. If it is based on smart matting, it means that the user turns on the smart matting function and then performs custom matting. At this time, it can determine whether a smart matting mask exists for the current frame. If it is YES, it directly reads the smart matting mask, i.e., the smart matting mask is successfully read. If it is NO, it obtains the smart matting mask, for example, inputs the current frame into the algorithm used by the smart matting function (the preset matting algorithm described above), and obtains the smart matting mask through the preset matting algorithm. This smart matting mask can represent the area information of the area to be superimposed. The smart matting mask can be obtained in the above manner, and here, this smart matting mask may be obtained through the preset matting algorithm or may be empty data.The smart matting mask, the original image of the video (i.e., the current frame), and the interaction result mask corresponding to the first frame are input into the tracking algorithm, and the final mask for the current frame is obtained through the tracking algorithm. It should be noted that the interaction result mask corresponding to the first frame in FIG. 7 is the user's custom mask described above, which can represent the area information of the object area specified by the user. Furthermore, the first frame in FIG. 7 is not the first frame of the video, but the first frame of the tracking algorithm, such as the target image on which the user performs a smear operation. When performing the tracking algorithm on the current frame, it can first determine whether the current frame is the first frame. If yes, the interaction result mask corresponding to the first frame is read; if no, the interaction result mask corresponding to the first frame can be directly obtained. In actual application, if the current frame is the first frame, the input parameters of the tracking algorithm are the interaction result mask corresponding to the first frame, the current frame image, and the smart matting mask, so that the tracking algorithm performs target tracking for subsequent frame images based on parameters such as the interaction result mask corresponding to the first frame; if the current frame is not the first frame, the input parameters of the tracking algorithm may also be the current frame image and the smart matting mask.
[0066] The tracking algorithm obtains a final mask for the current frame (corresponding to the target area to be extracted as mentioned above) based on the smart matting mask and the interaction result mask corresponding to the first frame, and the final mask may include the smart matting mask and the user's custom mask. It also determines whether the current frame is the end frame, which may be the last frame of the video or a frame image set by the user. If YES, it ends; if NO, it sequentially obtains and processes the next frame image, and executes the flow shown in Figure 7 for each frame image until it ends.
[0067] The execution flow of the reverse tracking algorithm is similar to that of the tracking algorithm and will not be further described here.
[0068] In summary, the video processing method according to the embodiment of the present disclosure provides a way for a user to select an image matting area by themselves through a user terminal, and based on the smeared area determined by the user's smearing operation, a segmentation algorithm can be used to determine the target area (including the object area specified by the user) that the user needs to extract. That is, the user terminal can combine front-end interactions (e.g., the user's smearing operation) with background algorithms (e.g., segmentation algorithms, tracking algorithms, preset matting algorithms, etc.), thereby reliably realizing a user-customized matting function locally on the user terminal. Furthermore, the user can perform custom matting in addition to smart matting (i.e., the preset matting function), and can extract user-customized targets from the entire video based on the tracking algorithm / reverse tracking algorithm. In practical application, an SDK can be configured in the user terminal, and the above method logic can be executed by the SDK to efficiently and reliably realize the combination of front-end interactions with background algorithms and achieve the effects of custom matting and / or smart matting.
[0069] Furthermore, the embodiments of the present disclosure are not limited to extracting only specified types of objects in an image as in related technologies, and can fully meet users' image matting needs. They also eliminate the need to upload images to a server for processing, avoiding problems such as time consumption and communication volume consumption in interactions between the server and mobile terminals, ensuring information security, and effectively ensuring real-time processing, which helps improve matting efficiency and reduce matting costs.
[0070] Corresponding to the above-mentioned image processing method, Figure 8 is a structural schematic diagram of a video processing device according to an embodiment of the present disclosure, which may be realized by software and / or hardware, and may generally be integrated into an electronic device, which may be a user terminal such as a mobile phone, a computer, etc. As shown in Figure 8, a smear area acquisition module 802 for acquiring a smear area of a target image in a target video, the smear area being an area determined based on a smearing operation performed by a user on the target image; a first target area determination module 804 for inputting area information of the target image and a smear area corresponding to the target image into a segmentation algorithm to obtain a target area in the target image output by the segmentation algorithm, the target area including an object area designated by a user, the object area designated by the user being the area of a main object to which the smear area belongs or the smear area; The second target area determination module 806 includes: a second target area determination module 806 for tracking a first image in the target video based on a target area in the target image to determine a target area in the first image; and / or reverse-tracking a second image in the target video based on the target area in the target image to determine a target area in the second image, wherein the first image is a frame image that follows the target image in the target video, and the second image is a frame image that precedes the target image in the target video.
[0071] In the above technical solution according to the embodiments of the present disclosure, a user can flexibly specify the required target area (which may include the smear area or the area of the main object to which the smear area belongs) in the target image in the video by using the smearing method, and then can easily and quickly determine the target areas of other frame images in the video by using the tracking and / or reverse tracking method, and the target areas determined by the above methods can well meet the user's individual needs.
[0072] It should also be noted that in the related art, the matting function needs to be realized via a server, and users need to upload images to the server to realize the matting function, which not only raises certain concerns about information security but also requires interaction flows such as image transmission between the server and the system, which consumes time and communication traffic, has poor real-time capabilities, low matting efficiency, and high matting costs. The above method according to the embodiments of the present disclosure may be executed by a user terminal, and can realize a method in which the user selects the target area by the user terminal. Based on the smear area determined by the user's smearing operation, a segmentation algorithm is used to determine the target area that the user needs to extract (including the object area specified by the user). That is, the user terminal can combine front-end interaction (e.g., the user's smearing operation) with a background algorithm (e.g., the segmentation algorithm), thereby reliably realizing a matting function customized by the user locally on the user terminal. This not only fully meets the user's image matting needs, but also eliminates the need to upload images to a server for processing, thereby avoiding problems such as time consumption and communication traffic consumption in the interaction between the server and the mobile terminal, ensuring information security, and effectively ensuring real-time processing, which is helpful in improving matting efficiency and reducing matting costs.
[0073] In some embodiments, the device further includes a region-to-be-overlapped acquisition module for acquiring a region-to-be-overlapped corresponding to the target image, where the region-to-be-overlapped corresponding to the target image is a region that needs to be acquired from the target image before the smearing operation. The first target region determination module 804 is specifically used for inputting the target image, region information of the smeared region corresponding to the target image, and region information of the region-to-be-overlapped into a segmentation algorithm, where the target region further includes the region-to-be-overlapped.
[0074] In some embodiments, the area to be superimposed acquisition module is specifically used for: when a preset matting function is on, based on a preset matting algorithm, to acquire an area of a preset type object in the target image, and set the area of the preset type object as an area to be superimposed, where the preset type object is an object type included in the training sample of the preset matting algorithm; when the preset matting function is off or there is no preset type object in the target image, to determine the area to be superimposed corresponding to the target image as an empty area.
[0075] In some embodiments, the device further includes a segmentation algorithm module, which is used to determine a target area in the target image according to the steps of: determining an object area specified by a user based on area information of the smear area; and obtaining a target area in the target image based on area information of the object area corresponding to the target image and area information of an area to be superimposed corresponding to the target image.
[0076] In some embodiments, the smear area acquisition module 802 is specifically used to acquire the brush type of the virtual brush selected by the user, and if the brush type is a first type, determine the object area specified by the user as the area of the main object to which the smear area belongs, and if the brush type is a second type, determine the object area specified by the user as the smear area.
[0077] In some embodiments, the region information of the object region includes a mask of the object region, and the region information of the region to be superimposed includes a mask of the region to be superimposed, and the segmentation algorithm module is specifically used to combine the mask of the object region corresponding to the target image and the mask of the region to be superimposed corresponding to the target image, and determine the target region in the target image based on the combined mask.
[0078] In some embodiments, the smearing operation is an operation in which the user uses a virtual brush to smear on the target frame image, and the smearing area acquisition module 802 is specifically used to acquire the brush thickness of the virtual brush selected by the user and smearing trajectory information of the user smearing on the target frame image using the virtual brush, and to determine the smearing area of the target image based on the brush thickness and the smearing trajectory information.
[0079] In some embodiments, the second target area determination module 806 is specifically used to sequentially select a frame image in the target video that follows the target image from front to back as a first image, obtain an area to be superimposed corresponding to the first image, input area information of the first image and the area to be superimposed corresponding to the first image into a tracking algorithm, and determine a target area in the first image by the tracking algorithm based on the area information of the object area of the target image and the area information of the area to be superimposed corresponding to the first image.
[0080] In some embodiments, the second target area determination module 806 is specifically used to sequentially select frame images in front of the target image in the target video as second images from back to front, obtain the area to be superimposed corresponding to the second image, input area information of the second image and the area to be superimposed corresponding to the second image into a reverse tracking algorithm, and determine the target area in the second image by the reverse tracking algorithm based on the area information of the object area of the target image and the area information of the area to be superimposed corresponding to the second image.
[0081] In some embodiments, the target area corresponding to any frame image in the target video includes an area to be superimposed corresponding to this frame image and an object area corresponding to this frame image, and the object area corresponding to this frame image is an area in this frame image of the main object to which the smear area belongs.
[0082] In some embodiments, the region information of the region to be superimposed included in the target region corresponding to an arbitrary frame image in the target video is: If the preset matting function is on for any frame image in the target video, determine whether or not the area information of the area to be superimposed corresponding to this frame image has been successfully read; If not, input the frame image into a preset matting algorithm to obtain region information of a preset type object corresponding to the frame image through the preset matting algorithm, where the preset type object is an object type included in the training sample of the preset matting algorithm; The area information of the area to be superimposed corresponding to the read frame image or the area information of the preset type object corresponding to the frame image obtained by the preset matting algorithm is used as the area information of the area to be superimposed corresponding to the frame image, If the preset matting function is off or if there is no preset type object in this frame image, empty data is constructed and the empty data is used as area information for the area to be superimposed corresponding to this frame image. It is obtained based on the method.
[0083] In some embodiments, the target region corresponding to any frame image in the target video includes only the object region corresponding to this frame image, where the object region corresponding to this frame image is the region in this frame image of the main object to which the smear region belongs.
[0084] The device further includes a matting module, which, when the preset matting function is on, inputs each frame image in the target video into a preset matting algorithm to determine the area of a preset type object in each frame image using the preset matting algorithm, where the preset type object is an object type included in the training sample of the preset matting algorithm, and when performing the matting process, integrates the mask of the target area corresponding to each frame image with the mask of the area of the preset type object, and performs the matting process based on the integrated mask corresponding to each frame image.
[0085] A video processing device according to an embodiment of the present disclosure can execute a video processing method according to any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.
[0086] As can be clearly understood by those skilled in the art, for convenience and conciseness of description, the specific operating processes of the above-described apparatus embodiments may refer to the corresponding processes in the method embodiments, and will not be described here.
[0087] An embodiment of the present disclosure further provides an electronic device, the electronic device including a processor and a memory for storing instructions executable by the processor, the processor being used to read the executable instructions from the memory and execute the instructions to realize any one of the above video processing methods.
[0088] 9 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. As shown in FIG. 9, the electronic device 900 includes one or more processors 901 and a memory 902.
[0089] The processor 901 may be a central processing unit (CPU) or other type of processing unit having data processing and / or instruction execution capabilities, and may control other units in the electronic device 900 to perform desired functions.
[0090] The memory 902 may include one or more computer program products, which may include various types of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 901 may execute the program instructions to implement the video processing method according to the embodiments of the present disclosure described above and / or other desired functions. The computer-readable storage media may also store various contents, such as an input signal, a signal component, and a noise component.
[0091] In one example, the electronic device 900 may further include an input device 903 and an output device 904, which are connected to each other via a bus system and / or other type of connection mechanism (not shown).
[0092] The input device 903 may further include, for example, a keyboard, a mouse, and the like.
[0093] The output device 904 can output various information including determined distance information, direction information, etc. The output device 904 may include, for example, a display, a speaker, a printer, a communication network, and a remote output device connected thereto.
[0094] Of course, for the sake of simplicity, Fig. 9 shows only some of the units according to the present disclosure in the electronic device 900, omitting units such as a bus and an input / output interface, etc. Besides, the electronic device 900 may further include any other appropriate units according to specific application situations.
[0095] In addition to the above methods and apparatus, embodiments of the present disclosure may also be a computer program product, which includes computer program instructions that, when executed by a processor, cause the processor to perform a video processing method according to embodiments of the present disclosure.
[0096] The computer program product may organize program code for carrying out operations of embodiments of the present disclosure using any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., as well as conventional procedural programming languages such as "C" or similar. The program code may run entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or a server.
[0097] In addition, an embodiment of the present disclosure may be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, causes the processor to perform a video processing method according to an embodiment of the present disclosure.
[0098] The computer-readable storage medium may employ one or any combination of multiple computer-readable media. The computer-readable medium may be a readable signal medium or a readable storage medium. The computer-readable storage medium may include, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include an electrical connection having one or more conductors, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical memory device, a magnetic memory device, or any suitable combination of the above.
[0099] An embodiment of the present disclosure further provides a computer program product including a computer program / instruction, which, when executed by a processor, realizes the video processing method in the embodiment of the present disclosure.
[0100] It should be noted that, in this specification, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply that such an actual relationship or order exists between those entities or operations. Furthermore, the terms "comprises," "including," or any other variant thereof, are intended to cover a non-exclusive "comprise," whereby a process, method, article, or device comprising a set of elements not only includes those elements, but also other elements not expressly listed, or elements inherent in such process, method, article, or device. Absent further limitations, an element qualified by the phrase "comprises one of" does not exclude the presence of other identical elements in the process, method, article, or device that includes said element.
[0101] The above description is merely a specific embodiment of the present disclosure to enable those skilled in the art to understand or realize the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. 1. A video processing method comprising: Obtaining a smear area of a target image in a target video, the smear area being an area determined based on a smear operation on the target image by a user; inputting area information of the target image and a smear area corresponding to the target image into a segmentation algorithm to obtain a target area in the target image output by the segmentation algorithm, wherein the target area includes an object area designated by a user, and the object area designated by the user is an area of a main object to which the smear area belongs or the smear area; Tracking a first image in the target video based on a target area in the target image to determine a target area in the first image, and / or reverse-tracking a second image in the target video based on the target area in the target image to determine a target area in the second image, wherein the first image is a frame image that follows the target image in the target video, and the second image is a frame image that precedes the target image in the target video. A video processing method comprising:
2. The method further includes obtaining a region to be superimposed corresponding to the target image, wherein the region to be superimposed corresponding to the target image is a region that needs to be obtained from the target image before the smearing operation; The step of inputting area information of the target image and the smear area corresponding to the target image into a segmentation algorithm includes: The method of claim 1 , further comprising inputting the target image, area information of a smear area corresponding to the target image, and area information of an area to be superimposed into a segmentation algorithm, wherein the target area further includes the area to be superimposed.
3. The step of acquiring a region to be superimposed corresponding to the target image includes: When the preset matting function is on, obtain a region of a preset type object in the target image based on a preset matting algorithm, and set the region of the preset type object as a region to be superimposed, where the preset type object is an object type included in a training sample of the preset matting algorithm; If the preset matting function is off or there is no preset type object in the target image, the area to be superimposed corresponding to the target image is determined as an empty area. The method of claim 2 , comprising:
4. The division algorithm is determining an object region designated by a user based on region information of the smear region; A target area in the target image is obtained based on area information of an object area corresponding to the target image and area information of an area to be superimposed corresponding to the target image. The method of claim 2 , further comprising determining a target region in the target image according to steps.
5. The smearing operation is an operation in which a user smears on the target image using a virtual brush, and the step of determining an object area designated by the user based on area information of the smeared area includes: obtaining a user-selected brush type for the virtual brush; If the brush type is a first type, determining an object area designated by the user as an area of a main object to which the smear area belongs; If the brush type is a second type, determining the object area designated by the user as the smear area; The method of claim 4, comprising:
6. the region information of the object region includes a mask of the object region, and the region information of the region to be superimposed includes a mask of the region to be superimposed; The step of obtaining a target area in the target image based on area information of an object area corresponding to the target image and area information of an area to be superimposed corresponding to the target image includes: The method according to claim 4, further comprising: merging a mask of an object region corresponding to the target image with a mask of a region to be superimposed corresponding to the target image; and determining a target region in the target image based on the merged mask.
7. the smearing operation is an operation in which a user smears on the target image using a virtual brush, The step of acquiring a smear region of a target image in a target video includes: Acquiring information about the brush thickness of the virtual brush selected by the user and the smearing trajectory of the virtual brush used by the user to smear on the target image in the target video; determining a smear area of the target image based on the brush thickness and the smear trajectory information; The method of claim 1 , comprising:
8. The step of tracking a first image in the target video based on a target region in the target image to determine a target region in the first image includes: a frame image following the target image in the target video is set as a first image in order from front to back; obtaining a region to be superimposed corresponding to the first image; 2. The method of claim 1, further comprising inputting region information of the first image and the region to be superimposed corresponding to the first image into a tracking algorithm, and determining a target region in the first image by the tracking algorithm based on region information of the object region of the target image and region information of the region to be superimposed corresponding to the first image.
9. reverse tracking a second image in the target video based on a target region in the target image to determine a target region in the second image, frame images in front of the target image in the target video are sequentially set as second images in order from back to front; obtaining a region to be superimposed corresponding to the second image; inputting region information of the second image and the region to be superimposed corresponding to the second image into a reverse tracking algorithm, and determining a target region in the second image by the reverse tracking algorithm based on region information of the object region of the target image and region information of the region to be superimposed corresponding to the second image; The method of claim 1 , comprising:
10. The method of claim 1, wherein the target region corresponding to any frame image in the target video includes a region to be superimposed corresponding to this frame image and an object region corresponding to this frame image, and the object region corresponding to this frame image is a region in this frame image of a main object to which the smear region belongs.
11. The region information of the region to be superimposed included in the target region corresponding to an arbitrary frame image in the target video is If the preset matting function is on for any frame image in the target video, determine whether or not the area information of the area to be superimposed corresponding to this frame image has been successfully read; If not, input the frame image into a preset matting algorithm to obtain region information of a preset type object corresponding to the frame image through the preset matting algorithm, where the preset type object is an object type included in the training sample of the preset matting algorithm; The area information of the area to be superimposed corresponding to the read frame image or the area information of the preset type object corresponding to the frame image obtained by the preset matting algorithm is used as the area information of the area to be superimposed corresponding to the frame image, If the preset matting function is off or if there is no preset type object in this frame image, empty data is constructed, and the empty data is used as area information of the area to be superimposed corresponding to this frame image. The method of claim 10, wherein the method is obtained based on a formula.
12. a target region corresponding to an arbitrary frame image in the target video includes only an object region corresponding to this frame image, wherein the object region corresponding to this frame image is a region in this frame image of a main object to which the smear region belongs; When the preset matting function is on, input each frame image in the target video into a preset matting algorithm, and determine a region of a preset type object in each frame image by the preset matting algorithm, where the preset type object is an object type included in a training sample of the preset matting algorithm; When performing the matting process, the mask of the target area corresponding to each frame image and the mask of the area of the preset type object are integrated, and the matting process is performed based on the integrated mask corresponding to each frame image. The method of claim 1 further comprising:
13. a smear area acquisition module for acquiring a smear area of a target image in a target video, the smear area being an area determined based on a smearing operation performed by a user on the target image; a first target area determination module for inputting area information of the target image and a smear area corresponding to the target image into a segmentation algorithm to obtain a target area in the target image output by the segmentation algorithm, the target area including an object area designated by a user, the object area designated by the user being the area of a main object to which the smear area belongs or the smear area; a second target area determination module for tracking a first image in the target video based on a target area in the target image to determine a target area in the first image and / or reverse-tracking a second image in the target video based on the target area in the target image to determine a target area in the second image, wherein the first image is a frame image that follows the target image in the target video and the second image is a frame image that precedes the target image in the target video; a video processing device comprising:
14. An electronic device, the electronic device comprising: a processor; a memory for storing instructions executable by said processor; An electronic device, wherein the processor is adapted to read the executable instructions from the memory and execute the instructions to implement the video processing method of any one of claims 1 to 12.
15. 13. A non-transitory computer-readable storage medium having stored thereon a computer program, the computer program being used to execute the video processing method of any one of claims 1 to 12.
Citation Information
Patent Citations
Image processing method and device, equipment and medium
CN114222181A
Image matting method and device, storage medium and electronic equipment
CN114549547A
Object region extraction processing program, object region extraction device, and object region extraction method
JP2009080660A
Image tracking device and image tracking method
JP2012252601A
Image processing system, image processing method, and program storage medium
WO2015186341A1