A video processing method, device, terminal, medium and program product
By automatically identifying and removing watermarks and other elements in videos, the problem of watermarks affecting video quality is solved, achieving efficient video processing and improving user experience and processing efficiency.
Patent Information
- Application Number
- CN202111198349.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-14
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2041-10-14
AI Technical Summary
In existing technologies, watermarks and elements that do not meet display requirements in videos affect video quality and transmission frequency, resulting in a poor user experience. Furthermore, existing methods rely on manual annotation by users, which is inefficient.
A video processing method and apparatus are provided, which automatically identify and remove target visual elements in the video, such as watermarks, and trigger the removal process by gesture operation, audio signal, vibration operation or timing operation, and achieve automated removal through template matching and mask area processing.
It improves the efficiency of identifying and removing watermarks and other target visual elements in videos, enhances user experience, simplifies operation processes, and increases intelligence and flexibility.
Smart Images

Figure CN115988259B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to the field of image processing, and more particularly to a video processing method, a video processing device, a terminal, a computer-readable storage medium, and a computer program product. Background Technology
[0002] Videos have gained widespread popularity due to their advantages such as being intuitive, vivid, and convenient.
[0003] Practice has shown that videos often contain elements that affect video quality or hinder its dissemination; for example, watermarks may cause users to hesitate to share the video due to their presence. Therefore, improving the efficiency of watermark removal in videos has become a hot research topic. Summary of the Invention
[0004] This application provides a video processing method, apparatus, terminal, medium, and program product that can automatically identify and remove watermarks in videos, thereby improving the efficiency of watermark removal in videos.
[0005] On one hand, embodiments of this application provide a video processing method, the method comprising:
[0006] Acquire the video to be processed, which contains the target visual elements;
[0007] Eliminate target visual elements in the video;
[0008] The processed video is displayed on the playback interface, and the target visual elements in the processed video have been eliminated.
[0009] On the other hand, embodiments of this application provide a video processing apparatus, the apparatus comprising:
[0010] The acquisition unit is used to acquire the video to be processed, which contains the target visual elements.
[0011] The processing unit is used to eliminate target visual elements in the video;
[0012] The processing unit is also used to display the processed video on the playback interface, where the target visual elements in the processed video are eliminated.
[0013] In one implementation, the playback interface includes processing options; if a processing option is selected, a step to eliminate the target visual element in the video is triggered.
[0014] The target visual elements include watermarks, or the target visual elements include graphics and text that do not meet the display requirements.
[0015] In one implementation, if an elimination processing trigger operation exists, then the step of eliminating the target visual element in the video is triggered.
[0016] The elimination processing trigger operation includes any of the following: gesture operation, audio signal input operation, vibration operation, and timing operation; timing operation refers to the operation where the display duration of the video on the playback interface exceeds the duration threshold.
[0017] In one implementation, the processing unit is further used for:
[0018] When the elimination process for a target visual element in the video is triggered, pause video playback; and...
[0019] The playback interface displays a prompt message indicating that the target visual element in the video is being eliminated.
[0020] In one implementation, the playback interface includes an information prompt window, in which prompt information is displayed; the information prompt window includes a close control; the processing unit is further configured to:
[0021] During the process of eliminating target visual elements in the video, if the close control is selected, the information prompt window in the playback interface will be closed; and,
[0022] Interrupt the process of eliminating target visual elements in the video;
[0023] If the elimination process is interrupted, the target visual elements in the processed video will fall into any of the following categories: not eliminated, completely eliminated, or partially eliminated.
[0024] In one implementation, the processing unit is further used for:
[0025] When the elimination process is finished, a processing result notification is output. The processing result notification is used to notify the elimination result of the target visual element in the video. The elimination result includes any of the following: not eliminated, completely eliminated, or not completely eliminated. The end of the elimination process includes: the elimination process is completed; or the elimination process is interrupted.
[0026] In one implementation, the playback interface includes an element selection entry point; the processing unit is further used for:
[0027] In response to the element selection entry being triggered, the obscured area is displayed in the playback interface;
[0028] Move the occluded area and cover the reference visual elements in the processed video;
[0029] Play the overlay-processed video, in which reference visual elements are removed.
[0030] In one implementation, the playback interface includes a publishing control, a processing unit, and is further used for:
[0031] In response to the selection of the publish control, the processed video is published.
[0032] In one implementation, when the processing unit acquires the video to be processed, it is also used for:
[0033] Displays the service interface of the social application, which includes a video access point;
[0034] When the video acquisition entry is triggered, the video selection interface is displayed, which includes one or more candidate videos;
[0035] In response to a selection operation on one or more candidate videos, the selected candidate video is taken as the video to be processed.
[0036] In one implementation, the video to be processed is the video playing in the video browsing interface. When the processing unit obtains the video to be processed, it uses the following methods:
[0037] Displays a video browsing interface, which is used to display different videos;
[0038] In response to a selected video being chosen in the video browsing interface, the selected video is treated as the video to be processed.
[0039] In one implementation, when the processing unit performs elimination processing on target visual elements in the video, it specifically performs the following:
[0040] The video is processed by detection, segmentation and tracking to obtain at least one frame containing the target visual element, and the mask area corresponding to the target visual element in each frame.
[0041] The mask region in each frame of the image is filled to obtain at least one frame of the image after filling.
[0042] Based on the playback order of the frames in the video, at least one filled frame and at least one frame that does not contain the target visual element are fused together to generate the processed video.
[0043] In one implementation, the detection, segmentation, and tracking process includes element localization or element tracking; the video contains N frames, where N is a positive integer;
[0044] The processing unit is used to perform detection, segmentation, and tracking processing on each frame of the video to obtain at least one frame of the video containing the target visual element, and the mask region corresponding to the target visual element in each frame of the video. Specifically, it is used for:
[0045] For the i-th frame of the video, if there exists an i-1-th frame and the i-1-th frame contains a target visual element, then perform element tracking processing on the i-th frame based on the i-1-th frame to determine the target visual element contained in the i-th frame and the mask region corresponding to the target visual element in the i-th frame; i is a positive integer and 1≤i≤N;
[0046] If there is no i-1 frame image, or if there is an i-1 frame image but the i-1 frame image does not contain the target visual element, then perform element localization processing on the i-1 frame image to determine the target visual element contained in the i-1 frame image, and the mask area corresponding to the target visual element in the i-1 frame image.
[0047] In one implementation, the processing unit is used to perform element localization processing on the i-th frame image, determining the target visual elements contained in the i-th frame image, and the target visual elements in the mask region corresponding to the i-th frame image, specifically used for:
[0048] Obtain a template image containing the target visual elements;
[0049] The template image is used to perform traversal detection on the i-th frame;
[0050] If the traversal detection results indicate that the template image matches the i-th frame image, then the target visual element contained in the i-th frame image is determined based on the template image, and the mask region of the target visual element in the i-th frame image is determined.
[0051] In one implementation, when the processing unit performs traversal detection on the i-th frame image using the template image, it is specifically used for:
[0052] Align the reference vertices of the template image with the reference vertices of the i-th frame image;
[0053] The template image is moved multiple times on the i-th frame image according to the set direction and set step size;
[0054] Based on each movement, obtain the overlapping region in the i-th frame image that coincides with the template image, and calculate the similarity between the overlapping region and the template image;
[0055] Obtain the maximum similarity among multiple similarities corresponding to multiple moves. If the maximum similarity is greater than the similarity threshold, it indicates that the template image matches the i-th frame image.
[0056] In one implementation, the processing unit is used to determine the target visual element contained in the i-th frame image based on the template image, and to determine the mask region of the target visual element in the i-th frame image, specifically for:
[0057] Locate the overlapping region corresponding to the maximum similarity in the i-th frame image, and determine that the overlapping region corresponding to the maximum similarity contains the target visual element; and,
[0058] Based on the position of the target visual element in the overlapping region corresponding to the maximum similarity, the mask region of the target visual element in the i-th frame image is determined.
[0059] In one implementation, the processing unit is used to perform element tracking processing on the i-th frame image based on the (i-1)-th frame image, to determine the target visual elements contained in the i-th frame image, and the target visual elements in the mask region corresponding to the i-th frame image, specifically for:
[0060] Based on the mask region corresponding to the target visual element in the (i-1)th frame image, outline the similar region in the i-th frame image;
[0061] Calculate the similarity between the similar region and the mask region in the (i-1)th frame image;
[0062] If the similarity result meets the similarity condition, then the i-th frame image is determined to contain the target visual element, and the mask region in the (i-1)-th frame image is determined as the mask region corresponding to the target visual element contained in the i-th frame image.
[0063] In one implementation, if the video contains adjacent first and second images, and the mask region corresponding to the target visual element in the first image is filled as a first mask-filled region, and the mask region corresponding to the target visual element in the second image is filled as a second mask-filled region, then the processing unit is further configured to:
[0064] In the second image, a reference area with a display area larger than the mask area is drawn out, centered on the mask area corresponding to the target visual element.
[0065] Extract the target visual elements from the reference area to obtain the non-element region;
[0066] Calculate the blending coefficient between the non-elemental region and the first mask-filled region;
[0067] By using a fusion coefficient, the first mask filling region and the second mask filling region are fused to obtain the filling result of the mask region corresponding to the target visual element in the updated second image.
[0068] On the other hand, embodiments of this application provide a terminal, which includes:
[0069] A processor is used to load and execute computer programs;
[0070] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the video processing method described above.
[0071] On the other hand, embodiments of this application provide a computer-readable storage medium storing a computer program adapted to be loaded by a processor and executed by the above-described video processing method.
[0072] On the other hand, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. The terminal's processor reads the computer instructions from the computer-readable storage medium, and when executed by the processor, the computer instructions implement the aforementioned video processing method.
[0073] In this embodiment, target visual elements contained in a video can be automatically identified; and after the target visual elements are identified, the target visual elements contained in the video are eliminated so that the video after elimination does not contain the target visual elements. This method of automatically identifying and eliminating target visual elements in a video does not rely on the user's annotation of the target visual elements in the video, improves the user experience, enhances the intelligence and flexibility of the identification and elimination of target visual elements, and improves the efficiency of the elimination of target visual elements. Attached Figure Description
[0074] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0075] Figure 1 This illustration shows a scenario for removing watermarks from a video, provided by an exemplary embodiment of this application.
[0076] Figure 2 A flowchart illustrating a video processing method provided in an exemplary embodiment of this application is shown.
[0077] Figure 3a A schematic diagram of a playback interface provided in an exemplary embodiment of this application is shown;
[0078] Figure 3bThis illustration shows a schematic diagram of an exemplary embodiment of the present application, which illustrates how inputting a gesture operation in a playback terminal can trigger the elimination of target visual elements in a video.
[0079] Figure 4a This illustration shows a schematic diagram of outputting prompt information in a playback interface according to an exemplary embodiment of this application;
[0080] Figure 4b This illustration shows a schematic diagram of outputting prompt information in a playback interface according to an exemplary embodiment of this application;
[0081] Figure 4c This illustration shows a schematic diagram of an information prompt window output in a playback interface according to an exemplary embodiment of this application;
[0082] Figure 5 This illustration shows a schematic diagram of an exemplary embodiment of the present application, which allows a target user to perceive the processing progress of a watermark by updating the display state of the watermark.
[0083] Figure 6 This illustration shows a schematic diagram of outputting a processing result notification in a playback interface according to an exemplary embodiment of this application;
[0084] Figure 7 A flowchart illustrating a video processing method provided in an exemplary embodiment of this application is shown.
[0085] Figure 8 This illustration shows a schematic diagram of an exemplary embodiment of the present application for acquiring a video to be processed;
[0086] Figure 9 This illustration shows a schematic diagram of obtaining a video to be processed through a video browsing interface, provided by an exemplary embodiment of this application.
[0087] Figure 10 This illustration shows a schematic diagram of playing a video to be processed in a playback interface according to an exemplary embodiment of this application;
[0088] Figure 11a This illustration shows a schematic diagram of a reference visual element annotated by a target user, provided by an exemplary embodiment of this application.
[0089] Figure 11b This illustration shows a schematic diagram of an exemplary embodiment of the present application for acquiring a video to be processed;
[0090] Figure 12 This illustration shows a schematic diagram of a method for publishing a processed video according to an exemplary embodiment of this application;
[0091] Figure 13 A flowchart illustrating a video processing method provided in an exemplary embodiment of this application is shown.
[0092] Figure 14 This illustration shows a schematic diagram of a mask area corresponding to a watermark provided in an exemplary embodiment of this application;
[0093] Figure 15 A flowchart illustrating a video processing method provided in an exemplary embodiment of this application is shown.
[0094] Figure 16 This illustration shows a template matching method provided in an exemplary embodiment of this application;
[0095] Figure 17 This illustration shows a comparison diagram of the filling results of a mask region without distinguishing black boundary regions, and the filling results of a mask region with distinguishing black boundary regions, provided by an exemplary embodiment of this application.
[0096] Figure 18 This invention provides a schematic diagram of the structure of a video processing apparatus according to an exemplary embodiment of the present application.
[0097] Figure 19 A schematic diagram of the structure of a terminal provided in an exemplary embodiment of this application is shown. Detailed Implementation
[0098] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0099] This application proposes a video processing scheme. A video is composed of at least two sequentially linked frames; that is, an image is the smallest or most basic unit of a video. When playing a video, multiple frames are output continuously in the playback order. When the continuous image changes exceed 24 frames per second, based on the principle of visual persistence, the human eye obtains a smooth and continuous visual effect from each frame. Visual persistence refers to the ability of the human eye to retain the image for approximately 0.1-0.4 seconds after the image (such as the scene contained in a single frame) disappears when an object is moving rapidly (e.g., multiple frames of a video are played at a speed exceeding 24 frames per second). With the development and widespread use of the internet, videos are widely filmed and disseminated due to their intuitiveness, interactivity, and rich information content. For example, users can upload their videos to video platforms (such as applications or websites with video dissemination capabilities) and share them with other users; users can also browse, download, or forward videos uploaded by other users on the same platform; and so on.
[0100] However, practice has shown that videos circulating on the internet may contain elements that affect video quality, violate laws or ethical standards, or negatively impact the viewing experience. These elements can reduce the frequency of video dissemination and affect users' experience of downloading or sharing the video. For example, videos may contain targeted visual elements, which may include watermarks, or graphics that do not meet display requirements (such as graphics depicting knives or guns) and text (such as text expressing morally offensive meanings). Specifically: ① Watermarks in videos may be added by the video platform after the video is uploaded. When the video is downloaded or forwarded from the platform, it will carry the added watermark. Watermarks added by video platforms include, but are not limited to, the platform's icon (such as a logo) and / or the publisher's account identifier (such as an account ID or nickname). ② Elements in the video that do not meet display requirements can refer to elements that violate legal requirements, ethical standards, or negatively impact the viewing experience; for example, bloody scenes, knives, guns, etc., in the video. For ease of explanation, the following description will use the target visual element as the watermark as an example.
[0101] To alleviate users' concerns about sharing videos (such as the presence of watermarks from video publishers), increase the frequency (or number of times) of video sharing, and improve the quality of video dissemination, the video processing solution proposed in this application supports: automatically identifying target visual elements in the video and eliminating these elements, so that the processed video does not contain the target visual elements. This method of automatically identifying target visual elements in the video without relying on user annotation enhances the user experience, improves the intelligence and flexibility of target visual element identification and elimination, and increases the efficiency of target visual element elimination processing.
[0102] The video processing solution proposed in this application can be executed by a target terminal or by an application running on the target terminal. The target terminal can be any terminal, including but not limited to: smartphones (such as Android phones, iOS phones, etc.), tablets, personal computers, portable personal computers, mobile internet devices (MIDs), smart TVs, in-vehicle devices, head-mounted devices, and other touchscreen smart devices. This application does not limit the type of terminal; this is not stated herein. An exemplary scenario diagram illustrating the removal (including identification and removal) of watermarks in a video can be found [link to relevant documentation]. Figure 1 ,like Figure 1 As shown, assuming the target terminal is a personal computer 101, when a target user (such as a user using the target terminal) uploads a video through the target terminal, the target terminal can automatically remove the watermark from the video, thus removing the watermark from the processed video. Taking the removal of the watermark in the first frame of the video as an example, assuming a watermark 1021 is detected in the lower right corner of the first frame 102 of the video, and then the watermark 1021 is removed, the processed first frame image is obtained, and the watermark 1021 is removed from the processed first frame image. Based on the above description, it can be seen that the video processing scheme mentioned in the embodiments of this application can automatically identify and remove watermarks in the video, improve the intelligence and flexibility of watermark identification and removal, and improve the efficiency of watermark removal processing.
[0103] Based on the video processing scheme described above, this application proposes a more detailed video processing method. The video processing method proposed in this application will be described in detail below with reference to the accompanying drawings.
[0104] Figure 2 The illustration shows a flowchart of a video processing method provided in an exemplary embodiment of this application; the video processing method can be executed by a target terminal (such as any terminal), and the video processing method may include, but is not limited to, steps S201-S203:
[0105] S201: Obtain the video to be processed.
[0106] The video to be processed contains a target visual element, which can be a watermark as described above or other elements that do not meet display requirements (such as graphics, text, etc.). The following discussion will use a watermark as an example. Specifically, the acquired video to be processed can be played in the playback interface. Besides playing the video, the playback interface also supports processing the video. This processing includes, but is not limited to: matching templates to the video, editing the video, adjusting filters, adding text, and removing watermarks. The following is a brief introduction to some of the processing methods described above. For example, matching templates to the video can include scene templates (such as beach templates, desert templates, basin templates, etc.). When any scene template (such as a beach template) is added to the video, the video will display a beach scene when played. Editing the video refers to cutting and / or merging the video so that the edited video is different from the original video. Adding text to a video refers to adding text elements (such as text elements, animation elements, image elements, etc.) to one or more frames of images contained in the video. Removing watermarks from a video refers to identifying watermarks in the video and removing them, so that the processed video does not contain watermarks (or other elements that do not meet display requirements).
[0107] In its implementation, the playback interface includes an editing area and a playback area. The playback area is used to play the video; the editing area may contain one or more processing options for manipulating the video, such as template options for matching templates, editing options for trimming the video, filter options for adjusting filters, and removal options for removing watermarks from the video. When any processing option in the editing area is triggered, the video can be processed according to the function corresponding to that triggered option. A schematic diagram of an exemplary playback interface can be found [link to schematic diagram]. Figure 3a ,like Figure 3aAs shown, the playback interface 301 includes a playback area 3011 and an editing area 3012. The playback area 3011 plays a video, and the editing area 3012 displays one or more processing options for processing the video, such as template option 30121, editing option 30122, sticker option 30123, and image enhancement processing option 30124. During the playback of the video to be processed, in response to the selection of any processing option in the editing area 3012, the video can be processed according to the function corresponding to the triggered processing option; for example, during the playback of the video to be processed, in response to the selection of image enhancement processing option 30124, the step of eliminating target visual elements in the video is triggered, i.e., step S202 is executed.
[0108] Of course, in addition to triggering the elimination of target visual elements in the video to be processed, the image enhancement processing option 30124 can also be used to trigger other processing of the video to be processed, such as enhancing the image clarity of the video to be processed. In other words, the embodiments of this application support the fusion of the ability to eliminate target visual elements in the video with the ability to perform other processing on the video. In this way, after triggering the same option (such as the image enhancement processing option), multiple processing can be performed on the video to be processed at the same time. For example, after triggering the image enhancement processing option, the target visual elements in the video to be processed can be eliminated at the same time, and the image clarity of the video to be processed can be enhanced at the same time.
[0109] It is worth noting that the method of triggering the elimination of target visual elements in a video, in addition to the image enhancement options included in the playback interface described above, may also include: during the playback of the video to be processed, if an elimination trigger operation exists, then the step of eliminating the target visual elements in the video is triggered. The elimination trigger operation existing in the playback interface may include any of the following: gesture operation, audio signal input operation, vibration operation, and timing operation, etc. Gesture operation can refer to double-clicking, long-pressing, pinching with two fingers, swiping, etc., input by the target user on the target terminal's screen, or the aforementioned operations input by the target user through an external device (such as a mouse or keyboard). An exemplary process of triggering the elimination of target visual elements in a video by inputting a gesture operation on the playback terminal can be found in [link to relevant documentation]. Figure 3b ,like Figure 3bAs shown, assuming the gesture operation that triggers the elimination of the target visual element in the video is drawing an "L" shape in the playback interface, when the target user draws an "L" with one finger (or two fingers, three fingers, etc.) in the playback interface, a gesture operation is confirmed, and the elimination of the target visual element in the video is triggered. The so-called audio signal input operation refers to the operation where, with the target terminal's microphone on, the target terminal receives a voice signal input by the target user to trigger the elimination process. The so-called vibration operation refers to the operation generated by the target user shaking the target terminal; this embodiment does not limit the vibration frequency (or shaking frequency) of the target terminal. The so-called timing operation refers to the operation generated when the display duration of the video in the playback interface exceeds the duration threshold. The display duration includes the playback duration of the video when it is playing in the playback interface, or the display duration of the video (such as the first frame image contained in the video) when it is not playing in the playback interface. For example, if the duration threshold is 3 seconds, when the playback duration of the video in the playback interface reaches 3 seconds, it is determined that a timing operation is generated, and the elimination processing of the target visual element in the video is triggered. The specific number of the duration threshold is not limited in the embodiments of this application.
[0110] The above describes how the elimination of target visual elements in a video is triggered by executing an elimination processing trigger operation or triggering a screen enhancement processing option during the playback of the video to be processed. However, it is understood that this application embodiment also supports triggering the elimination of target visual elements in the video when the video in the playback interface is paused. That is to say, regardless of whether the video to be processed is playing or paused, this application embodiment supports triggering the elimination of target visual elements in the video. For ease of explanation, the following description will take triggering the elimination of target visual elements in the video during the playback of the video to be processed as an example.
[0111] In summary, during the playback of the video to be processed, if there is an option to enhance the image or an operation to trigger the elimination process (such as a timing operation) in the playback interface, the target terminal can be triggered to automatically identify and eliminate the target visual elements in the video. Compared with relying on the target user to mark the target visual elements in the image frames of the video, this simplifies the operation of triggering the elimination process, improves the user experience, and increases the efficiency of eliminating the target visual elements.
[0112] S202: Eliminate target visual elements in the video.
[0113] S203: Display the processed video in the playback interface.
[0114] In steps S202-S203, during the playback of the video to be processed in the playback interface, when the elimination processing of target visual elements in the video is triggered, the target terminal begins to perform the identification and elimination of target visual elements in the video, so that the processed video does not contain target visual elements; and the processed video will also be displayed in the playback interface. Displaying the processed video in the playback interface may include: displaying the video identifier of the processed video (such as the first frame image contained in the processed video, a thumbnail of the first frame image, or a video account, etc.); or playing the processed video in the playback interface. During the process of the target terminal eliminating target visual elements in the video, changes may occur in the playback interface, thereby making the target user aware that the target terminal is eliminating target visual elements in the video. The following is an exemplary description of the interface changes (or changes in elements within the interface) that may occur in the playback interface during the process of the target terminal eliminating target visual elements in the video:
[0115] 1) Pause video playback and output a prompt message to inform the target user that a target visual element in the video is being eliminated. In a specific implementation, when elimination of a target visual element in the video is triggered, video playback is paused on the playback interface; and a prompt message is output on the playback interface. For example, assuming the video contains 100 frames, when the elimination of a target visual element in the video is triggered at frame 20, video playback is paused, and frame 20 is continuously displayed on the playback interface; and a prompt message is output as a floating layer above frame 20. An exemplary interface diagram illustrating the output of a prompt message on the playback interface can be found here. Figure 4a ,like Figure 4a As shown, when the elimination process for target visual elements in the video is triggered, a prompt message 401 is output in the playback interface 301. The prompt message 401 is used to indicate that the elimination process for target visual elements in the video is in progress.
[0116] It should be noted that the specific content of the prompts in the playback interface can change in real time according to the progress of the target terminal in eliminating the target visual elements in the video; see [link / reference]. Figure 4bIf the target terminal is recognizing a watermark (i.e., a target visual element) in the video, the specific content of the prompt message 401 output in the playback interface can be similar to "Recognizing"; if the target terminal recognizes the watermark in the video, the specific content of the prompt message 401 can be similar to "Watermark detected"; if the target terminal is removing the watermark in the video, the specific content of the prompt message 401 can be similar to "Removing watermark". This method of outputting prompt messages with different meanings based on the target terminal's processing progress of the target visual element in the video can inform the target user of the current processing progress in real time, improving the target user's experience. In addition to outputting prompt message 401 in the playback interface, this application embodiment also supports outputting other content in the playback interface at the same time to indicate the target terminal's current progress in removing the target visual element in the video, such as outputting a processing progress bar and / or animation; such as Figure 4c As shown, an information prompt window 402 is displayed in the playback interface. This information prompt window includes prompt information 401, a processing progress bar 403, and an animation 404. The display form of the processing progress bar 403 and / or the animation 404 can change to provide real-time prompts to the target user on the current processing progress of the target visual element. This application embodiment does not limit the type and quantity of content included in the information prompt window 402. For example, the information prompt window may include one or more of the processing progress bar 403 and the animation 404.
[0117] 2) Without pausing video playback, and based on the target terminal's processing progress of the target visual elements in the video, the display status of the target visual elements in the playback interface is updated in real time. This allows the target user to perceive that the target terminal is processing the removal of the target visual elements in the video. In other words, when the removal of the target visual elements in the video is triggered, the video continues to play in the playback interface, and during the process of the target terminal removing the watermark from the video, the display status of the watermark in the playback interface is updated in real time, thereby allowing the target user to perceive the progress of the watermark processing based on the watermark display status.
[0118] For example, when the target terminal removes the watermark at the first moment, the displayed area of the watermark on the playback interface is greater than when the target terminal removes the watermark at the second moment; wherein, the first moment is earlier than the second moment. An exemplary process of allowing the target user to perceive the progress of watermark processing by updating the watermark's display state can be found in [reference needed]. Figure 5 ,like Figure 5As shown, assuming the target terminal removes the watermark from the video at the first moment and the second moment respectively, and the first moment is earlier than the second moment, then when the target terminal removes the watermark from the target video at the first moment, the display area of the watermark on the playback interface is greater than the display area of the watermark on the playback interface when the target terminal removes the watermark from the target video at the first moment.
[0119] Furthermore, this application embodiment also supports interrupting the removal process of target visual elements in the video during the watermark removal process on the target terminal. The interruption of the removal process can be performed by the target user or by the target terminal (e.g., due to network disconnection or power outage). The following example illustrates how the removal process of target visual elements in the video can be interrupted by the target user, using the scenario of an information prompt window displayed in the playback interface as an example. In the specific implementation, the playback interface includes an information prompt window, and the prompt information is displayed in the information prompt window; the information prompt window also includes a close control (e.g., ...). Figure 4c This includes a close control (405); if the close control is selected during the removal process of the target visual element in the video, the information prompt window is closed in the playback interface; and the removal process of the target visual element in the video is interrupted; wherein, if the removal process is interrupted, the target visual element in the processed video may fall into any of the following categories: not removed, completely removed, or only partially removed. In short, throughout the entire process of the target terminal removing the watermark from the video, the target user can interrupt the watermark removal process by triggering the close control included in the information prompt window; depending on when the close control in the information prompt window is triggered, the progress of watermark removal at the time the close control is triggered will also be different, such as the watermark not being removed, completely removed, or partially removed at the time the close control is triggered.
[0120] It should be understood that during the process of removing target visual elements in a video by the target terminal, the operation of interrupting the removal process is not limited to the methods described above. For example, when the removal of target visual elements in a video is triggered by triggering the image enhancement option in the playback interface, the removal of target visual elements in the video can also be interrupted when the image enhancement option is triggered again. This application embodiment does not limit the specific implementation method of interrupting the removal process. In summary, this application embodiment supports the target user to trigger the close control to interrupt the removal of target visual elements in the target video, simplifying the operation of interrupting (or canceling) the removal process, improving the target user's experience with watermark removal, and thus increasing user stickiness.
[0121] This application embodiment also supports outputting a processing result notification to inform the user of the elimination result of the target visual element in the video when the elimination process is completed. Specifically, when the elimination process is completed, a processing result notification is output to inform the user of the elimination result of the target visual element in the video; wherein, the elimination result may include, but is not limited to, any of the following: not eliminated, completely eliminated, or not completely eliminated. The completion of the elimination process mentioned above may include: completing the elimination process, or the elimination process being interrupted; that is, whether the target terminal completes the elimination process of the target visual element in the video, or the target terminal is interrupted during the elimination process, it can be understood as the elimination process being completed; this application embodiment does not limit which of the above-mentioned types of elimination process completion is specified. An exemplary schematic diagram of outputting a processing result notification in the playback interface can be found in [reference needed]. Figure 6 When the removal process is complete, a processing result notification 601 can be output in the middle of the playback interface; or, when the removal process is complete, a processing result notification can be displayed in the associated area of the playback interface where the target visual element (such as a watermark) was originally displayed. This application embodiment does not limit the specific display position of the processing result notification in the playback interface, which is only described here.
[0122] Based on the above description, it can be seen that, depending on the implementation method of ending the elimination process, the presence of target visual elements in the processed video (or the display state as described above) is not the same when playing the processed video. For example, if ending the elimination process means completing the elimination process, then when playing the processed video, the target visual elements in the processed video are eliminated (i.e., completely eliminated). As another example, if ending the elimination process means the elimination process is interrupted, then when playing the processed video, the target visual elements in the processed video may be in any of the following situations: not eliminated, completely eliminated, or partially eliminated. The specific implementation process of playing the processed video in this application embodiment is not described in detail.
[0123] In this embodiment, target visual elements contained in a video can be automatically identified; and after the target visual elements are identified, the target visual elements contained in the video are eliminated so that the video after elimination does not contain the target visual elements. This method of automatically identifying and eliminating target visual elements in a video does not rely on the user's annotation of the target visual elements in the video, improves the user experience, enhances the intelligence and flexibility of the identification and elimination of target visual elements, and improves the efficiency of the elimination of target visual elements.
[0124] Figure 7The illustration shows a flowchart of a video processing method provided in an exemplary embodiment of this application; the video processing method can be executed by a target terminal (such as any terminal), and the video processing method may include, but is not limited to, steps S701-S704:
[0125] S701: Obtain the video to be processed.
[0126] As described above, the video processing method proposed in this application embodiment can be executed by an application running on the target terminal. An application, often simply referred to as an application, is a computer program designed to perform one or more specific tasks. Applications can be categorized by their running method, including but not limited to: clients installed on the terminal; or, installation-free applications, i.e., applications that can be used without downloading and installation, commonly known as mini-programs; web applications opened through a browser; and so on. Applications can also be categorized by their functional type, including but not limited to: IM (Instant Messaging) applications, which can include but are not limited to: social applications, map applications with social interaction functions, game applications, etc.; and content interaction applications, such as online banking, personal spaces, news, etc. It should be noted that the above are only two exemplary dimensions for classifying applications, and the types of applications under each dimension are exemplary. For ease of explanation, this application embodiment uses a social application as an example to describe the subsequent related content; in other words, a social application has the function of eliminating target visual elements in a video.
[0127] Below are several ways to obtain videos for processing through social applications:
[0128] (1) Obtaining the video to be processed through the service interface of the social application; in this implementation, the method of obtaining the video to be processed through the social application may include: displaying the service interface of the social application, which contains a video acquisition entry; when the video acquisition entry is triggered, displaying a video selection interface, which includes one or more candidate videos; in response to the selection operation of one or more candidate videos, the selected candidate video is used as the video to be processed. Optionally, the service interface of the social application may include a social dynamic interface, which may refer to an interface including a social information stream, which may refer to a feed stream, a kind of information stream that continuously updates and displays social information to users. Optionally, the service interface of the social application may include a service selection interface, which may include service identifiers of various services provided by the social application. When the service identifier of any service is triggered in the service selection interface, the service interface corresponding to the triggered service may be displayed.
[0129] Combination Figure 8 Using the service selection interface of a social application as an example, the process of acquiring videos to be processed is introduced, such as... Figure 8 As shown, a service interface 801 of a social application is displayed, which includes a video acquisition entry 8011. When the video acquisition entry 8011 is triggered, a video selection interface 802 is displayed, which contains one or more candidate videos, such as candidate video 1, candidate video 2, candidate video 3, and so on. In response to the selection operation of one or more candidate videos, the selected candidate video is used as the video to be processed. For example, if candidate video 1 in the video selection interface 802 is selected, then candidate video 1 is used as the video to be processed. The video identifier of the selected candidate video (such as the first frame of the video) can be displayed in the selection area 803 of the video selection interface 802 to facilitate the target user to edit the selected candidate video (such as deleting, adding, etc.). For example, the selection area 803 contains the selected candidate video 1, and the display area of candidate video 1 contains a close option 8031. When the close option 8031 is selected, it means that the target user does not want candidate video 1 to be used as the video to be processed, so candidate video 1 is deleted in the selection area 803. In addition, the selection area 8031 also includes a next step option 8032. If the next step option 8032 is triggered (such as by clicking), it means that the target user has completed the selection of the video to be processed in the video selection interface 802, and the elimination process of the video to be processed can be triggered.
[0130] It should be noted that: ① The elimination processing of target visual elements in a video proposed in this application embodiment is essentially the elimination processing of target visual elements in each frame of the video; based on this, this application embodiment also supports the elimination processing of target visual elements in a single frame image; therefore, in addition to candidate videos, the video selection interface 802 can also contain candidate images, such as candidate image 1, candidate image 2, ...; when any candidate image is triggered, it indicates that the target user wants to eliminate the target visual elements contained in the triggered candidate image. ② This application embodiment also supports the simultaneous selection of multiple candidate videos, that is, multiple candidate videos as videos to be processed. In this way, the target terminal can simultaneously or sequentially eliminate the target visual elements in the selected multiple candidate videos to obtain the candidate video after elimination processing for each candidate video; this method of selecting multiple candidate videos as videos to be processed at one time makes it convenient for the target user to select multiple videos to be processed at one time, simplifies the number of times to select videos to be processed, and thus improves the efficiency of eliminating multiple videos to be processed. ③ To facilitate the selection of candidate images or videos by target users, this embodiment also supports displaying multimedia type options in the video selection interface 802, such as type option 8021, type option 8022, and type option 8023. When type option 8021 is selected, candidate videos and candidate images are displayed in the video acquisition interface 802; when type option 8022 is selected, candidate videos are displayed in the video acquisition interface 802; and when type option 8023 is selected, candidate images are displayed in the video acquisition interface 802. By displaying candidate images and candidate videos separately, it is beneficial for target users to select candidate images or candidate videos in categories, thereby improving the convenience of selecting candidate images or candidate videos. ④ The candidate videos and / or candidate images included in the video selection interface 802 can be offline data stored in the local storage space of the target terminal, such as videos shot or downloaded by the target user; or, the candidate videos and / or candidate images can also be online data stored in a database on a server. This embodiment does not limit the source of candidate videos and / or candidate images.
[0131] (2) Obtain the video to be processed through a video browsing interface; wherein, the video browsing interface can be an interface for playing videos, and when a switching operation is performed in the video browsing interface, different videos can be switched to play. The switching operation may include, but is not limited to: the operation generated when swiping along any direction of the terminal screen; for example: the operation generated when swiping up along the terminal screen, at which time the next video of the current video is played in the video browsing interface; or the operation generated when swiping down along the terminal screen, at which time the previous video of the current video (i.e., the video played in the historical time period) is played in the video browsing interface. In short, the video to be processed can be the video that is playing in the video browsing interface. In this implementation, the implementation method of obtaining the video to be processed through the video browsing interface may include: displaying the video browsing interface, which is used to display different videos; in response to the video being selected in the video browsing interface, the selected video is used as the video to be processed.
[0132] The methods for selecting a video playing in the video browsing interface may include, but are not limited to: selecting a video by performing gesture operations in the video browsing interface (such as performing double-click, swipe, long-press, or drawing a specified shape such as "S"); selecting a video by inputting an audio signal (such as the target terminal acquiring the audio signal of the selected video as the video to be processed through a microphone); and selecting a video when a confirmation option is selected in the video browsing interface (such as when the video browsing interface includes a confirmation option, and when the confirmation option is selected, the currently playing video is determined to be the video to be processed). It should be noted that the above are only some exemplary methods for selecting a video in the video browsing interface. In actual application scenarios, the method for selecting the video to be processed may vary, and this application embodiment does not limit it. The following is combined with Figure 9 This section briefly introduces the process of acquiring a video to be processed through a video browsing interface, using the example of selecting a video by performing a gesture operation in the video browsing interface. (See [link to relevant documentation]). Figure 9 Assuming that video 1 is playing in the video browsing interface 901, and the specified gesture operation is: draw an "S" shape with two fingers in the video browsing interface; then when the target user draws an "S" shape with two fingers in the video browsing interface, it means that the target user wants to eliminate the target visual element in the currently playing video, and the currently playing video is determined as the video to be processed.
[0133] It should be understood that the aforementioned video browsing interface can be provided by a video application with video playback functionality, or it can be provided by the aforementioned social application; this application embodiment does not limit the specific type of application to which the video browsing interface belongs. When the video browsing interface is provided by a social application, the process of triggering the display of the video browsing interface in the social application may include: displaying the service interface of the social application, which contains a video acquisition entry; and displaying the video browsing interface when the video acquisition entry is triggered. Further, the method of determining the video to be processed through the video browsing interface may include: using the process described above to select the video currently playing in the video browsing interface as the video to be processed; or, the video browsing interface may also include selection options (see...). Figure 9 Option 902); when the selection option is triggered, a video selection interface is displayed, which includes one or more candidate videos; thus, the video to be processed is determined from the video selection interface in the aforementioned manner, and after the video to be processed is determined, it is played in the playback interface. The interface flow described above, which triggers the display of the video browsing interface from the service interface of the social application, triggers the display of the video selection interface from the video browsing interface, and selects the video to be processed from the video selection interface and plays the video to be processed in the playback interface, can be found in [reference missing]. Figure 10 This will not be discussed in detail here.
[0134] It is worth noting that the above are only two exemplary implementations of obtaining videos to be processed; in actual application scenarios, the implementation of obtaining videos to be processed can vary. For example, this application embodiment also supports using videos sent by any participant in a social conversation as videos to be processed within the conversation interface provided by a social application. The specific process of this implementation can be simply included as follows: displaying the conversation interface of the social conversation, which contains conversation video information sent by the participants in the social conversation, and the conversation video information contains the conversation video; if there is a selection operation for the conversation video information, then the conversation video is used as the video to be processed. In this way, the target user can conveniently select the video in the conversation as the video to be processed during the participation in the social conversation, shortening the path of selecting the video to be processed, thereby improving the speed and efficiency of obtaining the video to be processed.
[0135] It should be noted that other contents included in step S701 can be found in the foregoing. Figure 2 The specific implementation process of step S201 in the illustrated embodiment will not be repeated here.
[0136] S702: Eliminate target visual elements in the video.
[0137] S703: Displays the processed video in the playback interface.
[0138] It should be noted that the specific implementation process of steps S703-S704 can be found in the aforementioned... Figure 2 The specific implementation process of steps S202-S203 in the illustrated embodiment will not be repeated here.
[0139] S704: Overlay the reference visual elements contained in the processed video to obtain an overlaid video in which the reference visual elements are eliminated.
[0140] It is understood that this application embodiment also supports the target user marking elements that need to be eliminated in the images contained in the video to be processed, so that the target terminal (or the application running on the target terminal (such as a social application)) can eliminate the elements marked by the target user. In one possible scenario, if the processed video obtained after the target terminal automatically identifies and eliminates the target visual elements contained in the video still contains the elements that the target user wants to eliminate, the target user can manually mark the elements to be eliminated in the images, so that the target terminal can eliminate the elements marked by the target user. Here, to facilitate the distinction from the aforementioned target visual elements, this application embodiment refers to the elements that the target user wants to eliminate as reference visual elements, which will be explained here.
[0141] In the specific implementation, in response to the element selection entry being triggered, the occluded area is displayed in the playback interface; the occluded area is moved and covers the reference visual element in the processed video to eliminate the reference visual element; the video after coverage is played, and the reference visual element in the covered video is eliminated. An exemplary interface diagram showing the reference visual element annotated by the target user is shown below. Figure 11a As shown, as described above, the playback interface includes processing options, such as sticker option 30123. The element selection entry is set in sticker option 30123. When sticker option 30123 is triggered, it is determined that the element selection entry is triggered. The occlusion area 1101 is displayed in the playback interface. When the occlusion area 1101 is moved and covers the reference visual element in the processed video, it is determined that the reference visual element is eliminated to obtain the covered video. In the covered video, the reference visual element is eliminated, that is, the reference visual element is integrated with the background of the image, and the reference visual element cannot be seen from the image by the human eye.
[0142] It's easy to understand that, besides being directly set within the processing options included in the playback interface, the element selection entry point can also be indirectly set within the processing options; see [link to relevant documentation]. Figure 11bAfter triggering the sticker option 30123 in the playback interface, an option window 1102 is displayed. This option window 1102 contains at least one option associated with the sticker option, and at least one option includes a smart elimination option 1103. The element setting entry can be set in the smart elimination option 1103, that is, when the smart elimination option 1103 is triggered, the element setting entry is triggered. This application embodiment does not limit the specific setting location and display method of the element setting entry, which is not described here. In addition, besides supporting the target user to annotate reference visual elements during the playback of the processed video, this application embodiment also supports the target user to annotate reference visual elements during the playback of the video to be processed. This allows the target terminal to eliminate reference visual elements simultaneously during the identification and elimination of target visual elements, improving the efficiency of eliminating both target and reference visual elements in the video.
[0143] In summary, the embodiments of this application not only support the automatic identification and elimination of target visual elements contained in the video by the target terminal, but also support the elimination of reference visual elements marked in the video by the target user according to their own needs. This improves the efficiency of eliminating target visual elements while meeting the target user's needs for eliminating reference visual elements, thereby improving the quality of the elimination process.
[0144] In addition, this application embodiment also supports publishing and / or sharing the processed video; publishing refers to publishing the processed video to the social dynamics interface associated with the target user's user account; sharing refers to sharing the processed video to other users' accounts, such as sharing the processed video to other users' accounts in the form of links, identification codes (such as QR codes), etc. The aforementioned accounts or user accounts can refer to the accounts used by users when logging into social applications, which will not be described in detail here. Taking publishing the processed video as an example, the specific process of publishing the processed video may include: including a publishing control (or option, component, button, etc.) in the playback interface; responding to the selection operation of the publishing control, publishing the processed video. An exemplary interface diagram for publishing a processed video can be found here. Figure 12As shown in Figure 12, the playback interface 301 includes a publishing control 1201. In response to a selection operation on the publishing control 1201, the processed video is published in the social dynamics interface 1202. The social dynamics interface 1202 includes not only the social information of the processed video but also social information published by other users. Hidden social information can be displayed by swiping within the social dynamics interface 1202. Of course, after triggering the publishing control 1201, the target user can also edit the processed video, such as adding descriptive text or setting which users have permission to view the processed video. The processed video is only published to the social dynamics interface after the target user has finished editing. The above is an exemplary implementation of publishing a processed video. This application embodiment does not limit the specific implementation process of publishing a processed video, but it is described here. This method of directly triggering the publishing process of the processed video in the playback interface improves the speed of publishing the processed video and is beneficial for target users to quickly publish the processed video.
[0145] The video processing method provided in this application can be executed by a target terminal or various applications running on the target terminal. For example, it can be executed by a social application. This allows target users to directly select the video to be processed from the social application while participating in social conversations or using the social application, and eliminate target visual elements in the video to obtain the processed video. This shortens the path to obtaining the video to be processed and simplifies the operations of obtaining the video and eliminating target visual elements. Furthermore, this application also supports automatic identification and elimination of target visual elements in the video after it has been obtained, without relying on user annotation of target visual elements, thus improving user experience, enhancing the intelligence and flexibility of target visual element identification and elimination, and increasing the efficiency of target visual element elimination. Additionally, this application also supports target users annotating the video to be processed with reference visual elements they want to eliminate, meeting their personalized needs for element elimination and increasing user engagement.
[0146] The above Figure 2 and Figure 7 The relevant content of the illustrated embodiment mainly describes the interface diagram of the video processing method proposed in this application embodiment. The following is a detailed description in conjunction with... Figure 13 The flowchart shown is used to briefly introduce the technical processing flow of the embodiments of this application. For example... Figure 13 The flowchart of the video processing method shown mainly involves three functional parts, among which:
[0147] The first point is: frame extraction detection to determine whether the video to be processed contains the target visual element. Specifically, after the target terminal acquires the video to be processed, it extracts at least two frames from the video (e.g., extracting a portion of images at equal intervals in chronological order), and checks whether at least one of these two frames contains the target visual element. If none of the frames in the two images contain the target visual element, it means the video to be processed does not contain the target visual element, and therefore no elimination processing is needed; the original video to be processed is then directly output. Conversely, if one or more frames in the two images contain the target visual element, it means the video to be processed does contain the target visual element, and therefore elimination processing is required. This is to reduce the processing time for videos without the target visual element; that is, when it is detected that the video to be processed does not contain the target visual element, the video to be processed can be returned directly, avoiding further processing and thus shortening the processing time.
[0148] The second point is: traverse each frame of the video to be processed, and perform detection, segmentation, and tracking processing on each video frame (such as the images contained in the video to be processed) to obtain the mask region corresponding to the target visual element in the video to be processed. Here, the mask region corresponding to the target visual element can be called an element mask; when the target visual element is a watermark, this element mask can be called a watermark mask. The area occupied by the watermark mask corresponding to the watermark refers to the area composed of pixels used to present the watermark, and the shape of this area is consistent with the shape of the watermark. An exemplary watermark mask region can be found in [reference needed]. Figure 14 ,like Figure 14 As shown, the image contains the watermark "@61359". The mask area corresponding to the watermark "@61359" is the region composed of the pixels corresponding to the characters "@", "6", "1", "3", "5", and "9". The detection, segmentation, and tracking can be simply understood as follows: For a single frame of the video to be processed, element template matching is used to obtain the type and display position of the target visual element in that single frame. Then, tracking logic is added to the multiple frames of the video to be processed. This tracking logic means that if the difference between the previous frame and the current frame is small, or if the mask area of the target visual element in the previous frame has a certain similarity to the corresponding area in the current frame, then the target visual element and its mask area in the previous frame can be used as the target visual element and mask area in the current frame. This tracking logic shortens the recognition time of target visual elements in multiple frames and ensures the stability of the mask areas between different images, making subsequent filling processing less noticeable.
[0149] The third point is: after determining the mask region of the target visual element in the video to be processed based on the second point, the target visual element can be removed. Furthermore, the mask region of the target visual element is filled, so that the filled mask region blends seamlessly with the surrounding background in the image, making the processed video appear as if the target visual element has been removed without any trace. Specifically, the filling of the mask region in a single frame image in this application embodiment includes: using the feature information of similar blocks in the image that are similar to the mask region, and other areas within the region associated with the target visual element excluding the mask region, to fill the mask region; and adding a certain degree of blurring so that the filling effect of the mask region is not significantly different from the pixels of the surrounding area, thus making the filled mask region look more natural. For filling the mask region in multiple frames, this application embodiment also supports a smooth transition between the current frame image and the previous frame image based on the similarity of the background between the previous frame image and the current frame image, making the mask regions in consecutive frames of the processed video look more coherent and natural, thereby improving the quality of the processed video.
[0150] The following is combined with Figure 15 The more detailed flowchart shown below introduces the technical processing flow of the video processing method; Figure 15 The illustration shows a flowchart of a video processing method provided in an exemplary embodiment of this application; the video processing method can be executed by a target terminal (such as any terminal), and the video processing method may include, but is not limited to, steps S1501-S1505:
[0151] S1501: Obtain the video to be processed.
[0152] It should be noted that the specific implementation method shown in step S1501 can be found in the aforementioned... Figure 7 The specific implementation details of step S702 shown are not elaborated here.
[0153] S1502: Perform detection, segmentation and tracking processing on each frame of the video to obtain at least one frame of the video containing the target visual element, and the mask area corresponding to the target visual element in each frame.
[0154] Considering that the video to be processed contains multiple frames, this application embodiment divides the element localization process into two steps for stability and performance considerations between different frames: element localization processing for a single frame and element tracking processing between multiple frames. Specifically, assuming the video to be processed contains N frames, where N is a positive integer, the implementation of detection, segmentation, and tracking processing for each frame in the video can include: For the i-th frame in the video, if there exists an (i-1)-th frame and the (i-1)-th frame contains a target visual element, then element tracking processing is performed on the i-th frame based on the (i-1)-th frame to determine the target visual element contained in the i-th frame and the mask region corresponding to the target visual element in the i-th frame; i is a positive integer and 1≤i≤N. If there is no (i-1)-th frame, or if there is an (i-1)-th frame but the (i-1)-th frame does not contain a target visual element, then element localization processing is performed on the i-th frame to determine the target visual element contained in the i-th frame and the mask region corresponding to the target visual element in the i-th frame.
[0155] The following sections will elaborate on the element localization processing for a single frame image and the element tracking processing between multiple frames image, as mentioned above.
[0156] (1) Perform element localization processing on a single frame image, that is, identify (or detect) the target visual elements in the single frame image and the mask area corresponding to the target visual elements.
[0157] The general approach may include: using template matching, collecting platform watermarks (or logos) from mainstream video platforms (or websites) to form a template library (or watermark template library, including watermark templates); performing template matching between images in the video and the template library; if a watermark template that meets a certain matching degree is found, the template matching is terminated, and the watermark contained in the found watermark template is determined as the watermark contained in the image. The template library contains at least one template visual image (such as the aforementioned watermark template), and each template visual image contains one candidate visual element (such as a candidate watermark).
[0158] The detection process of target visual elements (or the element localization process) can be simply divided into two parts: template library establishment and template matching. The following sections will introduce these two parts in detail.
[0159] 1) Template library creation.
[0160] The template library contains at least one template visual image, and each template visual image contains visual elements that need to be removed (such as watermarks, or graphics or text that do not meet display requirements). This application embodiment supports collecting images containing visual elements to construct the template library. For example, images containing visual elements may include images containing platform logos of mainstream video platforms. When generating template visual images based on images containing visual elements after collection, the following points should be noted:
[0161] a. Based on the screen size of the target terminal, crop the image containing visual elements to ensure that the generated template visual image has a certain proportional size information. For example, if the target terminal is a smartphone, and the screen size of a smartphone is approximately 720*1280 pixels, meaning the shorter side of the screen is 720 pixels, then the image containing visual elements can be cropped to 720 pixels, so that the shorter side of the cropped image containing visual elements is 720 pixels.
[0162] b. Considering that most visual elements in images collected from mainstream video platforms are white, to facilitate highlighting visual elements in template visual images, this embodiment also supports first extracting the visual elements from a easily distinguishable background (such as a completely black background), then using an image editing tool (such as Photoshop) to cut out the foreground of the visual elements, and placing the cut-out foreground of the visual elements onto the completely black background image to generate a template visual image. In this template visual image, the visual elements are white and the background is black, thus highlighting the visual elements. Of course, if the visual elements are other colors, the cut-out visual elements can be placed on background images of other easily distinguishable colors. This embodiment does not limit the color of the visual elements or the corresponding background image color, which is not stated here.
[0163] c. Adjust the size of visual elements in the template visual image so that the resized visual elements can occupy a larger display area in the template visual image (e.g., occupy 80% of the display area of the template visual image). This can make the visual elements stand out in the template visual image, improve the accuracy of template matching, and reduce the template matching time.
[0164] 2) Template matching.
[0165] Template matching refers to matching images contained in a video to be processed with template visual images in a template library to find the coordinates of the most matching template visual image in the original image, and then determining the target visual element and its position in the image based on the most matching coordinates. Assuming the video to be processed contains N frames, where N is a positive integer, the specific implementation of template matching may include: First, obtaining a template image from the template library, which contains the target visual element; the template image can be any one of at least one template visual image in the template library; then, using the template image to perform traversal detection on the i-th frame, where i is a positive integer and 1≤i≤N; if the traversal detection result indicates that the template image matches the i-th frame, then the target visual element contained in the i-th frame is determined based on the template image, and the mask region of the target visual element in the i-th frame is determined.
[0166] The template matching process described above will be described in more detail below, which may include steps s11-s13:
[0167] s11: Scale the single-frame image to the same size as the template image, eliminating instability caused by size inconsistencies; for example, if the short side of the template image is 720 pixels, then the short side of the single-frame image will be scaled to 720 pixels. During the scaling process, methods that minimize distortion of the scaled single-frame image should be used, such as linear interpolation, to ensure the quality of the scaled single-frame image.
[0168] s12: Considering that target visual elements are often displayed in the four corners of an image (top left, top right, bottom left, and bottom right), this embodiment also supports cropping a single-frame image first, such as cropping the top and bottom two strip-shaped areas of the single-frame image, and then performing watermark template matching only on these two areas. Of course, target visual elements can also be displayed in the middle of the image. Therefore, this embodiment also supports not cropping the image, but performing template matching between the complete image and the template visual image, which helps to match all target visual elements from the image. In practical application scenarios, whether to crop the image can be chosen according to business needs; this embodiment does not limit this. For ease of explanation, the following describes the process of performing template matching between a single-frame image and the template visual image, using the example of performing template matching between the complete image and the template visual image. See step c for details.
[0169] s13: Template matching is performed using a single-frame image and a template visual image. The specific template matching process may include: First, aligning the reference vertices of the template image with the reference vertices of the i-th frame image. Both the reference vertices of the template image and the i-th frame image can refer to the image itself (e.g., the top left corner of the template image or the i-th frame image). Then, moving the template image multiple times on the i-th frame image according to a set direction (e.g., along the horizontal or vertical direction of the i-th frame image) and a set step size (e.g., one pixel). Based on each movement, obtaining the overlapping region in the i-th frame image that coincides with the template image, and calculating the similarity between the overlapping region and the template image. Next, obtaining the maximum similarity among the similarities obtained from multiple movements. If the maximum similarity is greater than a similarity threshold, it indicates a match between the template image and the i-th frame image. Finally, locating the overlapping region corresponding to the maximum similarity in the i-th frame image, and determining that the overlapping region contains the target visual element. And, based on the position of the target visual element in the overlapping region corresponding to the maximum similarity, determining the mask region of the target visual element in the i-th frame image.
[0170] Combined with appendix Figure 16 A general description of the template matching process described above is provided; such as... Figure 16 As shown, assuming the short side of the single-frame image is 29 pixels, the short side of the template image is 10 pixels, and the similarity threshold is 70%, the template image is first aligned with the top-left corner (reference vertex) of the single-frame image. The template image is then moved horizontally from the top-left corner of the single-frame image, traversing each row and column of the single-frame image in increments of a set step size (e.g., one pixel). At each position reached, the overlapping area between the template image and the single-frame image, and their similarity, are calculated. Specifically, when moving the template image to the right to traverse each row of the single-frame image, if the top-right corner of the template image coincides with the top-right corner of the single-frame image, the traversal stops on that row, and the traversal restarts from the next row, repeating this process until the entire single-frame image has been traversed. Using the above method, multiple similarities of the template image can be obtained by traversing a single frame image. Then, the maximum similarity among multiple pixel values is selected. If the maximum similarity is greater than the similarity threshold, the target visual element contained in the template image is taken as the target visual element of the single frame image. Furthermore, during the traversal process of obtaining the maximum similarity, the position of the template image in the single frame image is determined as the position of the target visual element in the single frame image. Then, the mask region corresponding to the target visual element in the region of the overlapping area corresponding to the maximum value in the currently traversed image is taken as the mask region corresponding to the target visual element.
[0171] See appendix Figure 16It can be seen that when the template image moves to the right for the second time along the first row of the single-frame image, that is, during the second traversal, the similarity of the target visual elements in the overlapping region of the template image and the single-frame image of the same size is the greatest. Figure 16 This is manifested when the target visual elements in the template image and the target visual elements in the single-frame image have a high degree of overlap, and the similarity is 80%, which is greater than the similarity threshold of 70%. In this case, it can be determined that the single-frame image contains the target visual elements, and the mask regions corresponding to the target visual elements are located within the overlapping regions. Of course, if, after traversal, the maximum similarity value among the multiple similarity scores is less than the similarity threshold, it is determined that the target visual elements contained in the template image do not exist in the single-frame image.
[0172] The following are some exemplary similarity calculation methods that can be used to calculate the similarity between visual elements in overlapping regions of the same size as the template image and a single frame image:
[0173] Squared difference matching: Similarity is represented by the sum of the squared differences of all pixels in the overlapping region between the template image and the single-frame image; the smaller the value, the better the matching. When cropping a single-frame image, the similarity is represented by the sum of the squared differences of all pixels in the overlapping region between the cropped strip region and the template image. Neither method is limited here. The formula for calculating squared difference matching is as follows.
[0174] R(x, y) = ∑ x′,y′ (T(x′,y′)-I(x+x′,y+y′)) 2 Formula 1
[0175] Where x and y are the coordinates of the top-left corner of the single-frame image, and x′ and y′ are the coordinates of the top-left corner of the template image. T(x′, y′) is the average pixel value of all pixels in the template image; I(x+x′, y+y′) is the average pixel value of all pixels in the overlapping area between the template image and the single-frame image. (T(x′, y′) - I(x+x′, y+y′)) 2 It is the squared difference of all pixels within the overlapping region of the template image and the single frame image during a single traversal; ∑ x′,y′ (T(x′,y′)-I(x+x′,y+y′)) 2 It is the sum of the squared differences of all pixels in the overlapping area between the template image and the single-frame image after multiple traversals. Based on the above calculation process, the matching coefficient R(x, y) for aligning each pixel in the template image to the top left corner of the single-frame image can be obtained, which is a (xT) x )*(yT y The image of T; where T x It is the width of a single frame image, T y It is the height of a single frame image.
[0176] Relevance matching: This method uses a multiplication operation between the template image and the single-frame image. A higher value indicates a better match. The formula for calculating relevance matching is as follows.
[0177] R(x, y) = ∑ x′,y′ (T′(x′,y′)·I′(x+x′,y+y′)) Formula 2
[0178] The meanings of each parameter and function can be found in the description of each parameter and function in Formula 1, and will not be repeated here.
[0179] Coefficient matching: This involves matching the relative value of the template image to its mean with the correlation value between the template image and the relative value of the single-frame image to its mean. A higher value indicates a better match.
[0180] R(x, y) = ∑ x′,y′ (T′(x′,y′)·I′(x+x′,y+y′)) Formula 3
[0181] in, w is the width of the template image, h is the height of the template image; x″ and y″ are the coordinates when traversing the template image, 1 / (w·h)·∑ x″,y″ T(x″,y″) is the pixel mean of all pixels in the template image. It should be noted that I′ is calculated in a similar way to T′, but will not be described in detail here.
[0182] In practice, the last coefficient matching method yields better results and can record the normalized coefficient matching values. Therefore, in practical applications, this application prefers to use coefficient matching to calculate the similarity between visual elements in overlapping regions of the same size as the template image and a single frame image. Furthermore, the above only provides three exemplary similarity calculation methods. In real-world applications, other similarity calculation methods can be used, and this application does not limit the specific similarity calculation method employed.
[0183] It should be noted that the above description uses the platform watermarks (such as logos) of mainstream video platforms as an example of the visual elements contained in the template library. However, the visual elements contained in the template library can also be other elements, such as text watermarks (such as user accounts), or graphics and text that do not meet display requirements. In this way, the template library can include a relatively rich set of visual elements to facilitate a more comprehensive identification of the target visual elements contained in the video to be processed. Optionally, considering that platform watermarks of mainstream video platforms are relatively easy to obtain, while text watermarks are not easy to obtain, the visual elements contained in the template library usually include platform watermarks of mainstream video platforms, but do not include text watermarks. In this implementation, the embodiments of this application also support the identification of text watermarks contained in a single frame image after the platform watermark is identified. Furthermore, since text watermarks are often displayed in the area associated with the platform watermark, such as the area next to the platform watermark, the embodiments of this application support outlining the area where the text watermark may appear in the vicinity of the platform watermark, converting the image of this area to YCbCr space, and then identifying the text watermark. YCbCr is generally used in continuous image processing in video, where Y represents the brightness of the color, and Cb and Cr are the concentration offset components of blue and red. This color space is more in line with human visual perception, and almost all watermarks are relatively bright areas. A threshold of 240 (or other values) on the brightness channel Y can be used as the cutoff value for text watermarks. This allows for better detection of the mask area (or mask region) where the text watermark is located in the vicinity of the platform watermark. Based on the steps of obtaining the platform mask using template matching and finding the text watermark near the platform watermark using brightness, the watermarks (including platform watermarks, text watermarks, etc.) and the corresponding mask areas of a single frame image can be obtained. Through the above process, not only can some fixed watermarks (such as platform watermarks) be detected, but also irregular watermarks with user IDs (i.e., personalized watermarks, such as user IDs) can be detected, improving the comprehensiveness of watermark detection in videos.
[0184] Furthermore, the embodiments of this application are not limited to using only the above-described methods to detect target visual elements contained in the video. For example, the embodiments of this application also support the use of some deep network segmentation methods to detect target visual elements in the video, thereby improving the robustness of the detection; this is hereby stated.
[0185] (2) Element tracking processing between multiple frames in the image to be processed.
[0186] In the specific implementation, taking a video containing N frames of images, where N is a positive integer, as an example, the process of element tracking is given as follows: Based on the mask region corresponding to the target visual element in the (i-1)th frame image, a similar region is bounded in the i-th frame image; the similarity between the similar region and the mask region in the (i-1)th frame image is calculated; if the similarity result satisfies the similarity condition, it is determined that the i-th frame image contains the target visual element, and the mask region in the (i-1)th frame image is determined as the mask region corresponding to the target visual element contained in the i-th frame image.
[0187] The above process is described in more detail below. First, in the (i-1)th frame image, a similar region (or square region) associated with the mask region is outlined. The mask region can occupy 70% (or other value) of the similar region to make it stand out. Second, within this square region, the similarity between the i-th image and the (i-1)th frame image is calculated to determine whether to use the information from the (i-1)th frame image for tracking. If the similarity condition is met, the mask region of the (i-1)th frame image can be directly used. If the similarity condition is not met, the mask region of the (i-1)th frame image is discarded, and the element localization processing steps described above are directly performed on the i-th frame image. This ensures the stability of the mask region across multiple frames, making subsequent mask region filling less noticeable, while also speeding up the detection time and achieving a better elimination processing effect.
[0188] The following are two exemplary methods for calculating the similarity between similar regions in the (i-1)th frame and similar regions in the i-th frame:
[0189] Mean-square error (MSE): Calculates the distance difference between pixels in two frames (or similar regions in two frames) to obtain the similarity between similar regions in the (i-1)th frame and the i-th frame. The smaller the calculated value, the higher the similarity between similar regions in the (i-1)th frame and the i-th frame. The formula for calculating the mean-square error is as follows.
[0190]
[0191] Where w and h are the width and height of the image; i and j are the coordinates of pixels in the image; I t-1 It is the (i-1)th frame image, I t It is the i-th frame image, I t-1 (i, j)-I t (i, j) is the distance difference between pixels in the (i-1)th frame and the i-th frame.
[0192] Structural Similarity (SSIM): This measure uses three comparisons—brightness, contrast, and structure—to assess the similarity between similar regions in frame x (i-1) and similar regions in frame y (i-1).
[0193] brightness:
[0194] Contrast:
[0195] structure:
[0196] Where μ and σ represent the mean and variance respectively, and c is a constant to prevent division by zero. Here, we take c³ = c² / 2, and multiplying the above three values gives the calculation formula:
[0197]
[0198] It is worth noting that an N*N sliding window is used for each calculation, and then the average is taken to obtain the similarity result between the similar regions in the (i-1)th frame image x and the similar regions in the i-th frame image y. Furthermore, this application embodiment supports using the above method to calculate similarity. In this case, when the result calculated by MSE is less than the first threshold, and the result calculated by SSIM is greater than the second threshold, the similar regions in the (i-1)th frame image x and the similar regions in the i-th frame image y are determined to be similar. The values of the first threshold and the second threshold can be the same or different; this application embodiment does not limit this.
[0199] In summary, through the implementation methods (1) and (2) described above, the target visual elements contained in each frame of the video to be processed and the mask regions corresponding to the target visual elements can be detected.
[0200] Furthermore, before traversing the video to be processed to detect the target visual elements contained in each frame, this embodiment of the application also supports first determining whether the video to be processed contains the target visual elements; if the video to be processed contains the target visual elements, then the target visual elements in the video to be processed are removed; otherwise, if the video to be processed does not contain the target visual elements, then the video to be processed is not removed. This is to reduce the processing time of videos without target visual elements, avoid subsequent processing of the video to be processed, and thus shorten the processing time of the video to be processed; and, it reduces the fault tolerance rate of video watermark detection, providing a better user experience.
[0201] The specific implementation process for detecting whether a video to be processed contains a target visual element may include: extracting N frames of images from the video according to extraction rules, where N is a positive integer; performing template matching on each of the N frames with template visual images in a template library; if any image in the N frames matches any template visual image in the template library, then the video is determined to contain a target visual element, and subsequent elimination processing can be performed on the video. Depending on the definition of the extraction rules, the number and type of images extracted from the video will vary; for example, extraction rules may include, but are not limited to: randomly extracting N frames of images from the video, or uniformly extracting N frames of images from the video according to the duration. Taking the extraction rule of uniformly extracting N frames of images from the video according to the duration as an example, assuming that 6 frames of images need to be extracted from the video, and the video contains a total of 1200 frames, then one frame can be extracted every 200 frames, thus extracting 6 frames of images from 1200 frames. It should be noted that the process of performing template matching on each of the extracted images to identify whether it contains the target visual element can be found in the relevant description above, and will not be repeated here.
[0202] S1503: Fill the mask area in each frame of the image to obtain at least one frame of the image after filling.
[0203] As described above, considering that the video to be processed contains multiple frames, this embodiment divides the filling process into two steps for stability and performance considerations between different frames: filling the mask region corresponding to the target visual element in a single frame image, and smooth filling of the mask region in multiple frames; wherein:
[0204] (1) Filling processing of target visual elements in a single frame image.
[0205] Since the mask region corresponding to the target visual element is often a small portion of a single frame image, this embodiment uses background pixels surrounding the mask region to fill it. This allows the filled mask region to blend better with the background. Compared to simply removing the mask region from a single frame image, the filled mask region in this embodiment looks more natural. Furthermore, this embodiment supports a fast inpainting algorithm based on fast movement (Fine Metal Mask, FMM) to fill the mask region in the image, resulting in a filled image.
[0206] The following is a brief explanation of the main principles of the Fast Model (FMM) based fast-moving inpainting algorithm. The core idea is to first process the pixels on the edges of the region to be repaired (e.g., the mask region), and then proceed layer by layer inwards (towards the interior of the mask region) until all the regions to be repaired are repaired. For filling an unknown pixel p on the edge of the mask region, we take pixels within a certain radius r and its neighborhood ε to fill it. For a known pixel q within the neighborhood, we can calculate its color value based on the brightness gradient.
[0207] Based on the above description, it is clear that in the process of calculating unknown pixels in the mask region using all pixels in the neighborhood ε, the roles of each pixel in the neighborhood ε are not the same. The following uses several weight values to measure the contribution of the calculation results of known pixels to the unknown pixels to be solved. The weights mainly consist of three parts:
[0208] Directional factor: Where N represents the normal direction, meaning that the closer a pixel is to the normal direction, the greater its contribution, which can preserve the boundary information on the edge of the mask area.
[0209] Geometric distance factor: Where d is the distance parameter, which is usually set to 1. That is to say, the closer the pixel is to the unknown pixel, the greater its contribution. This can maintain the consistency between the edge color of the filled mask area and the background color.
[0210] Level set distance factor: T represents the distance from the contour line. This means that the closer a known pixel is to the contour line of the area to be repaired passing through point p, the greater its contribution to the result, which can ensure the boundary continuity of the mask area.
[0211] Multiplying these three factors gives the weight of each point in the region to be repaired. Normalizing this weight using the formula yields the final pixel result.
[0212]
[0213] This application embodiment can use the calculation formula obtained above to fill the mask area. During the filling process, once an unknown pixel within the mask area is filled, that unknown pixel can be treated as a known pixel to fill new unknown pixels. Therefore, the order in which the mask area is filled is particularly important. To ensure that the filled mask area is consistent with the surrounding background, this application embodiment uses a filling order from the outside in, that is, first filling the unknown pixels on the edge of the mask area, and then advancing inward layer by layer. In this way, each filling (or repair) operation targets the outermost pixels.
[0214] The following describes how to determine the outermost pixels of the area to be repaired after each filling operation. Specifically, a narrow edge is constructed for the edge of the area to be repaired. This narrow edge is obtained by expanding the mask area by one ring and then subtracting the original mask area. After determining the narrow edge, the pixels to be repaired are obtained. The purpose is to find the following types of pixels: pixels on the narrow edge, pixels outside the narrow edge that do not need repair, and pixels inside the narrow edge that need repair. In this embodiment, markers are used to identify the above three types of pixels; where BAND indicates pixels on the narrow edge; KNOWN indicates pixels outside the narrow edge that do not need repair; and INSIDE indicates pixels inside the narrow edge that need repair. Each pixel needs to record two values: T represents the distance of the pixel from the narrow edge, and I is the grayscale value. The following is a general description of the movement method of the repaired pixels:
[0215] 1) Initialization; initialize the T value of BAND and KNOWN type pixels to 0, and set the T value of INSIDE type pixels to infinity.
[0216] 2) Define a doubly linked list, NarrowBand. Sort the pixels of BAND in ascending order of their T values and add them to NarrowBand sequentially. Assuming we process point p, change the type of p to KNOWN, remove it from NarrowBand, and then process the four neighboring points Pi of p sequentially. If Pi is of type INSIDE, recalculate the pixel value I, repair the point, update its T value, change its type to BAND, and add it to NarrowBand (still in order, i.e., always keeping NarrowBand in ascending order). Continue this process, processing the pixel with the smallest T value in NarrowBand each time, until there are no pixels left in NarrowBand.
[0217] By following the above repair steps, we can obtain the repair result (or filling result) of the mask area in a single frame image after it has been filled.
[0218] In a scenario where a target visual element is filled, the masked area of the target visual element spans a black boundary region; that is, part of the masked area is located within a colored region of the image, and the other part is located in a black background. All pixels within the black boundary region have a pixel value of 0. To better handle the situation where the target visual element spans a black boundary region, this embodiment also supports using Hough transform to detect straight lines in a single frame image. If a highly reliable straight line is detected, the single frame image is divided into two parts, and the target visual element in each part is filled separately. Then, the two filled images are merged together to obtain the final filled image. The Hough transform is a feature extraction technique used to identify features in an image, such as lines. This embodiment does not provide a detailed description of the Hough transform.
[0219] The method described above, which divides a single-frame image into two parts using a straight line as the boundary, fills the mask regions in each part separately, and then merges the two parts to generate the final filled single-frame image, avoids the cumbersome method of assigning values to pixels in the mask regions within black areas, thus improving filling efficiency. Furthermore, it avoids assigning values to unknown pixels that should have color in black areas, resulting in a more natural-looking repaired mask region. A comparison diagram showing the filling results without distinguishing black boundary regions and with distinguishing black boundary regions can be found in [link to diagram]. Figure 17 ;like Figure 17 As shown, the upper part of the mask area is located within the black boundary area, while the lower part of the mask area is located in the single frame image. If the black boundary area is not distinguished when filling the mask area, the lower part of the filled mask area will be filled with black. However, when the black boundary area is distinguished when filling the mask area, the mask area of the filled single frame image blends into the background of the single frame image, that is, the mask area can be better hidden in the single frame image after the filling process.
[0220] (2) Smooth filling between multiple frames in the image to be processed.
[0221] Considering that the restoration result of the above single-frame image is closely related to the surrounding pixels of the target visual element, and the frequent switching of surrounding pixels will cause jitter in the mask area, the embodiments of this application add inter-frame smoothing logic to the mask area filling process in multi-frame images, which can make the mask area filling result of the multi-frame images after filling process look smoother and more natural.
[0222] The inter-frame smoothing approach mentioned in this embodiment is as follows: calculate the similarity between the non-watermarked region near the mask area of the current frame and the previous frame image, and use the similarity coefficient to fuse the filling result of the current frame image and the filling result of the previous frame image, making the inter-frame transition between the filling result of the previous frame image and the filling result of the current frame image more natural. In a specific implementation, if the video contains adjacent first and second images, and the mask area corresponding to the target visual element in the first image is filled as the first mask filling area, and the mask area corresponding to the target visual element in the second image is filled as the second mask filling area, then the method of obtaining the filling result of the mask area in the second image through inter-frame smoothing may include: taking the mask area corresponding to the target visual element as the center in the second image, outlining a reference area with a display area larger than the display area of the mask area; extracting the target visual element from the reference area to obtain the non-element area; calculating the fusion coefficient between the non-element area and the first mask filling area; using the fusion coefficient, fusing the first mask filling area and the second mask filling area to obtain the updated filling result of the mask area corresponding to the target visual element in the second image.
[0223] The above-described inter-frame smoothing process can be simply summarized as follows: First, a region of X times (e.g., 2 times) the width and height of the mask area is selected from the second image as a reference region for similarity calculation. Then, the similarity metric between two frames mentioned in the visual element tracking section is used to calculate the similarity between the first mask-filled region in the first image and the reference region (i.e., the non-element region) after the mask area is extracted in the second image. Next, the similarity coefficient between 0 and 1 of the SSIM similarity algorithm is used to calculate the fusion coefficient; two thresholds are set: when the similarity is less than the minimum threshold Tmin, the fusion coefficient a = 0.0; when the similarity is greater than the maximum threshold Tmax, the fusion coefficient a = 1.0; when it is between Tmin and Tmax, the similarity is linearly mapped to [0.0, 1.0] as the fusion coefficient a. Finally, the filling result of the first image (e.g., the first mask-filled region) is fused with the filling result of the second image according to the aforementioned single-frame image filling method (e.g., the second mask-filled region) using the fusion coefficient a, resulting in the final filling result of the mask area in the second image. By filling the mask area in the second image using the above-mentioned inter-frame smoothing method, the continuity of the tracking video can be maintained, and there will be no problem of the previous frame remaining when the video switches scenes.
[0224] S1504: According to the playback order of each frame in the video, merge at least one filled frame and at least one frame that does not contain the target visual element to generate the processed video.
[0225] Specifically, assuming the video contains 200 frames, where frames 3, 45, and 89 contain the target visual element, while the other frames do not; after performing the removal process on the target visual element contained in frames 3, 45, and 89 according to the aforementioned steps, we can obtain the processed images of frames 3, 45, and 89, respectively; then, by fusing the 200 frames according to the playback order of the frames in the video, we can generate the processed video.
[0226] S1505: Displays the processed video on the playback interface.
[0227] It should be noted that the specific implementation process of step S1505 can be found in the aforementioned... Figure 7 The specific implementation process of step S704 in the illustrated embodiment will not be repeated here.
[0228] In this embodiment, after acquiring the video to be processed, the system automatically identifies and eliminates target visual elements within the video, without relying on user annotations of these elements. This enhances user experience, improves the intelligence and flexibility of target visual element identification and elimination, and increases the efficiency of target visual element elimination. Furthermore, when identifying target visual elements in each frame of the video, multi-frame visual element tracking is employed. This ensures the stability of the mask region across multiple frames, making subsequent mask region filling less noticeable, while also accelerating detection time and achieving better elimination results. Moreover, when filling the mask region in the multiple frames of the video, smooth filling between frames is used, resulting in a smoother and more natural filled appearance. Additionally, the entire video processing method proposed in this embodiment has a relatively simple computational process and can run quickly on target terminals (such as smartphones and personal computers). This allows for rapid elimination of target visual elements in videos through the target terminal, facilitating product deployment.
[0229] The methods of the embodiments of this application have been described in detail above. In order to facilitate better implementation of the above solutions of the embodiments of this application, the apparatus of the embodiments of this application is provided below.
[0230] Figure 18This illustration shows a schematic diagram of a video processing apparatus provided in an exemplary embodiment of this application; the video processing apparatus can be used as a computer program (including program code) running on a target terminal, for example, the video processing apparatus can be an application (such as a social application) on the target terminal; the video processing apparatus can be used to execute Figure 2 , Figure 7 as well as Figure 15 Some or all of the steps in the method embodiments shown. Please refer to [link / reference]. Figure 18 The video processing device includes the following units:
[0231] Acquisition unit 1801 is used to acquire the video to be processed, which contains target visual elements;
[0232] Processing unit 1802 is used to eliminate target visual elements in the video;
[0233] The processing unit 1802 is also used to display the processed video on the playback interface, in which the target visual elements in the processed video are eliminated.
[0234] In one implementation, the playback interface includes processing options; during video playback, if a processing option is selected, a step to eliminate the target visual element in the video is triggered.
[0235] The target visual elements include watermarks, or the target visual elements include graphics and text that do not meet the display requirements.
[0236] In one implementation, if an elimination processing trigger operation is triggered during video playback, the step of eliminating the target visual element in the video is executed.
[0237] The elimination processing trigger operation includes any of the following: gesture operation, audio signal input operation, vibration operation, and timing operation; timing operation refers to the operation where the display duration of the video on the playback interface exceeds the duration threshold.
[0238] In one implementation, the processing unit 1802 is further configured to:
[0239] When the elimination process for a target visual element in the video is triggered, pause video playback; and...
[0240] The playback interface displays a prompt message indicating that the target visual element in the video is being eliminated.
[0241] In one implementation, the playback interface includes an information prompt window, in which prompt information is displayed; the information prompt window includes a close control; the processing unit 1802 is further configured to:
[0242] During the process of eliminating target visual elements in the video, if the close control is selected, the information prompt window in the playback interface will be closed; and,
[0243] Interrupt the process of eliminating target visual elements in the video;
[0244] If the elimination process is interrupted, the target visual elements in the processed video will fall into any of the following categories: not eliminated, completely eliminated, or partially eliminated.
[0245] In one implementation, the processing unit 1802 is further configured to:
[0246] When the elimination process is finished, a processing result notification is output. The processing result notification is used to notify the elimination result of the target visual element in the video. The elimination result includes any of the following: not eliminated, completely eliminated, or not completely eliminated. The end of the elimination process includes: the elimination process is completed; or the elimination process is interrupted.
[0247] In one implementation, the playback interface includes an element selection entry; the processing unit 1802 is further configured to:
[0248] During the playback of the processed video, in response to the element selection entry being triggered, the obscured area is displayed in the processed playback interface;
[0249] Move the occluded area and cover the reference visual elements in the processed video;
[0250] Play the overlay-processed video, in which reference visual elements are removed.
[0251] In one implementation, the playback interface includes a publishing control, and the processing unit 1802 is further used for:
[0252] In response to the selection of the publish control, the processed video is published.
[0253] In one implementation, the processing unit 1802 is further configured to:
[0254] Displays the service interface of the social application, which includes a video access point;
[0255] When the video acquisition entry is triggered, the video selection interface is displayed, which includes one or more candidate videos;
[0256] In response to a selection operation on one or more candidate videos, the selected candidate video is taken as the video to be processed.
[0257] In one implementation, the video to be processed is the video playing in the video browsing interface, and the processing unit 1802 is further used for:
[0258] The video browsing interface is displayed, and it is used to display different videos.
[0259] In response to a video being selected in the video browsing interface, the selected video is treated as the video to be processed.
[0260] In one implementation, when the processing unit 1802 performs elimination processing on target visual elements in the video, it specifically performs the following:
[0261] The video is processed by detection, segmentation and tracking to obtain at least one frame containing the target visual element, and the mask area corresponding to the target visual element in each frame.
[0262] The mask region in each frame of the image is filled to obtain at least one frame of the image after filling.
[0263] Based on the playback order of the frames in the video, at least one filled frame and at least one frame that does not contain the target visual element are fused together to generate the processed video.
[0264] In one implementation, the detection, segmentation, and tracking process includes element localization or element tracking; the video contains N frames, where N is a positive integer;
[0265] The processing unit 1802 is used to perform detection, segmentation, and tracking processing on each frame of the video to obtain at least one frame of the video containing the target visual element, and the mask region corresponding to the target visual element in each frame of the video. Specifically, it is used for:
[0266] For the i-th frame of the video, if there exists an i-1-th frame and the i-1-th frame contains a target visual element, then perform element tracking processing on the i-th frame based on the i-1-th frame to determine the target visual element contained in the i-th frame and the mask region corresponding to the target visual element in the i-th frame; i is a positive integer and 1≤i≤N;
[0267] If there is no i-1 frame image, or if there is an i-1 frame image but the i-1 frame image does not contain the target visual element, then perform element localization processing on the i-1 frame image to determine the target visual element contained in the i-1 frame image, and the mask area corresponding to the target visual element in the i-1 frame image.
[0268] In one implementation, the processing unit 1802 is used to perform element localization processing on the i-th frame image, determining the target visual elements contained in the i-th frame image, and the target visual elements in the mask region corresponding to the i-th frame image, specifically for:
[0269] Obtain a template image containing the target visual elements;
[0270] The template image is used to perform traversal detection on the i-th frame;
[0271] If the traversal detection results indicate that the template image matches the i-th frame image, then the target visual element contained in the i-th frame image is determined based on the template image, and the mask region of the target visual element in the i-th frame image is determined.
[0272] In one implementation, when the processing unit 1802 performs traversal detection on the i-th frame image using the template image, it is specifically used for:
[0273] Align the reference vertices of the template image with the reference vertices of the i-th frame image;
[0274] The template image is moved multiple times on the i-th frame image according to the set direction and set step size;
[0275] Based on each movement, obtain the overlapping region in the i-th frame image that coincides with the template image, and calculate the similarity between the overlapping region and the template image;
[0276] Obtain the maximum similarity among multiple similarities corresponding to multiple moves. If the maximum similarity is greater than the similarity threshold, it indicates that the template image matches the i-th frame image.
[0277] In one implementation, the processing unit 1802 is used to determine the target visual element contained in the i-th frame image based on the template image, and to determine the mask region of the target visual element in the i-th frame image, specifically for:
[0278] Locate the overlapping region corresponding to the maximum similarity in the i-th frame image, and determine that the overlapping region corresponding to the maximum similarity contains the target visual element; and,
[0279] Based on the position of the target visual element in the overlapping region corresponding to the maximum similarity, the mask region of the target visual element in the i-th frame image is determined.
[0280] In one implementation, the processing unit 1802 is used to perform element tracking processing on the i-th frame image based on the (i-1)-th frame image, to determine the target visual elements contained in the i-th frame image, and the target visual elements in the mask region corresponding to the i-th frame image, specifically for:
[0281] Based on the mask region corresponding to the target visual element in the (i-1)th frame image, outline the similar region in the i-th frame image;
[0282] Calculate the similarity between the similar region and the mask region in the (i-1)th frame image;
[0283] If the similarity result meets the similarity condition, then the i-th frame image is determined to contain the target visual element, and the mask region in the (i-1)-th frame image is determined as the mask region corresponding to the target visual element contained in the i-th frame image.
[0284] In one implementation, if the video contains adjacent first and second images, and the mask region corresponding to the target visual element in the first image is filled as a first mask-filled region, and the mask region corresponding to the target visual element in the second image is filled as a second mask-filled region, then the processing unit 1802 is further configured to:
[0285] In the second image, a reference area with a display area larger than the mask area is drawn out, centered on the mask area corresponding to the target visual element.
[0286] Extract the target visual elements from the reference area to obtain the non-element region;
[0287] Calculate the blending coefficient between the non-elemental region and the first mask-filled region;
[0288] By using a fusion coefficient, the first mask filling region and the second mask filling region are fused to obtain the filling result of the mask region corresponding to the target visual element in the updated second image.
[0289] According to one embodiment of this application, Figure 18 The various units in the video processing apparatus shown can be individually or entirely combined into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the video processing apparatus may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by multiple units working together. According to another embodiment of this application, the video processing apparatus can be executed by running on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 2 , Figure 7 as well as Figure 15 The computer program (including program code) for each step involved in the corresponding method shown, to construct, as Figure 18The video processing apparatus shown herein, and the video processing method for implementing the embodiments of this application, are described. A computer program may be recorded on, for example, a computer-readable recording medium, loaded onto the aforementioned computing device via the computer-readable recording medium, and executed therein.
[0290] In this embodiment of the application, during the playback of the video to be processed, the processing unit 1802 can be used to automatically identify the target visual elements contained in the video; and after identifying the target visual elements, the target visual elements contained in the video are eliminated so that the video after elimination does not contain the target visual elements; this method of automatically identifying and eliminating target visual elements in the video does not rely on the user's annotation of the target visual elements in the video, improves the user experience, enhances the intelligence and flexibility of the identification and elimination of target visual elements, and improves the efficiency of the elimination of target visual elements.
[0291] Figure 19 A schematic diagram of the structure of a terminal provided in an exemplary embodiment of this application is shown. Please refer to... Figure 19 The terminal includes a processor 1901, a communication interface 1902, and a computer-readable storage medium 1103. The processor 1901, communication interface 1902, and computer-readable storage medium 1903 can be connected via a bus or other means. The communication interface 1902 is used to receive and send data. The computer-readable storage medium 1903 can be stored in the terminal's memory and is used to store computer programs, including program instructions. The processor 1901 is used to execute the program instructions stored in the computer-readable storage medium 1903. The processor 1901 (or CPU (Central Processing Unit)) is the terminal's computing and control core, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve corresponding method flows or functions.
[0292] This application embodiment also provides a computer-readable storage medium (Memory), which is a memory device in a terminal for storing programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal and extended storage media supported by the terminal. The computer-readable storage medium provides storage space that stores the processing system of the target terminal. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by the processor 1901, which can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device; optionally, it can also be at least one computer-readable storage medium located remotely from the aforementioned processor.
[0293] In one embodiment, the terminal may be the target terminal mentioned in the foregoing embodiments; the computer-readable storage medium stores one or more instructions; the processor 1901 loads and executes one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the above-described video processing method embodiments; specifically, the one or more instructions in the computer-readable storage medium are loaded by the processor 1901 and executed as follows:
[0294] Acquire the video to be processed, which contains the target visual elements;
[0295] Eliminate target visual elements in the video;
[0296] The processed video is displayed on the playback interface, and the target visual elements in the processed video have been eliminated.
[0297] In one implementation, the playback interface includes processing options; if a processing option is selected, a step to eliminate the target visual element in the video is triggered.
[0298] The target visual elements include watermarks, or the target visual elements include graphics and text that do not meet the display requirements.
[0299] In one implementation, if an elimination processing trigger operation exists, then the step of eliminating the target visual element in the video is triggered.
[0300] The elimination processing trigger operation includes any of the following: gesture operation, audio signal input operation, vibration operation, and timing operation; timing operation refers to the operation where the display duration of the video on the playback interface exceeds the duration threshold.
[0301] In one implementation, one or more instructions in a computer-readable storage medium are loaded by processor 1901 and further executed as follows:
[0302] When the elimination process for a target visual element in the video is triggered, pause video playback; and...
[0303] The playback interface displays a prompt message indicating that the target visual element in the video is being eliminated.
[0304] In one implementation, the playback interface includes an information prompt window, in which prompt information is displayed; the information prompt window has a close control; one or more instructions in the computer-readable storage medium are loaded by the processor 1901 and further execute the following steps:
[0305] During the process of eliminating target visual elements in the video, if the close control is selected, the information prompt window in the playback interface will be closed; and,
[0306] Interrupt the process of eliminating target visual elements in the video;
[0307] If the elimination process is interrupted, the target visual elements in the processed video will fall into any of the following categories: not eliminated, completely eliminated, or partially eliminated.
[0308] In one implementation, one or more instructions in a computer-readable storage medium are loaded by processor 1901 and further executed as follows:
[0309] When the elimination process is finished, a processing result notification is output. The processing result notification is used to notify the elimination result of the target visual element in the video. The elimination result includes any of the following: not eliminated, completely eliminated, or not completely eliminated. The end of the elimination process includes: the elimination process is completed; or the elimination process is interrupted.
[0310] In one implementation, the playback interface includes an element selection entry; one or more instructions in a computer-readable storage medium are loaded by the processor 1901 and further execute the following steps:
[0311] In response to the element selection entry being triggered, the obscured area is displayed in the processed playback interface;
[0312] Move the occluded area and cover the reference visual elements in the processed video;
[0313] Play the overlay-processed video, in which reference visual elements are removed.
[0314] In one implementation, the playback interface includes a publishing control, and one or more instructions in a computer-readable storage medium are loaded by the processor 1901 and further executed as follows:
[0315] In response to the selection of the publish control, the processed video is published.
[0316] In one implementation, one or more instructions in a computer-readable storage medium are loaded by processor 1901 and further executed as follows:
[0317] Displays the service interface of the social application, which includes a video access point;
[0318] When the video acquisition entry is triggered, the video selection interface is displayed, which includes one or more candidate videos;
[0319] In response to a selection operation on one or more candidate videos, the selected candidate video is taken as the video to be processed.
[0320] In one implementation, the video to be processed is the video playing in the video browsing interface, and one or more instructions in the computer-readable storage medium are loaded by the processor 1901 and further executed as follows:
[0321] The video browsing interface is displayed, and it is used to display different videos.
[0322] In response to a video being selected in the video browsing interface, the selected video is treated as the video to be processed.
[0323] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1901 and, when performing the removal process for target visual elements in the video, specifically execute the following steps:
[0324] The video is processed by detection, segmentation and tracking to obtain at least one frame containing the target visual element, and the mask area corresponding to the target visual element in each frame.
[0325] The mask region in each frame of the image is filled to obtain at least one frame of the image after filling.
[0326] Based on the playback order of the frames in the video, at least one filled frame and at least one frame that does not contain the target visual element are fused together to generate the processed video.
[0327] In one implementation, the detection, segmentation, and tracking process includes element localization or element tracking; the video contains N frames, where N is a positive integer;
[0328] When one or more instructions in a computer-readable storage medium are loaded by processor 1901 and executed to perform detection, segmentation, and tracking processing on each frame of a video to obtain at least one frame of the video containing a target visual element, and a mask region corresponding to the target visual element in each frame, the following steps are specifically performed:
[0329] For the i-th frame of the video, if there exists an i-1-th frame and the i-1-th frame contains a target visual element, then perform element tracking processing on the i-th frame based on the i-1-th frame to determine the target visual element contained in the i-th frame and the mask region corresponding to the target visual element in the i-th frame; i is a positive integer and 1≤i≤N;
[0330] If there is no i-1 frame image, or if there is an i-1 frame image but the i-1 frame image does not contain the target visual element, then perform element localization processing on the i-1 frame image to determine the target visual element contained in the i-1 frame image, and the mask area corresponding to the target visual element in the i-1 frame image.
[0331] In one implementation, when one or more instructions in a computer-readable storage medium are loaded by processor 1901 and executed to perform element localization processing on the i-th frame image, determine the target visual elements contained in the i-th frame image, and the mask region corresponding to the target visual elements in the i-th frame image, the following steps are specifically performed:
[0332] Obtain a template image containing the target visual elements;
[0333] The template image is used to perform traversal detection on the i-th frame;
[0334] If the traversal detection results indicate that the template image matches the i-th frame image, then the target visual element contained in the i-th frame image is determined based on the template image, and the mask region of the target visual element in the i-th frame image is determined.
[0335] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1901 and, when executing the traversal detection of the i-th frame image using the template image, specifically perform the following steps:
[0336] Align the reference vertices of the template image with the reference vertices of the i-th frame image;
[0337] The template image is moved multiple times on the i-th frame image according to the set direction and set step size;
[0338] Based on each movement, obtain the overlapping region in the i-th frame image that coincides with the template image, and calculate the similarity between the overlapping region and the template image;
[0339] Obtain the maximum similarity among multiple similarities corresponding to multiple moves. If the maximum similarity is greater than the similarity threshold, it indicates that the template image matches the i-th frame image.
[0340] In one implementation, when one or more instructions in the computer-readable storage medium are loaded by the processor 1901 and executed to determine the target visual element contained in the i-th frame image based on the template image, and to determine the mask region of the target visual element in the i-th frame image, the following steps are specifically performed:
[0341] Locate the overlapping region corresponding to the maximum similarity in the i-th frame image, and determine that the overlapping region corresponding to the maximum similarity contains the target visual element; and,
[0342] Based on the position of the target visual element in the overlapping region corresponding to the maximum similarity, the mask region of the target visual element in the i-th frame image is determined.
[0343] In one implementation, when one or more instructions in the computer-readable storage medium are loaded by the processor 1901 and executed to perform element tracking processing on the i-th frame image based on the (i-1)-th frame image to determine the target visual elements contained in the i-th frame image, and the mask region corresponding to the target visual elements in the i-th frame image, the following steps are specifically performed:
[0344] Based on the mask region corresponding to the target visual element in the (i-1)th frame image, outline the similar region in the i-th frame image;
[0345] Calculate the similarity between the similar region and the mask region in the (i-1)th frame image;
[0346] If the similarity result meets the similarity condition, then the i-th frame image is determined to contain the target visual element, and the mask region in the (i-1)-th frame image is determined as the mask region corresponding to the target visual element contained in the i-th frame image.
[0347] In one implementation, if the video contains adjacent first and second images, and the mask region corresponding to the target visual element in the first image is filled as the first mask-filled region, and the mask region corresponding to the target visual element in the second image is filled as the second mask-filled region, then one or more instructions in the computer-readable storage medium are loaded by the processor 1901 and the following steps are further executed:
[0348] In the second image, a reference area with a display area larger than the mask area is drawn out, centered on the mask area corresponding to the target visual element.
[0349] Extract the target visual elements from the reference area to obtain the non-element region;
[0350] Calculate the blending coefficient between the non-elemental region and the first mask-filled region;
[0351] By using a fusion coefficient, the first mask filling region and the second mask filling region are fused to obtain the filling result of the mask region corresponding to the target visual element in the updated second image.
[0352] In this embodiment of the application, during the playback of the video to be processed, the processor 1901 can be used to automatically identify the target visual elements contained in the video; and after identifying the target visual elements, the processor can eliminate the target visual elements contained in the video so that the video after elimination does not contain the target visual elements; this method of automatically identifying and eliminating target visual elements in the video does not rely on the user's annotation of the target visual elements in the video, improves the user experience, enhances the intelligence and flexibility of the identification and elimination of target visual elements, and improves the efficiency of the elimination of target visual elements.
[0353] This application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The terminal's processor reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the terminal to perform the aforementioned video processing method.
[0354] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0355] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0356] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this invention should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A video processing method, characterized in that, include: Obtain a video to be processed, the video containing target visual elements; the video includes N frames, where N is a positive integer; Each frame of the image is subjected to detection, segmentation and tracking processing to obtain at least one frame of the video containing the target visual element, and a mask region corresponding to the target visual element in each frame of the image; The detection, segmentation, and tracking process includes: if there is an (i-1)th frame image and the (i-1)th frame image contains a target visual element, then perform element tracking processing on the i-th frame image based on the (i-1)th frame image to determine the target visual element contained in the i-th frame image and the mask region corresponding to the target visual element in the i-th frame image; i is a positive integer and 1≤i≤N; if there is no (i-1)th frame image, or if there is an (i-1)th frame image but the (i-1)th frame image does not contain a target visual element, then perform element localization processing on the i-th frame image to determine the target visual element contained in the i-th frame image and the mask region corresponding to the target visual element in the i-th frame image. The mask region in each frame image is filled to obtain at least one filled frame image. The filling process includes: filling the mask region corresponding to the target visual element in a single frame image, and smooth filling the mask region in multiple frames. The smooth filling includes: if the video contains adjacent first and second images, and the mask region corresponding to the target visual element in the first image is filled as a first mask filling region, and the mask region corresponding to the target visual element in the second image is filled as a second mask filling region, then in the second image, with the mask region corresponding to the target visual element as the center, a reference region with a display area larger than the display area of the mask region is outlined; the target visual element is extracted from the reference region to obtain a non-element region; the fusion coefficient between the non-element region and the first mask filling region is calculated; using the fusion coefficient, the first mask filling region and the second mask filling region are fused to obtain the updated filling result of the mask region corresponding to the target visual element in the second image. According to the playback order of each frame in the video, the at least one filled frame and at least one frame that does not contain the target visual element are fused together to generate the processed video; The processed video is displayed on the playback interface, in which the target visual elements have been eliminated.
2. The method as described in claim 1, characterized in that, The playback interface includes processing options; if the processing option is selected, the step of eliminating the target visual elements in the video is triggered. The target visual element may include a watermark, or the target visual element may include graphics and text that do not meet display requirements.
3. The method as described in claim 1, characterized in that, If an elimination processing trigger operation exists, then the step of eliminating the target visual elements in the video is triggered. The elimination processing trigger operation includes any one of the following: gesture operation, audio signal input operation, vibration operation, and timing operation; the timing operation refers to the operation where the display duration of the video on the playback interface exceeds a duration threshold.
4. The method as described in claim 2 or 3, characterized in that, The method further includes: When the elimination process for target visual elements in the video is triggered, the video playback is paused; and... The playback interface outputs a prompt message indicating that the target visual element in the video is being eliminated.
5. The method as described in claim 4, characterized in that, The playback interface includes an information prompt window, and the prompt information is displayed in the information prompt window; the information prompt window has a close control; the method further includes: During the process of eliminating target visual elements in the video, if the close control is selected, the information prompt window is closed in the playback interface; and, The process of eliminating target visual elements in the video is interrupted; If the elimination process is interrupted, the target visual elements in the processed video will fall into any of the following categories: not eliminated, completely eliminated, or partially eliminated.
6. The method as described in claim 1, characterized in that, The method further includes: When the elimination process is completed, a processing result notification is output. The processing result notification is used to notify the elimination result of the target visual element in the video. The elimination result includes any of the following: not eliminated, completely eliminated, or not completely eliminated. The termination of the elimination process includes: completing the elimination process; or, the elimination process being interrupted.
7. The method as described in claim 1, characterized in that, The playback interface includes an element selection entry; the method further includes: In response to the element selection entry being triggered, the obscured area is displayed in the playback interface; Move the occluded area and cover the reference visual elements in the processed video; Play the overlay-processed video, in which the reference visual elements are eliminated.
8. The method as described in claim 1, characterized in that, The playback interface includes a publishing control, and the method further includes: In response to the selection operation of the publishing control, the processed video is published.
9. The method as described in claim 1, characterized in that, The process of acquiring the video to be processed includes: The service interface of the social application is displayed, and the service interface has a video acquisition entry point; When the video acquisition entry is triggered, a video selection interface is displayed, which includes one or more candidate videos; In response to the selection operation of the one or more candidate videos, the selected candidate video is taken as the video to be processed.
10. The method as described in claim 1, characterized in that, The video to be processed is the video currently playing in the video browsing interface. Obtaining the video to be processed includes: The video browsing interface is displayed, and the video browsing interface is used to display different videos; In response to the video selected in the video browsing interface, the selected video is used as the video to be processed.
11. The method as described in claim 1, characterized in that, The step of performing element localization processing on the i-th frame image to determine the target visual elements contained in the i-th frame image, and the mask region of the target visual elements in the i-th frame image, includes: Obtain a template image, wherein the template image contains target visual elements; The template image is used to perform traversal detection on the i-th frame image; If the traversal detection result indicates that the template image matches the i-th frame image, then the target visual element contained in the i-th frame image is determined based on the template image, and the mask region of the target visual element in the i-th frame image is determined.
12. The method as described in claim 11, characterized in that, The step of using the template image to perform traversal detection on the i-th frame image includes: Align the reference vertices of the template image with the reference vertices of the i-th frame image; The template image is moved multiple times on the i-th frame image according to a set direction and a set step size; Based on each movement, the overlapping region in the i-th frame image that coincides with the template image is obtained, and the similarity between the overlapping region and the template image is calculated; Obtain the maximum similarity among the multiple similarities corresponding to the multiple moves. If the maximum similarity is greater than the similarity threshold, then indicate that the template image matches the i-th frame image.
13. The method as described in claim 12, characterized in that, The step of determining the target visual element contained in the i-th frame image based on the template image, and determining the mask region of the target visual element in the i-th frame image, includes: Locate the overlapping region corresponding to the maximum similarity in the i-th frame image, and determine that the overlapping region corresponding to the maximum similarity contains the target visual element; and, Based on the position of the target visual element in the overlapping region corresponding to the maximum similarity, the mask region of the target visual element in the i-th frame image is determined.
14. The method as described in claim 1, characterized in that, The step of performing element tracking processing on the i-th frame image based on the (i-1)-th frame image to determine the target visual elements contained in the i-th frame image, and the mask region of the target visual elements in the i-th frame image, includes: Based on the mask region corresponding to the target visual element in the (i-1)th frame image, a similar region is outlined in the i-th frame image; Calculate the similarity between the similar region and the mask region in the (i-1)th frame image; If the similarity result meets the similarity condition, then the i-th frame image is determined to contain the target visual element, and the mask region in the (i-1)-th frame image is determined as the mask region corresponding to the target visual element contained in the i-th frame image.
15. A video processing apparatus, characterized in that, include: The acquisition unit is used to acquire the video to be processed, which contains the target visual elements. The video consists of N frames, where N is a positive integer; The processing unit is used to perform detection, segmentation and tracking processing on each frame of image to obtain at least one frame of image in the video containing the target visual element, and a mask area corresponding to the target visual element in each frame of image; The detection, segmentation, and tracking process includes: if there is an (i-1)th frame image and the (i-1)th frame image contains a target visual element, then perform element tracking processing on the i-th frame image based on the (i-1)th frame image to determine the target visual element contained in the i-th frame image and the mask region corresponding to the target visual element in the i-th frame image; i is a positive integer and 1≤i≤N; if there is no (i-1)th frame image, or if there is an (i-1)th frame image but the (i-1)th frame image does not contain a target visual element, then perform element localization processing on the i-th frame image to determine the target visual element contained in the i-th frame image and the mask region corresponding to the target visual element in the i-th frame image. The processing unit is further configured to fill the mask region in each frame image to obtain at least one filled frame image; the filling process includes: filling the mask region corresponding to the target visual element in a single frame image, and smooth filling the mask region in multiple frames; the smooth filling includes: if the video contains adjacent first and second images, and the mask region corresponding to the target visual element in the first image is filled as a first mask filling region, and the mask region corresponding to the target visual element in the second image is filled as a second mask filling region, then in the second image, with the mask region corresponding to the target visual element as the center, a reference region with a display area larger than the display area of the mask region is outlined; the target visual element is extracted from the reference region to obtain a non-element region; the fusion coefficient between the non-element region and the first mask filling region is calculated; using the fusion coefficient, the first mask filling region and the second mask filling region are fused to obtain the filling result of the mask region corresponding to the target visual element in the updated second image; The processing unit is further configured to merge the filled at least one frame image and the at least one frame image that does not contain the target visual element according to the playback order of each frame image in the video, and generate the processed video. The processing unit is also used to display the processed video on the playback interface, in which the target visual elements in the processed video are eliminated.
16. A terminal, characterized in that, include: A processor, adapted to execute computer programs; A computer-readable storage medium storing a computer program that, when executed by the processor, implements the video processing method as described in any one of claims 1-14.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and to execute the video processing method as described in any one of claims 1-14.
18. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a processor, implement the video processing method as described in any one of claims 1-14.
Citation Information
Patent Citations
Image processing method and device and electronic equipment
CN110636373A
Video watermark removing method, video data publishing method and related devices
CN110798750A
Video processing method and device, terminal and storage medium
CN112235650A
Video watermark detection method and device, electronic equipment and storage medium
CN112419132A
Method and apparatus for removing watermark from video
WO2017016294A1