Video object elimination method and device
By detecting interference objects in keyframes of video and eliminating interference objects in adjacent frames when necessary, the problem of low efficiency in eliminating interference objects in videos in existing technologies is solved, and fast and efficient video processing is achieved.
Patent Information
- Application Number
- CN202510938876.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies for eliminating interfering objects in videos are inefficient, requiring users to spend a considerable amount of time processing each frame.
By detecting interference objects in keyframes of the video, when it is determined that a keyframe and its adjacent non-keyframes contain interference objects, only that keyframe and its adjacent non-keyframes are eliminated, thus narrowing the detection range and avoiding detection of all video frames.
It enables the rapid elimination of interfering objects in videos, improving the efficiency of electronic devices.
Smart Images

Figure CN120876284A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of video processing technology, and specifically relates to a method and apparatus for eliminating video objects. Background Technology
[0002] Currently, users frequently record videos using electronic devices, such as travel videos. Because video recording requires capturing moving images, it's easier to include distracting objects like passersby or small animals in the recording compared to taking still photos.
[0003] In related technologies, users typically employ traditional selection tools to process recorded videos frame by frame to eliminate interfering objects. However, this process is time-consuming, requiring users to obtain a video free of interfering objects. Therefore, electronic devices are relatively inefficient at eliminating interfering objects from videos. Summary of the Invention
[0004] The purpose of this application is to provide a video object removal method and apparatus that can improve the efficiency of electronic devices in removing interfering objects from videos.
[0005] In a first aspect, embodiments of this application provide a video object removal method, which includes: performing interference object detection on N keyframes in a first video to obtain detection results; and, if, based on the detection results, it is determined that a first object in a first keyframe among the N keyframes is an interference object, and a first video frame corresponding to the first keyframe contains the first object, performing removal processing on the first keyframe and the first object contained in the first video frame to obtain a processed first video; wherein, the first video frame is a non-keyframe in the first video that is adjacent to the first keyframe.
[0006] Secondly, embodiments of this application provide a video object removal apparatus, which includes a detection module and a processing module. The detection module is used to detect interfering objects in N keyframes of a first video and obtain detection results. The processing module is used to, based on the detection results obtained by the detection module, determine that a first object in a first keyframe among the N keyframes is an interfering object, and that a first video frame corresponding to the first keyframe contains the first object, perform removal processing on the first keyframe and the first object contained in the first video to obtain a processed first video; wherein, the first video frame is a non-keyframe in the first video adjacent to the first keyframe.
[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0010] In a sixth aspect, embodiments of this application provide a computer program / program product stored in a storage medium, which is executed by at least one processor to implement the method as described in the first aspect.
[0011] In this embodiment, interference object detection is performed on N keyframes in the first video to obtain detection results. If, based on the detection results, it is determined that a first object in the first keyframe among the N keyframes is an interference object, and the first video frame corresponding to the first keyframe contains the first object, then the first keyframe and the first object contained in the first video frame are eliminated to obtain the processed first video. Here, the first video frame is a non-keyframe adjacent to the first keyframe in the first video. In this solution, by performing interference object detection on the keyframes of the first video, when it is determined that a certain keyframe and its adjacent non-keyframes contain interference objects, the interference objects in that keyframe and its adjacent non-keyframes are eliminated. That is, by narrowing the interference object detection range of the electronic device, interference objects in the video are eliminated, avoiding interference object detection on all video frames of the first video. This achieves rapid elimination of interference objects in the video, thereby improving the efficiency of the electronic device in eliminating interference objects in the video. Attached Figure Description
[0012] Figure 1 This is one of the flowcharts illustrating a video object removal method provided in an embodiment of this application;
[0013] Figure 2 This is a second schematic flowchart of a video object removal method provided in an embodiment of this application;
[0014] Figure 3(A) is one of the schematic diagrams of an example of eliminating interfering objects in a video according to an embodiment of this application;
[0015] Figure 3(B) is a second schematic diagram of an example of eliminating interfering objects in a video according to an embodiment of this application;
[0016] Figure 4This is the third example of an embodiment of the present application that illustrates the elimination of interfering objects in a video;
[0017] Figure 5 This is the third flowchart illustrating a video object removal method provided in this application embodiment;
[0018] Figure 6(A) is one of the schematic diagrams of an example of playing video provided in the embodiments of this application;
[0019] Figure 6(B) is a second schematic diagram of an example of playing video provided in the embodiments of this application;
[0020] Figure 7 This is the fourth flowchart of a video object removal method provided in the embodiments of this application;
[0021] Figure 8(A) is a third example of a video playback embodiment provided in this application;
[0022] Figure 8(B) is a fourth example of a video playback embodiment provided in this application.
[0023] Figure 9 This is a schematic diagram of the structure of a video object removal device provided in an embodiment of this application;
[0024] Figure 10 This is one of the hardware structure diagrams of an electronic device provided in the embodiments of this application;
[0025] Figure 11 This is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0026] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0027] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0028] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."
[0029] The following is a definition of the technical terms used in the embodiments of this application:
[0030] 1. Interference objects refer to objects or elements in the frame that are unrelated to the subject or core content of the shot and may distract the viewer. For example: (1) Items unrelated to the scene: such as trash cans, untidy tables and chairs, and stacked cardboard boxes in the background when shooting indoors; billboards, discarded items, and personal belongings of pedestrians when shooting outdoors. (2) Natural debris: such as leaves blown down by the wind, dust falling, and garbage floating on the water (common in natural scene shooting). (3) Shooting equipment leakage: such as fingers exposed when shooting handheld, a corner of a tripod, or shadows and wires of lighting equipment. (4) Interference objects in post-processing: such as watermarks that have not been cleaned, incorrect subtitles, and pixel noise left by special effects. (5) Clothing or accessory interference: such as exaggerated accessories unrelated to the theme, wrinkled or stained clothing on the person in the shot; the action of a person (such as waving casually) accidentally entering the frame when shooting multiple people.
[0031] 2. RGB images are digital images that present rich colors by superimposing the colors of the three primary color channels: red, green, and blue.
[0032] 3. A depth image is a special type of image that records the distance information between video objects in a scene and electronic devices. Each pixel value represents the depth (distance) from that pixel to the electronic device.
[0033] The video object removal method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0034] The video object removal method in this application embodiment can be applied to scenarios where interfering objects in a video need to be removed.
[0035] For example: After shooting a travel video, remove the scenes of passersby from the video.
[0036] For example, after shooting a home video, remove the scene of the dog from the video.
[0037] In the solution provided in this application embodiment, interference objects are detected on the key frames of the first video. When it is determined that a certain key frame and its adjacent non-key frames contain interference objects, the interference objects in the certain key frame and its adjacent non-key frames are eliminated. That is, the interference objects in the video are eliminated by narrowing the interference object detection range of the electronic device, avoiding interference object detection on all video frames of the first video, realizing the rapid elimination of interference objects in the video, thereby improving the efficiency of the electronic device in eliminating interference objects in the video.
[0038] The execution subject of the video object removal method provided in this application embodiment can be a video object removal device, which can be an electronic device, or a functional module or entity within an electronic device. The following uses an electronic device as an example to illustrate the technical solution provided in this application embodiment.
[0039] This application provides a method for video object removal. Figure 1 A flowchart of a video object removal method provided in an embodiment of this application is shown, which can be applied to electronic devices. Figure 1 As shown, the video object removal method provided in this application embodiment may include the following steps 201 and 202.
[0040] Step 201: The electronic device performs interference object detection on N key frames in the first video and obtains the detection results.
[0041] In the embodiments of this application, N is an integer greater than 1.
[0042] In some embodiments of this application, the first video can be any of the following: landscape video, travel video, animation video, character video, etc. The specific type can be determined according to actual usage requirements, and this application does not limit this.
[0043] In some embodiments of this application, the aforementioned first video can be a short video corresponding to a live photo.
[0044] In some embodiments of this application, the aforementioned N keyframes are video frames in the first video that contain major scene changes.
[0045] For example, when recording a video at the beach, a video frame of a "puppy" suddenly appears in the recording scene.
[0046] In some embodiments of this application, the aforementioned N keyframes may be video frames corresponding to shot switching points in the first video.
[0047] It is understandable that when the camera switches, the scene, perspective, or subject of the image will change abruptly, so these video frames can be used as keyframes of the video.
[0048] In some embodiments of this application, the electronic device can use a multimodal deviation model to detect interference objects in N keyframes and obtain the above detection results.
[0049] In some embodiments of this application, the above detection results are used to indicate interference objects in N keyframes.
[0050] In some embodiments of this application, the detection result may include M keyframes, each of which carries a second identifier for marking interference objects in the M keyframes, where M≤N.
[0051] In some embodiments of this application, the above-mentioned M keyframes are all or part of the N keyframes. It can be understood that the M keyframes are keyframes containing the interfering object among the N keyframes.
[0052] In some embodiments of this application, the second identifier can be any of the following: a rectangle, a dot, an irregular shape, etc. The specific identifier can be determined according to actual usage requirements, and this application does not limit this.
[0053] In some embodiments of this application, the aforementioned interference target can be at least one of the following: animals (e.g., cats), passersby, insects, plastic bags, vehicles, etc. The specific target can be determined based on actual usage needs, and this application does not limit this.
[0054] For example, if a keyframe out of N keyframes contains a "puppy", the detection result can be that keyframe, and that keyframe carries a rectangular box to identify the "puppy" as a distraction object. This detection result is used to indicate that the "puppy" is a distraction object.
[0055] In some embodiments of this application, when the electronic device displays the video playback interface of the first video, if a fourth input from the user is received, the electronic device can perform interference object detection on N key frames in the first video and obtain the detection result.
[0056] Step 202: If, based on the detection results, it is determined that the first object in the first key frame among N key frames is an interference object, and the first video frame corresponding to the first key frame contains the first object, the electronic device performs elimination processing on the first key frame and the first object contained in the first video to obtain the processed first video.
[0057] In some embodiments of this application, the first keyframe described above may be one or more keyframes among N keyframes.
[0058] In some embodiments of this application, the first object can be any of the following: animals, pedestrians, insects, plastic bags, vehicles, etc. The specific object can be determined based on actual usage needs, and this application does not limit this.
[0059] In some embodiments of this application, when the detection result indicates that the first object in the first keyframe is an interference object, the electronic device can determine that the first object is an interference object; or, when the detection result indicates that the first object in the first keyframe out of N keyframes is an interference object, and the electronic device receives user input on the first object in the first keyframe, the electronic device can determine that the first object is an interference object; or, when the detection result indicates that the first object in the first keyframe out of N keyframes is an interference object, and the electronic device detects that the first object gradually moves away from the subject being filmed or remains at a constant distance from the subject being filmed in a preset number of video frames after the first keyframe, the electronic device can determine that the first object is an interference object.
[0060] In some embodiments of this application, if the detection result indicates that the first object in the first keyframe is an interfering object, and at least one keyframe in a predetermined number of consecutive keyframes following the first keyframe does not contain the first object, the electronic device can determine that the first object is an interfering object. It can be understood that if the first object appears consecutively in keyframes, it is considered that the first object is not an interfering object.
[0061] It is understandable that, based on the detection results, if the first object in the first key frame among N key frames is determined to be an interfering object, the electronic device can determine whether the first video frame corresponding to the first key frame contains the first object. Then, if the first video frame contains the first object, the electronic device can perform elimination processing on the first key frame and the first object contained in the first video to obtain the processed first video.
[0062] In the embodiments of this application, the first video frame is a non-key frame in the first video that is adjacent to the first key frame.
[0063] In some embodiments of this application, the first video frame may be all non-key frames adjacent to the first key frame in the first video, or a preset number of non-key frames.
[0064] For example: Suppose that the N keyframes are the 10th, 20th and 30th frames in the first video, and the first keyframe is the 20th frame, then the first video frames can be the 11th to 19th frames and the 21st to 29th frames, or the first video frames can be the 15th to 19th frames and the 21st to 25th frames.
[0065] For example: Suppose that the N keyframes are the 5th, 10th and 20th frames in the first video, and the first keyframe is the 5th and 10th frames, then the first video frames can be the 1st to the 4th frames, the 6th to the 9th frames, and the 11th to the 19th frames.
[0066] In some embodiments of this application, the first video frame may be a non-key frame in the first video that is adjacent to both the first key frame and the frame before and after it, or it may be a non-key frame in the first video that is adjacent to the first key frame before it, or it may be a non-key frame in the first video that is adjacent to the first key frame after it.
[0067] For example: Suppose that the N keyframes are the 10th, 15th and 30th frames in the first video, and the first keyframe is the 15th frame, then the first video frame can be the 11th to the 14th frame, or the 16th to the 29th frame, or the 11th to the 14th frame, and the 16th to the 29th frame.
[0068] In some embodiments of this application, when the first object in the first keyframe is an interfering object and the first video frame contains the first object, the electronic device can also display thumbnails of the first keyframe and the first video frame on the video playback interface of the first video, as well as display prompt information. The prompt information is used to prompt the user that the first object in the video frame corresponding to these thumbnails is an interfering object. Then, when the electronic device receives the user's input of the prompt information, it can perform elimination processing on the first object contained in the first keyframe and the first video frame to obtain the processed first video.
[0069] In some embodiments of this application, the electronic device may also receive input from the user for any one of the thumbnails of the first keyframe and the first video frame, and display the video frame corresponding to the any one thumbnail, so that the user can determine whether to perform elimination processing on the first object in the video frame corresponding to the any one thumbnail.
[0070] In some embodiments of this application, the electronic device can acquire the depth image and motion vector data corresponding to the first keyframe, and then perform elimination processing on the first object contained in the first keyframe based on the depth image and motion vector data corresponding to the first keyframe.
[0071] Similarly, the electronic device can acquire the depth image and motion vector data corresponding to the first video frame, and then perform elimination processing on the first object contained in the first video frame based on the depth image and motion vector data corresponding to the first video frame.
[0072] Understandably, depth images provide distance information for various video objects in a scene, allowing electronic devices to determine which are foreground distracting objects and which are part of the background. Motion vector data, on the other hand, analyzes the movement of each video object within the frame, further distinguishing between dynamic video objects and static background. Therefore, based on the depth image and motion vector data corresponding to a video frame, electronic devices can effectively separate the distracting object layer from the background layer, thus eliminating distracting objects within that video frame.
[0073] In some embodiments of this application, after the first object contained in the first keyframe is eliminated, the electronic device can call the repair model to complete the background content of the first keyframe and the first video frame, and render realistic dynamic textures.
[0074] Specifically, electronic devices can undergo area-by-area repair to achieve a more natural effect. For static background areas, direct sampling and filling with surrounding pixels is used to make the transition smoother. When processing dynamic texture areas, such as flowing water or flames, a physical model is used to generate realistic dynamic effects, making the image more vivid. In addition, multi-level transparency adjustment is supported to ensure that the repaired edges are smooth and natural, without appearing harsh.
[0075] In the video object removal method provided in this application embodiment, interference objects are detected on key frames of the first video. When it is determined that a certain key frame and its adjacent non-key frames contain interference objects, the interference objects in the certain key frame and its adjacent non-key frames are removed. That is, interference objects in the video are removed by narrowing the interference object detection range of the electronic device, avoiding interference object detection on all video frames of the first video, realizing the rapid removal of interference objects in the video, thereby improving the efficiency of the electronic device in removing interference objects in the video.
[0076] In some embodiments of this application, if the electronic device does not completely eliminate the first object after the first object contained in the first keyframe and the first video frame has been eliminated, the user can manually select the area that has not been eliminated so that the electronic device can completely eliminate the first object.
[0077] In some embodiments of this application, users can use a manual smearing tool to finely adjust and modify any video frame in the first video to ensure that any residual interference objects are completely removed, ultimately obtaining a satisfactory picture effect.
[0078] In some embodiments of this application, combined with Figure 1 ,like Figure 2 As shown, step 201 can be implemented through steps 201a and 201b below.
[0079] Step 201a: The electronic device acquires N RGB images corresponding to N key frames, N depth images corresponding to N key frames, and N motion vector data corresponding to N key frames.
[0080] In some embodiments of this application, each of the above N keyframes corresponds to an RGB image, each keyframe corresponds to a depth image, and each keyframe corresponds to a motion vector data.
[0081] In the embodiments of this application, each of the above motion vector data is used to characterize the motion parameters of at least one video object in the keyframe corresponding to each motion vector data.
[0082] In some embodiments of this application, the video object can be any of the following: animals, passersby, insects, plastic bags, vehicles, plants, etc.
[0083] For example, if keyframe A, which contains the video object "pedestrian", is one of N keyframes, the motion vector data of keyframe A can be used to characterize the motion parameter of the pedestrian in keyframe A from the previous keyframe B to keyframe A, which is "the pedestrian moves 100 pixels to the right".
[0084] It is understandable that electronic devices can acquire the RGB image, depth image, and motion vector data corresponding to each keyframe during the recording of the first video.
[0085] Step 201b: The electronic device performs interference object detection on N keyframes based on N RGB images, N depth images and N motion vector data to obtain the detection results.
[0086] It is understandable that electronic devices can perform interference object detection on each keyframe based on the RGB image, the depth image, and the motion vector data corresponding to each keyframe, and obtain the above detection results.
[0087] Specifically, the electronic device can extract the RGB feature information of the RGB image corresponding to each key frame, the depth feature information of the depth image corresponding to each key frame, and the motion feature information of the motion vector data corresponding to each key frame. Then, the electronic device can perform fusion processing on the RGB feature information, the depth feature information, and the motion feature information corresponding to each key frame to obtain the multimodal feature data corresponding to each key frame. Based on the multimodal feature data corresponding to each key frame, the electronic device can determine the interference objects in N key frames and obtain the above detection results.
[0088] In some embodiments of this application, the electronic device can use a multimodal deviation model to detect interference objects in N keyframes and obtain detection results.
[0089] Specifically, the electronic device can input N RGB images, N depth images, and N motion vector data into a multimodal deviation model. The multimodal deviation model extracts the RGB feature information of the RGB image corresponding to each keyframe, the depth feature information of the depth image corresponding to each keyframe, and the motion feature information of the motion vector data corresponding to each keyframe. Then, the multimodal deviation model can fuse the RGB feature information, depth feature information, and motion feature information corresponding to each keyframe to obtain the multimodal feature data corresponding to each keyframe. Based on the multimodal feature data corresponding to each keyframe, the multimodal deviation model can determine the interference objects in the N keyframes and finally output the detection results.
[0090] It can be understood that when an interfering object is detected in M of the above N keyframes, the multimodal deviation model can output the M keyframes and carry a second identifier for marking the interfering object on the M keyframes.
[0091] It should be noted that for the detailed steps of electronic devices using a multimodal deviation model to detect interference objects in N keyframes and obtain detection results, please refer to the relevant descriptions in related technologies, which describe how electronic devices use a multimodal deviation model to detect interference objects in keyframes of video and obtain detection results. These will not be repeated here.
[0092] In some embodiments of this application, the electronic device can analyze the RGB image corresponding to each video frame, the depth image corresponding to each video frame, and the motion vector data corresponding to each video frame in the first video to construct a three-dimensional spatiotemporal grid. Then, by comparing the motion trajectory, depth changes, and texture changes of objects in consecutive frames, it can identify those "abnormal" interference objects that deviate from the norm, such as passersby, flying insects, pets, or suddenly appearing obstacles that accidentally enter the frame.
[0093] For example, if someone suddenly enters the scene during filming, the frame number Fx can be recorded. At this time, the RGB image displays the person's color outline information [(R1,G1,B1), (R2,G2,B2), (R3,G3,B3), ..., (Rn),Gn),Bn], the depth image displays the person's positional changes [(X1,Y1,Z1), (X2,Y2,Z2), (X3,Y3,Z3), ..., (Xn,Yn,Zn)], and the motion vector data displays the person's direction of movement (Vx,Vy,Vz). The electronic device can then determine whether the person exists in the subsequent keyframes Fy and Fz. If at least one keyframe does not contain the person, the person is identified as an interfering object, and the position of the person in the keyframe Fx image [(X1,Y1), (X2,Y2)] is recorded, triggering a person recognition marker.
[0094] In some embodiments of this application, after step 201 described above, the video object removal method provided in this application further includes steps 301 and 302 as described below.
[0095] Step 301: If the detection result indicates that the first object in the first keyframe is an interfering object, the electronic device displays the first keyframe.
[0096] In the embodiments of this application, the first keyframe carries a first identifier, which is used to mark the first object in the first keyframe as an interference object.
[0097] In some embodiments of this application, the first identifier can be any of the following: a rectangle, a dot, an irregular shape, etc. The specific identifier can be determined according to actual usage requirements, and this application does not limit this.
[0098] In some embodiments of this application, the electronic device may display a first keyframe carrying a first identifier on the video playback interface of the first video.
[0099] In some embodiments of this application, when the first keyframe is multiple keyframes, the electronic device can display any one of the keyframes.
[0100] Step 302: The electronic device responds to the user's first input on the first object in the first keyframe and determines that the first object is an interference object.
[0101] In some embodiments of this application, the electronic device can receive a first input from a user on a first object in a first keyframe, and in response to the first input, determine that the first object is an interference object.
[0102] In some embodiments of this application, the first input described above is used to eliminate a first object in a first video.
[0103] In some embodiments of this application, the aforementioned first input includes, but is not limited to: touch input from a user to the first object via a touch device such as a finger or stylus, a voice command input by the user, a specific gesture input by the user, a click input, or other feasible inputs. The specific input can be determined according to actual usage needs, and this application does not limit this.
[0104] In some embodiments of this application, the specific gesture mentioned above can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long-press gesture, an area change gesture, a double-press gesture, or a double-tap gesture.
[0105] For example, if a user wants to remove interfering objects from a travel video (i.e., the first video mentioned above), the user can trigger an electronic device to detect interfering objects in N keyframes of the travel video and obtain the detection results. If the detection results indicate that the bucket (i.e., the first object mentioned above) in the first keyframe of the travel video is an interfering object, as shown in Figure 3(A), the electronic device can display the first keyframe 11, which includes a rectangular border 12 (i.e., the first identifier mentioned above). Then, the user can long-press (i.e., the first input mentioned above) the bucket 13 in the rectangular border 12, so that the electronic device determines that the bucket 12 is an interfering object. Then, when the bucket 13 is contained in the non-keyframe (i.e., the first video frame mentioned above) adjacent to the first keyframe 11, the electronic device can remove the bucket 13 in the first keyframe 11 and the non-keyframe containing the bucket 13, as shown in Figure 3(B). The first keyframe 11 does not contain the bucket 13, thus obtaining the processed travel video.
[0106] Thus, after detecting the interfering object, the electronic device will require the user to confirm it again. For example, the user can select the area where the interfering object is located, which can ensure the accuracy of eliminating the interfering object.
[0107] In some embodiments of this application, the first input includes a first sub-input and a second sub-input; in response to the user's first sub-input on the first object in the first keyframe, the electronic device overlays and displays a first floating layer on the first keyframe, the first floating layer including the first object, and the display area of the first object including a cancellation control; then, in response to the user's second sub-input on the cancellation control, the electronic device determines that the first object is an interference object, and if the first object is contained in the first video frame corresponding to the first keyframe, performs cancellation processing on the first object contained in the first keyframe and the first video frame to obtain the processed first video.
[0108] For example, referring to Figure 3(A), such as Figure 4As shown, after the user long-presses (i.e., the first sub-input mentioned above) the bucket 13 in the rectangular border 12, the electronic device can display a first floating layer 21 on the first keyframe 11. The first floating layer 21 includes the bucket 13. The display area of the bucket 13 includes a removal control 22. The user can click (i.e., the second sub-input mentioned above) the removal control 22, so that the electronic device determines that the bucket 13 is an interference object, and when the bucket 13 is included in the non-keyframe adjacent to the first keyframe 11, the electronic device performs removal processing on the bucket 13 in the first keyframe 11 and the non-keyframe containing the bucket 13.
[0109] In some embodiments of this application, when displaying a first keyframe, if the first keyframe contains an interfering object (e.g., a third object) that is not detected by the electronic device, the electronic device can receive the user's selection input for the third object, so that the electronic device determines that the third object is an interfering object, and if the first video frame corresponding to the first keyframe contains the third object, the electronic device performs elimination processing on the third object contained in the first keyframe and the first video frame.
[0110] In some embodiments of this application, after receiving the user's selection input, the electronic device can expand the selected graphic into a rectangle according to the shape of the selected graphic. Then, the electronic device can identify the content within the rectangular area. When a third object is identified, the electronic device determines that the third object is an interference object.
[0111] In some embodiments of this application, after the third object contained in the first keyframe and the first video frame is eliminated, the electronic device can sample the color value of the outer edge of the rectangle as the fill color value of the removed area. At the same time, the user can manually sample the color of the surrounding pixels as the supplementary color value of the removed area.
[0112] In some embodiments of this application, if the detection result indicates that the second object in the second key frame out of N key frames is an interference object, the electronic device can determine that the second object is an interference object; or, if the detection result indicates that the second object in the second key frame out of N key frames is an interference object, and it is detected that the second object gradually moves away from the subject in subsequent video frames or the distance between the second object and the subject remains unchanged, the electronic device can determine that the second object is an interference object; or, if the detection result indicates that the second object in the second key frame out of N key frames is an interference object, and the user inputs the second object in the second key frame, the electronic device can determine that the second object is an interference object.
[0113] In some embodiments of this application, after step 201 described above, the video object removal method provided in this application further includes step 401 or step 402 as described below.
[0114] Step 401: If the detection result indicates that the second object in the second keyframe among N keyframes is an interference object, and if the first distance is less than the second distance, the electronic device determines that the second object is an interference object.
[0115] In the embodiments of this application, the first distance is the distance between the second object in the second keyframe and the shooting subject in the second keyframe, the second distance is the distance between the second object in the second video frame and the shooting subject in the second video frame, the second video frame is the video frame in the first video that is adjacent to the second keyframe, and the shooting time of the second video frame is later than the shooting time of the second keyframe.
[0116] In some embodiments of this application, the second object can be any of the following: animals, pedestrians, insects, plastic bags, vehicles, etc. The specific object can be determined based on actual usage needs, and this application does not limit this.
[0117] In some embodiments of this application, the subject of the photograph can be any of the following: animals, people, buildings, plants, vehicles, etc. The specific subject can be determined according to actual usage needs, and this application does not limit this.
[0118] In some embodiments of this application, the second video frame is a preset number of video frames in the first video that are adjacent to the second keyframe.
[0119] For example, when the second keyframe is the 10th video frame in the first video, the second video frame can be the 11th or 12th video frame in the first video.
[0120] In some embodiments of this application, the subject being photographed in the second keyframe is the same as the subject being photographed in the second video frame.
[0121] It is understandable that after obtaining the detection result of the interference object detection, if the detection result indicates that the second object in the second key frame out of N key frames is the interference object, the electronic device can obtain the first distance between the second object in the second key frame and the shooting subject in the second key frame, and obtain the second distance between the second object in the second video frame after the second key frame and the shooting subject in the second video frame. Then, when the first distance is less than the second distance, that is, when the second object is far away from the shooting subject, the second object is determined to be the interference object.
[0122] Step 402: If the detection result indicates that the second object in the second keyframe among N keyframes is an interfering object, and if the first distance is greater than or equal to the second distance, the electronic device determines that the second object is not an interfering object.
[0123] It is understandable that after obtaining the detection result of the interference object detection, if the detection result indicates that the second object in the second key frame out of N key frames is the interference object, the electronic device can obtain the first distance between the second object in the second key frame and the shooting subject in the second key frame, and obtain the second distance between the second object in the second video frame after the second key frame and the shooting subject in the second video frame. Then, when the first distance is greater than the second distance, that is, when the second object is close to the shooting subject, it is determined that the second object is not the interference object. Alternatively, when the first distance is equal to the second distance, that is, when the distance between the second object and the shooting subject remains unchanged, it is determined that the second object is not the interference object.
[0124] Thus, after the detection result indicates that the second object in the second keyframe is an interference object, the electronic device can further determine whether the second object is an interference object based on the distance between the second object and the shooting subject in a preset number of video frames after the second keyframe, thereby improving the accuracy of interference object determination.
[0125] In some embodiments of this application, the electronic device may determine that the second object is not an interfering object when the second object gradually approaches the shooting subject in a preset number of video frames after the second keyframe until the distance between the second object and the shooting subject remains unchanged.
[0126] In some embodiments of this application, when the detection result indicates that the first object in the first keyframe is an interfering object and the third distance is less than the fourth distance, the electronic device can determine that the first object is an interfering object.
[0127] In some embodiments of this application, the third distance is the distance between the first object in the first keyframe and the subject being filmed in the first keyframe. The fourth distance is the distance between the first object in the fifth video frame following the first keyframe and the subject being filmed in the fifth video frame.
[0128] In some embodiments of this application, the fifth video frame is a preset number of video frames adjacent to the first keyframe in the first video, and the shooting time of the fifth video frame is later than the shooting time of the first keyframe.
[0129] In some embodiments of this application, the subject being photographed in the first keyframe is the same as the subject being photographed in the fifth video frame.
[0130] In some embodiments of this application, before step 201 above, the video object removal method provided in this application further includes steps 501 to 503 as described below.
[0131] Step 501: The electronic device acquires the color histogram corresponding to each video frame in the first video.
[0132] In some embodiments of this application, the electronic device can convert each video frame to a suitable color space (e.g., HSV or RGB), then count the number of times each color value appears for each video frame, and generate a color histogram for each video frame.
[0133] It should be noted that for detailed steps on how an electronic device acquires the color histogram corresponding to each video frame in the first video, please refer to the description of the electronic device acquiring the color histogram corresponding to the video frame in related technologies, which will not be repeated here.
[0134] Step 502: The electronic device calculates the difference between the color histograms corresponding to each two adjacent video frames in the first video.
[0135] In some embodiments of this application, the electronic device may use any of the following to calculate the difference between the color histograms corresponding to every two adjacent video frames in the first video: chi-square distance, Bach distance, and Euclidean distance.
[0136] Step 503: If the difference between the color histogram corresponding to the third video frame and the color histogram corresponding to the fourth video frame is greater than or equal to a preset threshold, the electronic device uses the fourth video frame as a key frame in the first video.
[0137] In the embodiments of this application, the third video frame and the fourth video frame are two adjacent video frames in the first video, and the fourth video frame is captured later than the third video frame.
[0138] It is understandable that when the difference between the color histogram corresponding to the third video frame and the color histogram corresponding to the fourth video frame is greater than or equal to a preset threshold, the scene is considered to have changed. Therefore, the fourth video frame, which was captured later, can be used as a key frame in the first video.
[0139] In the embodiments of this application, the electronic device can use the color histogram method to determine key frames in the first video, calculate the color histogram of each video frame, and when the histograms of two frames differ significantly, it is considered that a scene change has occurred. At this time, the subsequent video frame can be used as the key frame. Scene changes are detected by analyzing the changes between video frames. Key frames (such as shot transition points, frames with sudden changes in object motion) are automatically extracted from the video, and an interference object tracking chain is established. The user only needs to select the interference object area in the key frame, and the system automatically maps it to the adjacent frames of the key frame, reducing manual operation. Interference objects in the key frame are identified and attracted, and the key frame is compared with the adjacent non-key frames in a round-robin fashion. The interference objects identified in the key frame are marked on the non-key frames to eliminate the interference objects in these frames.
[0140] In some embodiments of this application, combined with Figure 1 ,like Figure 5 As shown, after step 202 above, the video object removal method provided in this application embodiment further includes the following steps 601 and 602.
[0141] Step 601: The electronic device displays the video playback interface.
[0142] In the embodiments of this application, the video playback interface includes a first area and a second area. The first area displays the processed video frame of the first video, and the second area displays the video frame of the first video. The display area of the first area is larger than the display area of the second area.
[0143] It is understandable that after the electronic device receives the processed first video, it can display the aforementioned video playback interface, and while displaying the processed first video on the video playback interface, it can also display the original first video in a smaller area.
[0144] Step 602: In response to the user's second input to the second area, the electronic device displays the video frame of the first video in the first area and the processed video frame of the first video in the second area.
[0145] In some embodiments of this application, the second input described above is used to view the original first video.
[0146] In some embodiments of this application, the aforementioned second input includes, but is not limited to: touch input by the user on the second area using a touch device such as a finger or stylus, or a voice command input by the user, or a specific gesture input by the user, or a click input, or other feasible input. The specific input can be determined according to actual usage needs, and this application embodiment does not limit this.
[0147] It is understandable that when the electronic device receives input from the user into the second area, it can swap the display area of the processed first video with the display area of the first video.
[0148] For example, as shown in Figure 6(A), after obtaining the processed travel video, if a user wants to view the processed travel video, the user can trigger the electronic device to display the video playback interface 31. The video playback interface 31 includes a first area 32 and a second area 33. The first area 32 displays the video frame of the processed travel video, and the second area 33 displays the video frame of the travel video. Then, when the user wants to view the original travel video, the user can click on the second area 33, as shown in Figure 6(B), triggering the electronic device to display the video frame of the travel video in the first area 32 and the video frame of the processed travel video in the second area 33.
[0149] Thus, on the one hand, since the video frame of the first video can be displayed simultaneously when the processed video frame is displayed, it is convenient for users to compare the video before and after processing; on the other hand, since the display area of the processed video frame and the display area of the original video frame can be switched by user input, it is convenient for users to view the original video in a larger display area.
[0150] In some embodiments of this application, the area sizes of the first region and the second region may also be the same.
[0151] In some embodiments of this application, such as Figure 7 As shown, after step 202 above, the video object removal method provided in this application embodiment further includes the following steps 701 and 702.
[0152] Step 701: The electronic device displays the video list interface.
[0153] In the embodiments of this application, the video list interface includes a first video thumbnail.
[0154] In some embodiments of this application, the aforementioned first video thumbnail corresponds to the first video and the processed first video.
[0155] Step 702: The electronic device responds to the user's third input to the first video thumbnail. If the input features of the third input match the first preset input features, the processed first video is played; if the input features of the third input match the second preset input features, the first video is played.
[0156] In some embodiments of this application, the electronic device can receive a third input from a user on a first video thumbnail, and in response to the third input, play the processed first video if the input features of the third input match a first preset input feature; and play the first video if the input features of the third input match a second preset input feature.
[0157] In some embodiments of this application, the third input described above is used to play the first video or the processed first video.
[0158] In some embodiments of this application, the aforementioned third input includes, but is not limited to: touch input of the user onto the first video thumbnail via a touch device such as a finger or stylus, or voice commands input by the user, or specific gestures input by the user, or click input, or other feasible inputs. The specific input can be determined according to actual usage needs, and this application does not limit this.
[0159] It is understandable that different input features of the third input will result in different videos being played.
[0160] For example, as shown in Figure 8(A), after obtaining the processed travel video, the electronic device can display a first video thumbnail 42 in the video list interface 41. This first video thumbnail 42 corresponds to the processed travel video and the original travel video. When the user wants to play the processed travel video, the user can slide upwards on the first video thumbnail 42 (i.e., the third input mentioned above). Then, if the input matches the first preset input feature, the electronic device will play the processed travel video. As shown in Figure 8(B), when the user wants to play the original travel video, the user can slide downwards on the first video thumbnail 42 (i.e., the third input mentioned above). Then, if the input matches the second preset input feature, the electronic device will play the original travel video.
[0161] Thus, since the video played will be different when the input characteristics of the third input are different, the flexibility of electronic devices in playing videos is improved.
[0162] In some embodiments of this application, after obtaining the processed first video, when playing the processed first video, the electronic device can receive a fifth input from the user on the processed first video and play the original first video.
[0163] In some embodiments of this application, the electronic device may associate and store the processed first video with the original first video.
[0164] In some embodiments of this application, when displaying the video playback interface of the first video, the electronic device can display processing controls on the video playback interface. The electronic device can receive user input to the processing controls, perform interference removal processing on the first video, and obtain the processed first video.
[0165] Specifically, the aforementioned processing controls may include a first control, a second control, a third control, and a fourth control. The first control is used to detect interfering objects in N keyframes; the second control is used to separate the layers of the object selected by the user; the third control is used to detect whether adjacent video frames contain interfering objects; and the fourth control is used to eliminate interfering objects. When the electronic device receives user input for the first control, it can perform interfering object detection on the N keyframes in the first video and obtain the detection result. If the detection result indicates that the first object in the first keyframe is an interfering object, the electronic device can display the first keyframe. Then, when the electronic device receives user input for the second control and the first object, it can overlay and display the first keyframe. A first floating layer includes a first object, and the display area of the first object includes a removal control. After receiving user input to the removal control, the electronic device can determine that the first object is an interference object and mark the first object on the first keyframe. Then, when receiving user input to a third control, the electronic device can detect whether the first object is contained in the first video frame corresponding to the first keyframe. If the first object is contained in the first video frame, a thumbnail of the first keyframe and the first video frame is displayed on the video playback interface. Then, when receiving user input to a fourth control, the electronic device can perform removal processing on the first object contained in the first keyframe and the first video frame to obtain the processed first video.
[0166] For example, the implementation of this application can be applied to the process of taking live photos, and similarly to removing passersby in travel videos and removing flying insects in surveillance videos.
[0167] It should be noted that the above-described method embodiments, or the various possible implementations of the method embodiments, can be executed individually, or, provided there are no contradictions, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.
[0168] It should be noted that the video object removal method provided in this application embodiment can be executed by a video object removal device. This application embodiment uses a video object removal device executing the video object removal method as an example to illustrate the video object removal device provided in this application embodiment.
[0169] Figure 9 A schematic diagram of a possible structure of the video object removal device involved in an embodiment of this application is shown. For example... Figure 9 As shown, the video object removal device 70 may include a detection module 71 and a processing module 72;
[0170] Among them, the detection module 71 is used to detect interference objects in N key frames of the first video and obtain the detection results;
[0171] The processing module 72 is used to perform elimination processing on the first key frame and the first object contained in the first video in the first video when the detection result obtained by the detection module 71 determines that the first object in the first key frame among N key frames is an interference object and the first video frame corresponding to the first key frame contains the first object, so as to obtain the processed first video; wherein, the first video frame is a non-key frame in the first video that is adjacent to the first key frame.
[0172] This application provides a video object removal device. By detecting interference objects in key frames of a first video, when it is determined that a certain key frame and its adjacent non-key frames contain interference objects, the interference objects in the key frame and its adjacent non-key frames are removed. That is, interference objects in the video are removed by narrowing the interference object detection range of the electronic device, avoiding interference object detection in all video frames of the first video, realizing rapid removal of interference objects in the video, thereby improving the efficiency of the electronic device in removing interference objects in the video.
[0173] In one possible implementation, the video object removal device 70 provided in this application embodiment further includes: an acquisition module. The acquisition module is used to acquire N RGB images corresponding to N keyframes, N depth images corresponding to N keyframes, and N motion vector data corresponding to N keyframes, where each motion vector data is used to characterize the motion parameters of at least one video object in the keyframe corresponding to each motion vector data. The detection module 71 is specifically used to perform interference object detection on the N keyframes based on the N RGB images, N depth images, and N motion vector data acquired by the acquisition module, and obtain detection results.
[0174] In one possible implementation, the video object removal apparatus 70 provided in this application embodiment further includes a display module and a determination module. The display module is configured to, after the detection module 71 performs interference object detection on N keyframes in the first video and obtains the detection result, display the first keyframe if the detection result indicates that the first object in the first keyframe is an interference object. The first keyframe carries a first identifier, which is used to mark the first object in the first keyframe as an interference object. The determination module is configured to, in response to a user's first input regarding the first object in the first keyframe displayed by the display module, determine that the first object is an interference object.
[0175] In one possible implementation, the video object removal device 70 provided in this application embodiment further includes a determination module. The determination module is configured to, after the detection module 71 performs interference object detection on N keyframes in the first video and obtains the detection result, if the detection result indicates that a second object in a second keyframe among the N keyframes is an interference object, determine that the second object is an interference object if a first distance is less than a second distance; or, determine that the second object is not an interference object if the first distance is greater than or equal to the second distance; wherein the first distance is the distance between the second object in the second keyframe and the shooting subject in the second keyframe, the second distance is the distance between the second object in the second video frame and the shooting subject in the second video frame, the second video frame is a video frame in the first video adjacent to the second keyframe, and the shooting time of the second video frame is later than the shooting time of the second keyframe.
[0176] In one possible implementation, the video object removal device 70 provided in this application embodiment further includes: an acquisition module and a calculation module. The acquisition module is used to acquire the color histogram corresponding to each video frame in the first video before the detection module 71 performs interference object detection on N keyframes in the first video and obtains the detection result. The calculation module is used to calculate the difference between the color histograms corresponding to every two adjacent video frames in the first video. The determination module is further used to, if the calculation module calculates that the difference between the color histogram corresponding to the third video frame and the color histogram corresponding to the fourth video frame is greater than or equal to a preset threshold, identify the fourth video frame as a keyframe in the first video; wherein the third video frame and the fourth video frame are two adjacent video frames in the first video, and the shooting time of the fourth video frame is later than the shooting time of the third video frame.
[0177] In one possible implementation, the video object removal device 70 provided in this application embodiment further includes a display module. The display module is configured to, after the processing module 72 performs removal processing on the first keyframe and the first object contained in the first video frame to obtain the processed first video, display a video playback interface. The video playback interface includes a first area and a second area. The first area displays the video frame of the processed first video, and the second area displays the video frame of the first video. The display area of the first area is larger than the display area of the second area. In response to a second input from the user to the second area, the display module displays the video frame of the first video in the first area and the processed video frame of the first video in the second area.
[0178] In one possible implementation, the video object removal device 70 provided in this application embodiment further includes a display module and a playback module. The display module is used to display a video list interface, which includes thumbnails of the first video, after the processing module 72 performs removal processing on the first keyframe and the first object contained in the first video to obtain the processed first video. The playback module is used to respond to a third input from the user to the thumbnail of the first video displayed by the display module, playing the processed first video if the input features of the third input match a first preset input feature; and playing the first video if the input features of the third input match a second preset input feature.
[0179] The video object removal device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0180] The video object removal device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0181] The video object removal device provided in this application embodiment can implement all the processes implemented in the above method embodiments, and will not be described again here to avoid repetition.
[0182] Optionally, such as Figure 10As shown, this application embodiment also provides an electronic device 900, including a processor 901 and a memory 902. The memory 902 stores a program or instructions that can run on the processor 901. When the program or instructions are executed by the processor 901, they implement the various steps of the above method embodiments and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0183] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0184] Figure 11 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0185] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.
[0186] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 11 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0187] The processor 110 is used to detect interference objects in N key frames of the first video and obtain detection results; and if it is determined from the detection results that the first object in the first key frame among the N key frames is an interference object and the first video frame corresponding to the first key frame contains the first object, the processor performs elimination processing on the first key frame and the first object contained in the first video frame to obtain the processed first video; wherein the first video frame is a non-key frame in the first video that is adjacent to the first key frame.
[0188] This application provides an electronic device that performs interference object detection on key frames of a first video. When it is determined that a certain key frame and its adjacent non-key frames contain interference objects, the interference objects in the key frame and its adjacent non-key frames are eliminated. That is, interference objects in the video are eliminated by narrowing the interference object detection range of the electronic device, avoiding interference object detection on all video frames of the first video, realizing the rapid elimination of interference objects in the video, thereby improving the efficiency of the electronic device in eliminating interference objects in the video.
[0189] In some embodiments of this application, the processor 110 is specifically configured to acquire N RGB images corresponding to N key frames, N depth images corresponding to N key frames, and N motion vector data corresponding to N key frames, wherein each motion vector data is used to characterize the motion parameters of at least one video object in the key frame corresponding to each motion vector data; and based on the N RGB images, N depth images, and N motion vector data, to perform interference object detection on the N key frames and obtain detection results.
[0190] In some embodiments of this application, the display unit 106 is used to display the first key frame after the processor 110 performs interference object detection on N key frames in the first video and obtains the detection result, if the detection result indicates that the first object in the first key frame is an interference object. The first key frame carries a first identifier, which is used to mark the first object in the first key frame as an interference object.
[0191] The processor 110 is also configured to determine, in response to a first input from the user to a first object in the first keyframe, that the first object is an interfering object.
[0192] In some embodiments of this application, the processor 110 is further configured to, after detecting interference objects in N keyframes of the first video and obtaining the detection results, determine that the second object is an interference object if the first distance is less than the second distance, or if the first distance is greater than or equal to the second distance, determine that the second object is not an interference object; wherein, the first distance is the distance between the second object in the second keyframe and the shooting subject in the second keyframe, the second distance is the distance between the second object in the second video frame and the shooting subject in the second video frame, the second video frame is the video frame in the first video adjacent to the second keyframe, and the shooting time of the second video frame is later than the shooting time of the second keyframe.
[0193] In some embodiments of this application, the processor 110 is further configured to, before performing interference object detection on N key frames in the first video and obtaining the detection result, acquire a color histogram corresponding to each video frame in the first video; calculate the difference between the color histograms corresponding to every two adjacent video frames in the first video; and, if the difference between the color histogram corresponding to the third video frame and the color histogram corresponding to the fourth video frame is greater than or equal to a preset threshold, use the fourth video frame as a key frame in the first video; wherein the third video frame and the fourth video frame are two adjacent video frames in the first video, and the shooting time of the fourth video frame is later than the shooting time of the third video frame.
[0194] In some embodiments of this application, the display unit 106 is configured to display a video playback interface after the processor 110 performs elimination processing on the first keyframe and the first object contained in the first video frame to obtain the processed first video. The video playback interface includes a first area and a second area. The first area displays the video frame of the processed first video, and the second area displays the video frame of the first video. The display area of the first area is larger than the display area of the second area. In response to a second input from the user to the second area, the display unit 106 displays the video frame of the first video in the first area and the video frame of the processed first video in the second area.
[0195] In some embodiments of this application, the display unit 106 is used to display a video list interface after the processor 110 performs elimination processing on the first keyframe and the first object contained in the first video frame to obtain the processed first video. The video list interface includes a thumbnail of the first video.
[0196] The processor 110 is also configured to, in response to a third input from a user to a first video thumbnail displayed on the display unit 106, play the processed first video if the input features of the third input match a first preset input feature, and play the first video if the input features of the third input match a second preset input feature.
[0197] The electronic device provided in this application embodiment can implement the various processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0198] The beneficial effects of the various implementation methods in this embodiment can be found in the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, they will not be repeated here.
[0199] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.
[0200] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0201] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.
[0202] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0203] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0204] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0205] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0206] This application provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here.
[0207] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0208] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0209] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for eliminating video objects, characterized in that, The method includes: Interference object detection is performed on N keyframes in the first video to obtain the detection results, where N is an integer greater than 1; If, based on the detection results, it is determined that the first object in the first key frame among the N key frames is an interference object, and the first video frame corresponding to the first key frame contains the first object, then the first key frame in the first video and the first object contained in the first video frame are eliminated to obtain the processed first video. Wherein, the first video frame is a non-key frame in the first video that is adjacent to the first key frame.
2. The method according to claim 1, characterized in that, The detection of interference objects in N keyframes of the first video, and the resulting detection results, include: Obtain N RGB images corresponding to the N keyframes, N depth images corresponding to the N keyframes, and N motion vector data corresponding to the N keyframes. Each motion vector data is used to characterize the motion parameters of at least one video object in the keyframe corresponding to each motion vector data. Based on the N RGB images, the N depth images, and the N motion vector data, interference object detection is performed on the N keyframes to obtain the detection results.
3. The method according to claim 1 or 2, characterized in that, After detecting interference objects in N keyframes of the first video and obtaining the detection results, the method further includes: If the detection result indicates that the first object in the first keyframe is a interference object, the first keyframe is displayed, and the first keyframe carries a first identifier, which is used to mark the first object in the first keyframe as a interference object. In response to a user’s first input on the first object in the first keyframe, the first object is determined to be an interference object.
4. The method according to claim 1 or 2, characterized in that, After detecting interference objects in N keyframes of the first video and obtaining the detection results, the method further includes: If the detection result indicates that the second object in the second keyframe among the N keyframes is an interfering object, then if the first distance is less than the second distance, the second object is determined to be an interfering object; or, if the first distance is greater than or equal to the second distance, the second object is determined not to be an interfering object. Wherein, the first distance is the distance between the second object in the second keyframe and the shooting subject in the second keyframe, the second distance is the distance between the second object in the second video frame and the shooting subject in the second video frame, the second video frame is the video frame in the first video that is adjacent to the second keyframe, and the shooting time of the second video frame is later than the shooting time of the second keyframe.
5. The method according to claim 1 or 2, characterized in that, Before performing interference object detection on N keyframes in the first video and obtaining the detection results, the method further includes: Obtain the color histogram corresponding to each video frame in the first video; Calculate the difference between the color histograms corresponding to every two adjacent video frames in the first video; If the difference between the color histogram corresponding to the third video frame and the color histogram corresponding to the fourth video frame is greater than or equal to a preset threshold, the fourth video frame is used as a key frame in the first video. The third video frame and the fourth video frame are two adjacent video frames in the first video, and the fourth video frame was captured later than the third video frame.
6. The method according to claim 1, characterized in that, After removing the first keyframe and the first object contained in the first video frame to obtain the processed first video, the method further includes: The video playback interface includes a first area and a second area. The first area displays the video frame of the processed first video, and the second area displays the video frame of the first video. The display area of the first area is larger than the display area of the second area. In response to a second input from the user to the second area, the video frame of the first video is displayed in the first area, and the video frame of the processed first video is displayed in the second area.
7. The method according to claim 1, characterized in that, After removing the first keyframe and the first object contained in the first video frame to obtain the processed first video, the method further includes: The video list interface includes a first video thumbnail; In response to a third input from the user to the first video thumbnail, if the input features of the third input match a first preset input feature, the processed first video is played; if the input features of the third input match a second preset input feature, the first video is played.
8. A video object removal device, characterized in that, The device includes: a detection module and a processing module; The detection module is used to detect interference objects in N keyframes of the first video and obtain the detection results. The processing module is used to perform elimination processing on the first keyframe and the first object contained in the first video in the first video when the detection result obtained by the detection module determines that the first object in the first keyframe among the N keyframes is an interference object and the first video frame corresponding to the first keyframe contains the first object, so as to obtain the processed first video. Wherein, the first video frame is a non-key frame in the first video that is adjacent to the first key frame.
9. The apparatus according to claim 8, characterized in that, The device further includes: an acquisition module; The acquisition module is used to acquire N RGB images corresponding to the N key frames, N depth images corresponding to the N key frames, and N motion vector data corresponding to the N key frames. Each motion vector data is used to characterize the motion parameters of at least one video object in the key frame corresponding to each motion vector data. The detection module is specifically used to perform interference object detection on the N keyframes based on the N RGB images, the N depth images, and the N motion vector data acquired by the acquisition module, and obtain the detection results.
10. The apparatus according to claim 8 or 9, characterized in that, The device further includes: a display module and a determination module; The display module is used to display the first key frame after the detection module performs interference object detection on N key frames in the first video and obtains the detection result, when the detection result indicates that the first object in the first key frame is an interference object. The first key frame carries a first identifier, and the first identifier is used to mark the first object in the first key frame as an interference object. The determining module is configured to determine the first object as an interference object in response to a first input from the user to the first object in the first keyframe displayed by the display module.
11. The apparatus according to claim 8 or 9, characterized in that, The device further includes: a determining module; The determining module is configured to, after the detection module performs interference object detection on N keyframes in the first video and obtains the detection result, determine the second object as an interference object if the first distance is less than the second distance when the detection result indicates that the second object in the second keyframe among the N keyframes is an interference object; or, if the first distance is greater than or equal to the second distance, determine the second object as not an interference object. Wherein, the first distance is the distance between the second object in the second keyframe and the shooting subject in the second keyframe, the second distance is the distance between the second object in the second video frame and the shooting subject in the second video frame, the second video frame is the video frame in the first video that is adjacent to the second keyframe, and the shooting time of the second video frame is later than the shooting time of the second keyframe.
12. The apparatus according to claim 8 or 9, characterized in that, The device further includes: an acquisition module and a calculation module; The acquisition module is used to acquire the color histogram corresponding to each video frame in the first video before the detection module performs interference object detection on N key frames in the first video and obtains the detection result. The calculation module is used to calculate the difference between the color histograms corresponding to every two adjacent video frames in the first video; The determining module is further configured to, when the calculation module calculates that the difference between the color histogram corresponding to the third video frame and the color histogram corresponding to the fourth video frame is greater than or equal to a preset threshold, use the fourth video frame as a key frame in the first video. The third video frame and the fourth video frame are two adjacent video frames in the first video, and the fourth video frame was captured later than the third video frame.
13. The apparatus according to claim 8, characterized in that, The device further includes: a display module; The display module is configured to display a video playback interface after the processing module performs elimination processing on the first keyframe and the first object contained in the first video frame in the first video to obtain the processed first video. The video playback interface includes a first area and a second area. The first area displays the video frame of the processed first video, and the second area displays the video frame of the first video. The display area of the first area is larger than the display area of the second area. In response to a second input from the user to the second area, the display module displays the video frame of the first video in the first area and the video frame of the processed first video in the second area.
14. The apparatus according to claim 8, characterized in that, The device further includes: a display module and a playback module; The display module is used to display a video list interface after the processing module performs elimination processing on the first keyframe and the first object contained in the first video frame in the first video to obtain the processed first video. The video list interface includes a thumbnail of the first video. The playback module is configured to respond to a third input from the user to the thumbnail of the first video displayed by the display module, and to play the processed first video if the input features of the third input match a first preset input feature; and to play the first video if the input features of the third input match a second preset input feature.