Video processing method and device, storage medium and electronic equipment
By automatically generating video prompt words and cyclic shift strategies, and combining diffusion model to process video sequences, the problem of insufficient time and space domain consistency in traditional video processing technology is solved, and efficient completion of video after target removal in complex dynamic scenarios is achieved.
Patent Information
- Application Number
- CN202510482975.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-22
AI Technical Summary
When traditional video processing technology removes the target and completes the video, especially in complex dynamic scenarios, it cannot effectively maintain the consistency of time and space.
By automatically generating video prompt words that do not contain objects to be removed, combining cyclic shift strategies and diffusion models (such as Diffusion Transformer), the masked video sequence and video to be completed are processed, and video completion is completed in batches.
After removing the target, it can maintain the spatial and temporal consistency of the video in complex dynamic scenarios, reduce manual participation, and improve the alignment efficiency and accuracy of video content and descriptive text.
Smart Images

Figure CN120358318A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision and generative artificial intelligence technology. More specifically, it relates to a video processing method, device, storage medium, and electronic device. Background Art
[0002] Video processing has a wide range of applications in modern life, such as dynamic object removal in film production, obstacle cleaning in surveillance footage, and video denoising in historical video restoration.
[0003] However, traditional video processing techniques (such as optical flow methods or GAN-based methods) have limited performance in compensating for the semantic and temporal consistency of videos after object removal. Especially in complex dynamic scenes, these methods cannot maintain good spatio-temporal consistency.
[0004] Therefore, how to maintain spatio-temporal consistency during the process of compensating for videos after object removal is an urgent problem to be solved in this application. Summary of the Invention
[0005] In view of this, this application discloses a video processing method, device, storage medium, and electronic device, aiming to maintain spatio-temporal consistency during the process of compensating for videos after object removal.
[0006] To achieve the above objective, the disclosed technical solutions are as follows:
[0007] The first aspect of this application discloses a video processing method, and the method includes:
[0008] Obtain the name of the object to be removed and the target video prompt; wherein, the target video prompt is a video text prompt that does not include the object to be removed;
[0009] According to the name of the object to be removed and the target video prompt, obtain a masked video sequence and a video to be completed; wherein, the video to be completed is a video in which the object to be removed is covered with a mask;
[0010] Process the masked video sequence and the video to be completed corresponding to the video with the target object to be removed through a cyclic shift strategy to obtain video sequences to be completed for each batch;
[0011] Perform video completion on the video sequences to be completed for each batch through a preset video completion method.
[0012] Preferably, the obtaining the name of the object to be removed and the target video prompt includes:
[0013] Respond to user requirements to determine the video with the target object to be removed;
[0014] Obtain the names of each object and the video prompt words included in the video to be target-removed through the first-round question-and-answer information of the video understanding model;
[0015] Determine the name of the object to be removed from the names of each object;
[0016] Obtain the target video prompt words through the second-round question-and-answer information of the video understanding model; wherein, the second-round question-and-answer information is determined by the name of the object to be removed and the video prompt words.
[0017] Preferably, obtaining the masked video sequence and the video to be completed according to the name of the object to be removed and the target video prompt words includes:
[0018] Perform target segmentation on the video to be target-removed according to the name of the object to be removed to obtain a masked video sequence;
[0019] Determine the video to be completed according to the target video prompt words and the masked video sequence.
[0020] Preferably, processing the masked video sequence and the video to be completed corresponding to the video to be target-removed through a cyclic shift strategy to obtain each batch of video sequences to be completed includes:
[0021] Determine the video to be target-removed as the source video, and fix the source video frames of the source video to keep the source video frames unchanged;
[0022] Under the condition of keeping the source video frames unchanged, reverse-copy the frames of the source video that are not fixed in the video to be target-removed to obtain a copied video; wherein, the source video frames are from the 0th frame to the nth frame; the reverse copy is a copy method from the (n - 1)th frame to the 1st frame;
[0023] Stitch the source video and the copied video to obtain a cyclic video;
[0024] Split the cyclic video into each batch of video sequences to be completed according to preset processing parameters.
[0025] Preferably, video-completing each batch of video sequences to be completed through a preset video completion method includes:
[0026] For the first video segment in each batch of video sequences to be completed, perform video completion on the first video segment through a video diffusion module;
[0027] For the non-first video segments in each batch of video sequences to be completed, obtain the shifted number of frames of the video completion result of the previous batch;
[0028] Use the frame corresponding to the number of shifted frames as the starting part, and splice the starting part with the video sequence to be completed in the current batch in the time domain until all non-first video segments in the video sequences to be completed in each batch are spliced, so as to complete the video completion process of the video sequences to be completed in each batch.
[0029] The second aspect of the present application discloses a video processing device, the device includes:
[0030] A first acquisition unit, configured to acquire the name of the object to be removed and the target video prompt word; wherein, the target video prompt word is a video text prompt word that does not include the object to be removed;
[0031] A second acquisition unit, configured to acquire a masked video sequence and a video to be completed according to the name of the object to be removed and the target video prompt word; wherein, the video to be completed is a video in which the object to be removed is covered with a mask;
[0032] A processing unit, configured to process the masked video sequence and the video to be completed corresponding to the video to be removed by the target through a cyclic shift strategy to obtain video sequences to be completed in each batch;
[0033] A video completion unit, configured to complete the video of the video sequences to be completed in each batch through a preset video completion method.
[0034] Preferably, the first acquisition unit includes:
[0035] A response module, configured to respond to user requirements to determine the video to be removed by the target;
[0036] A first acquisition module, configured to acquire the names of each object and video prompt words included in the video to be removed by the target through the first-round question-and-answer information of the video understanding model;
[0037] A first determination module, configured to determine the name of the object to be removed from the names of each object;
[0038] A second acquisition module, configured to acquire the target video prompt word through the second-round question-and-answer information of the video understanding model; wherein, the second-round question-and-answer information is determined by the name of the object to be removed and the video prompt word.
[0039] Preferably, the second acquisition unit includes:
[0040] A segmentation module, configured to perform target segmentation on the video to be removed by the target according to the name of the object to be removed to obtain a masked video sequence;
[0041] A second determination module, configured to determine the video to be completed according to the target video prompt word and the masked video sequence.
[0042] In a third aspect of the present application, a storage medium is disclosed. The storage medium includes stored instructions, wherein when the instructions run, the device where the storage medium is located is controlled to execute the video processing method according to any one of the first aspects.
[0043] In a fourth aspect of the present application, an electronic device is disclosed, including a memory, and one or more instructions, wherein one or more instructions are stored in the memory and are configured to be executed by one or more processors to execute the video processing method according to any one of the first aspects.
[0044] As can be seen from the above technical solutions, the present application discloses a video processing method, device, storage medium and electronic device. The name of the object to be removed and the target video prompt word are obtained, wherein the target video prompt word is a video text prompt word that does not include the object to be removed. According to the name of the object to be removed and the target video prompt word, a masked video sequence and a video to be completed are obtained, wherein the video to be completed is a video in which the object to be removed is covered with a mask. The masked video sequence and the video to be completed corresponding to the target video to be removed are processed by a cyclic shift strategy to obtain video sequences to be completed for each batch, and the video sequences to be completed for each batch are video-completed by a preset video completion method.
[0045] Through the above solution, a video prompt description that does not include the object to be removed, that is, the target video prompt word, is automatically generated, reducing manual participation and ensuring a more efficient and accurate alignment between the video content and the descriptive text. A cyclic shift strategy is proposed during inference. Since the entire video is continuous, there will be no temporal inconsistency between the front and back frames during inference through the cyclic shift strategy. Therefore, under the current automatic generation of the target video prompt word and the cyclic shift strategy, by using the preset video completion method of the diffusion transformer, the video sequences to be completed for each batch are video-completed, and even in a complex dynamic scene, the purpose of maintaining spatio-temporal consistency during the process of completing the video after removing the target can still be achieved. Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0047] Figure 1 It is a schematic flowchart of a video processing method disclosed in an embodiment of the present application;
[0048] Figure 2 Schematic diagram of the process for obtaining the name of the object to be removed and the target video prompt disclosed in the embodiments of the present application;
[0049] Figure 3 Schematic diagram of the cyclic shift strategy disclosed in the embodiments of the present application;
[0050] Figure 4 Schematic diagram of the process for video completion of each batch of videos to be completed by a preset video completion method disclosed in the embodiments of the present application;
[0051] Figure 5 Schematic diagram of the entire video processing disclosed in the embodiments of the present application;
[0052] Figure 6 Schematic diagram of the structure of a video processing device disclosed in the embodiments of the present application;
[0053] Figure 7 Schematic diagram of the structure of an electronic device disclosed in the embodiments of the present application. Detailed implementation manners
[0054] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0055] In the present application, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0056] As can be seen from the background art, traditional video processing technologies (such as optical flow method or GAN-based method) are limited in terms of semantic and temporal consistency of video completion after removing the target. Especially in complex dynamic scenes, these methods cannot well maintain the spatio-temporal domain consistency. Therefore, how to maintain the spatio-temporal domain consistency in the process of video completion after removing the target is an urgent problem to be solved in the present application.
[0057] To solve the above problems, the present application discloses a video processing method, apparatus, storage medium, and electronic device, which automatically generate video prompt descriptions that do not include objects to be removed, that is, target video prompt words, reducing manual participation and ensuring more efficient and accurate alignment between video content and descriptive text. A cyclic shift strategy is proposed during inference. Since the entire video is continuous, there will be no temporal inconsistency between the front and back frames during inference through the cyclic shift strategy. Therefore, under the current automatic generation of target video prompt words and cyclic shift strategy, by using the preset video completion method of the diffusion model to perform video completion on each batch of video sequences to be completed, even in complex dynamic scenarios, the purpose of maintaining spatio-temporal consistency during the process of completing the video after removing the target can still be achieved. The specific implementation manner will be specifically described in the following embodiments.
[0058] Reference Figure 1 As shown, a video processing method disclosed in an embodiment of the present application mainly includes the following steps:
[0059] S101: Obtain the name of the object to be removed and the target video prompt (prompt); wherein, the target video prompt is a video text prompt that does not include the object to be removed.
[0060] In S101, the name of the object to be removed and the target video prompt can be obtained through a video understanding model.
[0061] Among them, the video understanding model is a module unit used to understand the video, output the objects included in the video, and perform prompt description on the video, that is, module A below Figure 5 in.
[0062] Automatically generate a prompt description of the video that does not include the object to be removed through the video understanding model, reducing manual participation, thereby ensuring more efficient and accurate alignment between video content and descriptive text.
[0063] The name of the object to be removed is determined according to user needs.
[0064] The specific process of obtaining the name of the object to be removed and the target video prompt is as Figure 2 shown.
[0065] S201: Respond to user needs to determine the target video to be removed.
[0066] It should be noted that the video to be targeted for removal is determined according to the user's needs, and it contains video segments with the objects to be removed (such as people, animals, etc.). For example, if there is a video containing a person and the user's requirement is to remove this person, then this video can be called the video to be targeted for removal, that is, the video that needs to be processed and has the target to be removed.
[0067] S202: Obtain the names of each object and the video prompt included in the video to be targeted for removal through the first-round Q&A information of the video understanding model.
[0068] Among them, the video prompt refers to the video description, that is, the description of the content shown in this video.
[0069] For example, the first-round Q&A information is: Question: What objects are included in the video and what kind of scene is shown; Answer: Girl, skirt, high heels, church, white bridge. This scene depicts a girl standing on a white rock. That is, the names of each object included in the video to be targeted for removal are girl, skirt, high heels, church, and white bridge; the video prompt is that this scene depicts a girl standing on a white rock.
[0070] S203: Determine the name of the object to be removed from the names of each object.
[0071] In S203, in response to the object to be removed selected by the user, the name of the object to be removed corresponding to it is determined accordingly.
[0072] S204: Obtain the target video prompt through the second-round Q&A information of the video understanding model; among them, the second-round Q&A information is determined by the name of the object to be removed and the video prompt.
[0073] Among them, the target video prompt is the video description that does not include the object to be removed.
[0074] For example, the second-round Q&A information is: Question: Remove the girl and describe the objects and the scene; Answer: White fence, walkway, building with arches and doors. This scene shows an elegant building entrance passage without anyone; the name of the object to be removed is girl. Then the target video prompt is to remove the girl, white fence, walkway, building with arches and doors. This scene shows an elegant building entrance passage without anyone.
[0075] Through the determined name of the object to be removed and the video prompt in S202 above, a second-round conversation is carried out, that is, through the second-round Q&A information, the target video prompt that does not include the object to be removed is obtained.
[0076] S102: Obtain the mask video sequence and the video to be completed according to the name of the object to be removed and the target video prompt; wherein, the video to be completed is the video with the object to be removed covered by the mask.
[0077] In S102, perform target segmentation on the target removal video according to the name of the object to be removed to obtain the mask video sequence, and determine the video to be completed according to the target video prompt and the mask video sequence.
[0078] It should be noted that the video to be completed cannot contain the object to be removed.
[0079] S103: Process the mask video sequence corresponding to the target removal video and the video to be completed through a cyclic shift strategy to obtain the video sequence to be completed for each batch.
[0080] In S103, send the video segment to be completed and the mask video sequence for each batch to the video diffusion (video diffusion transformer). The video diffusion transformer is an algorithm module for completing the mask area of the video to be completed.
[0081] Among them, the video sequence to be completed for each batch refers to multiple batches of video segments to be completed.
[0082] Due to hardware limitations, the algorithm cannot process the entire long video sequence and needs to be processed batch by batch. Therefore, the mask video sequence needs to be divided into multiple batches (batches) for processing. Each batch processes a fixed number of frames F. To better ensure the spatio-temporal consistency of multiple batches, the present application adopts a cyclic shift strategy.
[0083] Spatio-temporal consistency refers to the consistency of spatial objects, the consistency of the background, and the consistency of objects and the background in different video frames in terms of time. For example, in the videos shot in reality that we see, it can be clearly seen that the objects in the continuous videos do not deform or disappear, and the background does not deform or disappear. It is usually expressed as spatio-temporal consistency.
[0084] The cyclic shift strategy proposes a cyclic shift strategy during inference. Cyclic shift is because the entire video is continuous and there will be no temporal inconsistency between the front and back frames during inference, which is used to ensure the spatio-temporal consistency of long-sequence videos. The cyclic shift strategy is as Figure 3 shown.
[0085] Figure 3Among them, (1) Determine the target-removal video to be processed as the source video (i.e., the yellow video sequence in the figure), and fix the frames of the source video (V_ori) in the target-removal video to keep the source video frames unchanged, with the frame numbers ranging from frame 0 to frame n; the video remains unchanged from frame 0 to frame n, and then reverse copying is performed at the end of this video, that is Figure 3 from frame n - 1 to frame 1 in the second row; n is an integer greater than or equal to 1;
[0086] (2) Under the condition of keeping the source video frames unchanged, reverse copy the frames of the non-fixed source video in the target-removal video, that is, copy its frames from (n - 1) to 1, to obtain a copied video (i.e., the reverse-order video, i.e., the gray-blue video sequence in the figure), and name this copied video V_inv; the reverse copying is in the way of copying from frame n - 1 to frame 1;
[0087] (3) Circular shift: Concatenate the source video V_ori and the copied video V_inv to obtain a loopable video V_cycle; Split V_cycle into batch video sequences to be completed according to preset processing parameters (fixed number of frames F and shift S processed in each batch), and process one segment at a time, where the shift S belongs to hyperparameters and is a preset fixed value.
[0088] S104: Complete video for each batch of video sequences to be completed through a preset video completion method.
[0089] It should be noted that video completion is to fill out the holes in a video with holes.
[0090] The preset video completion method includes but is not limited to the Diffusion Transformer (DiT) video completion method. The preferred preset video completion method in this application is the DiT video completion method.
[0091] Based on DiT video completion, use the 3D full attention of DiT to maintain the spatio-temporal domain consistency of the current video segment to be completed, and at the same time automatically generate background textures to generate a more natural and high-quality video picture.
[0092] The specific process of completing video for each batch of video sequences to be completed through the preset video completion method is as Figure 4 shown.
[0093] S401: For the first video segment in each batch of video sequences to be completed, complete the video for the first video segment through a video diffusion module.
[0094] S402: For the non-first video segments in each batch of video sequences to be completed, obtain the shifted number of frames of the video completion result of the previous batch.
[0095] In S402, for non-first video segments, it is necessary to use the result that has been completed in the previous batch to obtain the shifted frame number (S frame) as the starting part, and splice the video sequence to be completed in the current batch in the time domain, and so on to obtain the completed result of the entire video.
[0096] It should be noted that the shift of the first batch and each batch after the first batch in the video sequences to be completed in each batch is different.
[0097] S403: Use the frame corresponding to the shifted frame number as the starting part, and splice the starting part with the video sequence to be completed in the current batch in the time domain until all non-first video segments in the video sequences to be completed in each batch are spliced, so as to complete the video completion process of the video sequences to be completed in each batch.
[0098] Under the current automated prompt and circular shift strategy, video object removal and completion can obtain better long-term temporal consistency.
[0099] To facilitate the understanding of the above video processing process, combined with Figure 5 an example is given below Figure 5 which shows a schematic diagram of the entire video processing.
[0100] Figure 5 In, Step 1: The video to be target-removed obtained in response to the user's demand, through two rounds of question-and-answer automation of the video understanding model, obtains a video prompt that does not contain the object to be removed. The two rounds of question-and-answer of the video understanding model are Figure 5 as shown in Module A in
[0101] Specifically, the execution process of Module A is as follows:
[0102] (1) Obtain the name of the object to be removed and the video prompt included in the video through the first round of question-and-answer information; for example, the first round of question-and-answer information is: What objects are included in the video and what kind of scene is shown? Answer: Girl, skirt, high heels, church, white bridge. This scene depicts a girl standing on a white rock;
[0103] (2) In response to the object name given by the user, select the object to be removed from the object name;
[0104] (3) Given the name of the object to be removed in the second round of question-and-answer information, conduct a second-round conversation to obtain a target video prompt that does not contain the object to be removed; for example, the question in the second round of question-and-answer information is: Remove the girl and describe the object and the scene; Answer: White fence, walkway, building with arches and doors. This scene shows an elegant building entrance passage without anyone;
[0105] Step 2: In response to the name of the object to be removed given according to the user's needs, obtain the corresponding mask video sequence and the video to be completed according to the name of the object to be removed, as shown in Module B in Figure 5 ;
[0106] Step 3: Due to hardware limitations, the algorithm cannot directly process a long video all at once and needs to process it batch by batch. Here, batch refers to how many batches to divide for processing. Therefore, the mask video sequence needs to be divided into multiple batches for processing, and each batch processes a fixed number of frames F. To better ensure the spatio-temporal consistency of multiple batches, this application adopts a cyclic shift strategy; usually, a long video sequence cannot be evenly divisible by the fixed number of frames F. Therefore, this application proposes a cyclic shift strategy, that is, as shown in Figure 3 ; specifically as follows:
[0107] (1) Determine the video with the target to be removed as the source video, and fix the frames of the source video (V_ori) in the video with the target to be removed to keep the frames of the source video unchanged, and its frame numbers range from frame 0 to the nth frame; the video from frame 0 to the nth frame remains unchanged, and reverse copying is performed behind this video, that is, Figure 3 the (n - 1)th frame to the 1st frame in the second row of
[0108] (2) Under the condition of keeping the frames of the source video unchanged, perform reverse copying on the frames of the source video that are not fixed in the video with the target to be removed, that is, copy its (n - 1)-1 frames to obtain a copied video, and this copied video is named V_inv; the reverse copying is in the way of copying from the (n - 1)th frame to the 1st frame;
[0109] (3) Concatenate the source video V_ori and the copied video V_inv to obtain a loopable video V_cycle; split V_cycle into batch video sequences to be completed according to the preset processing parameters (each batch processes a fixed number of frames F and a shift S), and process one segment each time, where the shift S belongs to a hyperparameter and is a preset fixed value;
[0110] Step 4: Send the video segment to be completed for each batch, the mask video sequence, and the prompt encoding corresponding to the target video prompt (the prompt encoding is the video prompt encoding that does not include the object to be removed, that is, Figure 5 the prompt word features in Figure 5 ) into the video diffusion transformer (i.e., Module C in
[0111] It should be noted that since the video object removal task is to remove a specified object in the same scene, the target video prompt can represent the prompt for each segment.
[0112] Step Five: Perform video completion on each batch of sequences to be completed.
[0113] Among them, video completion is to fill out the holes in a video with holes. For example, Figure 5 the algorithm in Figure 5 fills out the black humanoid holes in the video to be completed, obtaining the final
[0114] The specific process of video completion is as follows:
[0115] (1) For the first video segment in each batch of sequences to be completed, perform video completion on the first video segment through the video diffusion transformer.
[0116] (2) For non-first segments in each batch of sequences to be completed, it is necessary to use the completed result of the previous batch, obtain its shifted frame number (S frame) as the starting part, and splice the sequences to be completed in the current batch in the time domain, and so on to obtain the completed result of the entire video.
[0117] The beneficial effects of this application: Automatically generate video prompt descriptions that do not contain the object to be removed, that is, target video prompt words, reducing manual participation and ensuring a more efficient and accurate alignment between video content and descriptive text. A cyclic shift strategy is proposed during inference. Since the entire video is continuous, there will be no temporal inconsistency between front and back frames during inference through the cyclic shift strategy. Therefore, under the current automatic generation of target video prompt words and cyclic shift strategy, by using the preset video completion method of the diffusion transformer to perform video completion on each batch of sequences to be completed, even in complex dynamic scenes, the purpose of maintaining spatio-temporal consistency during the process of completing the video after removing the target can still be achieved.
[0118] Based on the above embodiments Figure 1 disclosed a video processing method, an embodiment of the present application also correspondingly discloses a video processing device, as Figure 6 shown, the video processing device includes:
[0119] The first acquisition unit 601 is used to acquire the name of the object to be removed and the target video prompt word; among them, the target video prompt word is a video text prompt word that does not contain the object to be removed.
[0120] A second acquisition unit 602, configured to acquire a masked video sequence and a video to be completed according to the name of the object to be removed and the target video prompt word; wherein, the video to be completed is a video with the object to be removed covered by a mask.
[0121] A processing unit 603, configured to process the masked video sequence and the video to be completed corresponding to the target video to be removed through a cyclic shift strategy to obtain video sequences to be completed in each batch.
[0122] A video completion unit 604, configured to complete the video sequences to be completed in each batch through a preset video completion method.
[0123] Further, the first acquisition unit 601 includes:
[0124] A response module, configured to respond to user requirements to determine the target video to be removed.
[0125] A first acquisition module, configured to obtain the names of each object and the video prompt word included in the target video to be removed through the first-round Q&A information.
[0126] A first determination module, configured to determine the name of the object to be removed from the names of each object.
[0127] A second acquisition module, configured to obtain the target video prompt word through the second-round Q&A information of the video understanding model; wherein, the second-round Q&A information is determined by the name of the object to be removed and the video prompt word.
[0128] Further, the second acquisition unit 602 includes:
[0129] A segmentation module, configured to perform target segmentation on the target video to be removed according to the name of the object to be removed to obtain a masked video sequence.
[0130] A second determination module, configured to determine the video to be completed according to the target video prompt word and the masked video sequence.
[0131] Further, the processing unit 603 includes:
[0132] A fixation module, configured to determine the target video to be removed as the source video and fix the source video frames of the source video to keep the source video frames unchanged.
[0133] A reverse copying module, configured to reversely copy the frames of the non-fixed source video of the target video to be removed under the condition of keeping the source video frames unchanged to obtain a copied video; wherein, the source video frames are from frame 0 to frame n; the reverse copying is a copying method from frame n-1 to frame 1.
[0134] A first splicing module, configured to splice the source video and the copied video to obtain a cyclic video.
[0135] A splitting module, configured to split a loop video into video sequence batches to be completed according to preset processing parameters.
[0136] Furthermore, the video completion unit 604 includes:
[0137] A video completion module, configured to perform video completion on the first video segment in each batch of video sequences to be completed through a video diffusion module;
[0138] A third acquisition module, configured to acquire the shifted frame number of the video completion result of the previous batch for non-first video segments in each batch of video sequences to be completed;
[0139] A second splicing module, configured to use the frame corresponding to the shifted frame number as the starting part, and splice the starting part with the video sequence to be completed in the current batch in the time domain until all non-first video segments in each batch of video sequences to be completed are spliced, so as to complete the video completion process for each batch of video sequences to be completed.
[0140] The beneficial effects of the embodiments of the present application: Automatically generate video prompt descriptions that do not contain objects to be removed, that is, target video prompt words, reduce manual participation, and ensure more efficient and accurate alignment between video content and descriptive text. A cyclic shift strategy is proposed during inference. Since the entire video is continuous, there will be no temporal inconsistency between the front and back frames during inference through the cyclic shift strategy. Therefore, under the current automatic generation of target video prompt words and cyclic shift strategy, by using the preset video completion method of the diffusion transformer to complete the video for each batch of video sequences to be completed, even in complex dynamic scenes, the purpose of maintaining spatio-temporal consistency during the process of completing the video after removing the target can still be achieved.
[0141] The embodiments of the present application also provide a storage medium, which includes stored instructions. When the instructions run, they control the device where the storage medium is located to execute the video processing method as described above.
[0142] The embodiments of the present application also provide an electronic device, the structural schematic diagram of which is as Figure 7 shown, specifically including a memory 701 and one or more instructions 702. One or more instructions 702 are stored in the memory 701 and are configured to be executed by one or more processors 703 to execute the above video processing method.
[0143] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0144] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the partial description of the method embodiments for the relevant parts.
[0145] The steps in the methods of the embodiments of this application can be adjusted, combined, and deleted according to actual needs.
[0146] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or sequence between these entities or operations.
[0147] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the scope of this application. Therefore, this application will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0148] The above is only the preferred embodiment of this application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A video processing method, characterized in that, The method includes: Obtaining the name of the object to be removed and the target video prompt; wherein, the target video prompt is a video text prompt that does not include the object to be removed; According to the name of the object to be removed and the target video prompt, obtaining a masked video sequence and a video to be completed; wherein, the video to be completed is a video with the object to be removed covered by a mask; Processing the masked video sequence and the video to be completed corresponding to the video to be removed by the loop shift strategy to obtain video sequences to be completed in each batch; Completing the video sequences to be completed in each batch through a preset video completion method.
2. The method according to claim 1, characterized in that, The obtaining the name of the object to be removed and the target video prompt includes: Responding to user requirements to determine the video to be removed; Obtaining the names of each object and video prompts included in the video to be removed through the first-round Q&A information of the video understanding model; Determining the name of the object to be removed from the names of each object; Obtaining the target video prompt through the second-round Q&A information of the video understanding model; wherein, the second-round Q&A information is determined by the name of the object to be removed and the video prompt.
3. The method according to claim 2, wherein The obtaining the masked video sequence and the video to be completed according to the name of the object to be removed and the target video prompt includes: Performing target segmentation on the video to be removed according to the name of the object to be removed to obtain a masked video sequence; Determining the video to be completed according to the target video prompt and the masked video sequence.
4. The method according to claim 1, characterized in that, The processing the masked video sequence and the video to be completed corresponding to the video to be removed by the loop shift strategy to obtain video sequences to be completed in each batch includes: Determining the video to be removed as the source video and fixing the source video frames of the source video to keep the source video frames unchanged; Under the condition of keeping the source video frames unchanged, reversely copying the frames of the non-fixed source video of the video to be removed to obtain a copied video; wherein, the source video frames are from the 0th frame to the nth frame; the reverse copying is a copying method from the n-1th frame to the 1st frame; Splicing the source video and the copied video to obtain a loop video; Splitting the loop video into video sequences to be completed in each batch according to preset processing parameters.
5. The method according to claim 1, wherein The completing the video sequences to be completed in each batch through a preset video completion method includes: For the first video segment in each batch of video sequences to be completed, completing the first video segment through a video diffusion module; For the non-first video segments in each batch of video sequences to be completed, obtaining the shifted number of frames of the video completion result of the previous batch; Using the frames corresponding to the shifted number of frames as the starting part, splicing the starting part with the video sequence to be completed in the current batch in the time domain until all non-first video segments in each batch of video sequences to be completed are spliced, so as to complete the process of completing the video sequences to be completed in each batch.
6. A video processing device, characterized in that, The device includes: A first acquisition unit, configured to acquire the name of the object to be removed and the target video prompt; wherein, the target video prompt is a video text prompt that does not include the object to be removed. A second acquisition unit, configured to acquire a masked video sequence and a video to be completed according to the name of the object to be removed and the target video prompt; wherein, the video to be completed is a video in which the object to be removed is covered with a mask. A processing unit, configured to process the masked video sequence and the video to be completed corresponding to the target video to be removed through a cyclic shift strategy to obtain video sequences to be completed in each batch. A video completion unit, configured to complete the video sequences to be completed in each batch through a preset video completion method.
7. The device according to claim 6, characterized in that, The first acquisition unit includes: A response module, configured to respond to a user requirement to determine the target video to be removed. A first acquisition module, configured to obtain the names of the objects included in the target video to be removed and the video prompt through the first-round question-and-answer information of the video understanding model. A first determination module, configured to determine the name of the object to be removed from the names of the objects. A second acquisition module, configured to obtain the target video prompt through the second-round question-and-answer information of the video understanding model; wherein, the second-round question-and-answer information is determined by the name of the object to be removed and the video prompt.
8. The device according to claim 7, characterized in that, The second acquisition unit includes: A segmentation module, configured to perform target segmentation on the target video to be removed according to the name of the object to be removed to obtain a masked video sequence. A second determination module, configured to determine the video to be completed according to the target video prompt and the masked video sequence.
9. A storage medium, characterized in that, The storage medium includes stored instructions, wherein, when the instructions run, the device where the storage medium is located is controlled to execute the video processing method according to any one of claims 1 to 5.
10. An electronic device, characterized in that, It includes a memory, and one or more instructions, wherein one or more instructions are stored in the memory and are configured to be executed by one or more processors to execute the video processing method according to any one of claims 1 to 5.