Video processing method, video processing device and storage medium
By performing frame processing and multimodal feature extraction on the video, combined with large language model and diffusion model for image editing, the problems of difficult use and reduced clarity of existing video editing software are solved, and low-threshold and high-quality video editing effects are achieved.
Patent Information
- Application Number
- CN202311586113.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
The threshold for using existing video editing software is high, which makes it difficult for beginners to master quickly, and the clarity of video frames after editing decreases, affecting the video quality.
By performing frame processing on the video to be processed, a multi-frame video frame image is obtained, and the image to be processed is determined based on the received selection instruction. Receive image editing instructions, determine text description information and multimodal features of the pending image, process the pending image using a large language model and a diffusion model, generate a target image, and insert it into the pending video.
It lowers the technical threshold for video editing, enables users to quickly edit video content, improves the picture quality of edited content, and ensures the quality of edited videos.
Smart Images

Figure CN120050488A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and particularly to a video processing method, a video processing apparatus, and a storage medium. Background Art
[0002] Video editing technology refers to the technology of performing operations such as editing, color correction, and special effect processing on original video materials to produce a video work that meets the user's needs.
[0003] In related technologies, users can perform video editing through a variety of video editing software. However, video editing software requires users to have certain usage experience, with a high usage threshold, and beginners cannot quickly master it to obtain a video work that meets their own needs through video editing. Moreover, in related technologies, the clarity of the video frames edited by users will decrease, affecting the quality of the final video work. Summary of the Invention
[0004] To overcome the problems existing in related technologies, the present disclosure provides a video processing method, a video processing apparatus, and a storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, a video processing method is provided, including: performing frame division processing on a video to be processed to obtain multiple video frame images corresponding to the video to be processed, and determining a to-be-processed image from the multiple video frame images based on a received selection instruction; receiving an image editing instruction, determining text description information of the to-be-processed image, and determining multi-modal features of the to-be-processed image, where the image editing instruction is text information; processing the to-be-processed image according to the image editing instruction, the text description information, and the multi-modal features to obtain a target image; and inserting the target image into the video to be processed to obtain a target video.
[0006] In an implementation manner, determining the multi-modal features of the to-be-processed image includes: obtaining a content segmentation image, a depth segmentation image, and a masking image corresponding to the to-be-processed image, and obtaining adjacent frame images of the to-be-processed image from the multiple video frame images; and determining the multi-modal features of the to-be-processed image according to the content segmentation image, the depth segmentation image, the masking image, and the adjacent frame images.
[0007] In an implementation manner, obtaining the content segmentation image corresponding to the to-be-processed image includes: determining multiple target subjects included in the to-be-processed image; performing image segmentation on the to-be-processed image according to the multiple target subjects to make the image regions where the multiple target subjects are located independent; and determining the to-be-processed image after completion of segmentation as the content segmentation image.
[0008] In one implementation, obtaining the depth segmentation image corresponding to the image to be processed includes: determining the depth image corresponding to the image to be processed, where the gray value of a pixel in the depth image represents the depth information of the pixel; segmenting the depth image according to the texture features and structural features in the depth image to make multiple gray-scale continuous regions of the depth image independent; and determining the segmented depth image as the depth segmentation map.
[0009] In one implementation, obtaining the masking image corresponding to the image to be processed includes: performing image masking processing on each of the multiple image regions included in the content segmentation image or the depth segmentation image, and determining the multiple images obtained by performing the masking processing multiple times as the masking image; where the image masking processing includes: retaining a single image region and masking the other multiple image regions.
[0010] In one implementation, obtaining the adjacent frame images of the image to be processed from the multiple video frame images includes: obtaining multiple video frame images with timings before the image to be processed and multiple video frame images with timings after the image to be processed according to the timings of the multiple video frame images; and determining the multiple video frame images with timings before the image to be processed and the multiple video frame images with timings after the image to be processed as the adjacent frame images.
[0011] In one implementation, determining the multi-modal features of the image to be processed according to the content segmentation image, the depth segmentation image, the masking image, and the adjacent frame images includes: performing unified feature encoding processing on the segmentation image, the depth segmentation image, the masking image, and the adjacent frame images using the same fusion feature encoder to obtain the image features corresponding to the segmentation image, the depth segmentation image, the masking image, and the adjacent frame images respectively; and performing feature fusion processing on the image features corresponding to the segmentation image, the depth segmentation image, the masking image, and the adjacent frame images respectively through a feature fusion module, and determining the fusion features obtained by processing the feature fusion module as the multi-modal features.
[0012] In one implementation, processing the image to be processed according to the image editing instruction, the text description information, and the multi-modal features to obtain a target image includes: inputting the image editing instruction and the text description information into a large language model to obtain a semantic embedding layer, where the semantic embedding layer contains the image features of the target image; performing noise addition processing on the image to be processed and adding the multi-modal features to the image to be processed to obtain a multi-modal image; and processing the semantic embedding layer and the multi-modal image through a diffusion model and an encoder to obtain a target image.
[0013] In one implementation, processing the semantic embedding layer and the multi-modal image through a diffusion model and an encoder to obtain a target image includes: obtaining a semantic vector corresponding to the semantic embedding layer through the diffusion model, and gradually removing noise from the multi-modal image according to the semantic vector to obtain a feature conversion image; through the encoder, adjusting the feature conversion image according to the image size of the image to be processed, and determining the feature conversion image after the size adjustment as the target image.
[0014] In one implementation, inserting the target image into the video to be processed to obtain a target video includes: replacing the video frame image corresponding to the processed image in the video to be processed with the target image to obtain the target video.
[0015] According to a second aspect of the embodiments of the present disclosure, there is provided a video processing apparatus, including: a selection unit configured to perform frame division on a video to be processed, obtain multiple video frame images corresponding to the video to be processed, and determine an image to be processed from the multiple video frame images based on a received selection instruction; a determination unit configured to receive an image editing instruction, determine text description information of the image to be processed, and determine multi-modal features of the image to be processed, where the image editing instruction is text information; a processing unit configured to process the image to be processed according to the image editing instruction, the text description information, and the multi-modal features to obtain a target image; and an insertion unit configured to insert the target image into the video to be processed to obtain a target video.
[0016] In one implementation, the determination unit determines the multi-modal features of the image to be processed in the following manner: obtaining a content segmentation image, a depth segmentation image, and a masking image corresponding to the image to be processed, and obtaining adjacent frame images of the image to be processed from the multiple video frame images; determining the multi-modal features of the image to be processed according to the content segmentation image, the depth segmentation image, the masking image, and the adjacent frame images.
[0017] In one implementation, the determination unit obtains the content segmentation image corresponding to the image to be processed in the following manner: determining multiple target objects included in the image to be processed; performing image segmentation on the image to be processed according to the multiple target objects to make the image regions where the multiple target objects are located independent; and determining the image to be processed after the segmentation as the content segmentation image.
[0018] In one implementation, the determining unit obtains the depth segmentation image corresponding to the image to be processed in the following manner: determining the depth image corresponding to the image to be processed, where the gray value of a pixel in the depth image represents the depth information of the pixel; segmenting the depth image according to the texture features and structural features in the depth image to make multiple regions with continuous gray levels in the depth image independent; and determining the segmented depth image as the depth segmentation map.
[0019] In one implementation, the determining unit obtains the masking image corresponding to the image to be processed in the following manner: for multiple image regions included in the content segmentation image or the depth segmentation image, performing image masking processing one by one, and determining the multiple images obtained by performing the masking processing multiple times as the masking image; where the image masking processing includes: retaining a single image region and masking other multiple image regions.
[0020] In one implementation, obtaining the adjacent frame images of the image to be processed from the multiple video frame images includes: according to the time sequence of the multiple video frame images, obtaining multiple video frame images whose time sequence is before the image to be processed, and obtaining multiple video frame images whose time sequence is after the image to be processed; and determining the multiple video frame images whose time sequence is before the image to be processed and the multiple video frame images whose time sequence is after the image to be processed as the adjacent frame images.
[0021] In one implementation, the determining unit determines the multi-modal features of the image to be processed according to the content segmentation image, the depth segmentation image, the masking image, and the adjacent frame images in the following manner: using the same fusion feature encoder to perform unified feature encoding processing on the segmentation image, the depth segmentation image, the masking image, and the adjacent frame images to obtain the image features corresponding to the segmentation image, the depth segmentation image, the masking image, and the adjacent frame images respectively; and performing feature fusion processing on the image features corresponding to the segmentation image, the depth segmentation image, the masking image, and the adjacent frame images respectively through a feature fusion module, and determining the fusion features obtained by processing the feature fusion module as the multi-modal features.
[0022] In one implementation, the processing unit processes the image to be processed according to the image editing instruction, the text description information, and the multi-modal features to obtain a target image in the following manner: inputting the image editing instruction and the text description information into a large language model to obtain a semantic embedding layer, where the semantic embedding layer contains the image features of the target image; performing noise addition processing on the image to be processed, and adding the multi-modal features to the image to be processed to obtain a multi-modal image; and processing the semantic embedding layer and the multi-modal image through a diffusion model and an encoder to obtain a target image.
[0023] In one implementation, the processing unit processes the semantic embedding layer and the multimodal image through a diffusion model and an encoder in the following manner to obtain a target image: obtaining a semantic vector corresponding to the semantic embedding layer through the diffusion model, and gradually removing noise from the multimodal image according to the semantic vector to obtain a feature conversion image; and adjusting the feature conversion image according to the image size of the image to be processed through the encoder, and determining the feature conversion image after the size adjustment as the target image.
[0024] In one implementation, the insertion unit inserts the target image into the video to be processed in the following manner to obtain a target video: replacing the video frame image corresponding to the processed image in the video to be processed with the target image to obtain the target video.
[0025] According to a third aspect of the embodiments of the present disclosure, there is provided a video processing apparatus, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the video processing method described in the first aspect or any one of the implementations of the first aspect.
[0026] According to a fourth aspect of the embodiments of the present disclosure, there is provided a storage medium storing instructions that, when executed by a processor, enable the processor to execute the video processing method described in the first aspect or any one of the implementations of the first aspect.
[0027] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: obtaining multiple video frame images corresponding to a video to be processed, and determining an image frame to be edited among the multiple video frame images. Receiving an image editing instruction in the form of text information, determining the text description information of the image frame, and the multimodal features of the image frame. Completing the editing of the image frame according to the image editing instruction, the text description information, and the multimodal features to obtain an edited image frame, and inserting the edited image frame into the video to be processed to complete the video editing. Through the present disclosure, the video frame is edited in an interactive text manner, which facilitates the user to quickly get started with video content editing, and based on the content generation method of a single video frame with multiple modalities, improves the picture quality of the edited content and ensures the quality of the edited video.
[0028] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0030] Figure 1 is a flowchart of a video processing method shown according to an exemplary embodiment.
[0031] Figure 2 is a block diagram of a video processing method shown according to an exemplary embodiment of the present disclosure.
[0032] Figure 3 is a flowchart of a method for determining multi-modal features of an image to be processed shown according to an exemplary embodiment.
[0033] Figure 4 is a flowchart of a method for obtaining a content segmentation image shown according to an exemplary embodiment.
[0034] Figure 5 is a flowchart of a method for obtaining a depth segmentation image shown according to an exemplary embodiment.
[0035] Figure 6 is a flowchart of a method for obtaining a masked image shown according to an exemplary embodiment.
[0036] Figure 7 is a flowchart of a method for obtaining adjacent frame images shown according to an exemplary embodiment.
[0037] Figure 8 is a flowchart of a method for determining multi-modal features of an image to be processed shown according to an exemplary embodiment.
[0038] Figure 9 is a flowchart of a method for obtaining a target image shown according to an exemplary embodiment.
[0039] Figure 10 is a flowchart of a method for obtaining a target image shown according to an exemplary embodiment.
[0040] Figure 11 is a flowchart of a method for obtaining a target video shown according to an exemplary embodiment.
[0041] Figure 12 is a block diagram of a video processing method shown according to an exemplary embodiment of the present disclosure.
[0042] Figure 13 is a block diagram of a video processing apparatus shown according to an exemplary embodiment.
[0043] Figure 14 is a block diagram of an apparatus for video processing shown according to an exemplary embodiment. Detailed implementation manners
[0044] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure.
[0045] The video processing method provided by the embodiments of the present disclosure is applied to a scenario of editing video frame images in a video.
[0046] The video processing method proposed by the present disclosure can be applied to a variety of scenarios, such as video editing within the album application of a terminal, video editing can be achieved through the video clip and editing functions of the terminal, video generation functions can be realized in the self-developed application APP of the terminal manufacturer, and video clips can be automatically realized according to text description information, etc.
[0047] The video processing method proposed by the present disclosure mainly relates to the fields of deep neural networks, video editing, diffusion models, and related artificial intelligence technologies. In particular, it relates to a video editing method based on frame-by-frame diffusion and single-frame multi-modal. It can be extended in technical fields such as animated video editing, video editing, and audio-visual production and generation.
[0048] Video editing technology refers to the technology of performing operations such as editing, color correction, special effects processing, and audio processing on original video materials, and finally producing a video work that meets the requirements. Video editing technology can be applied to various types of film and television works such as movies, TV dramas, advertisements, documentaries, and short videos. The development of video editing technology has made film and television production more efficient, accurate, and diversified, and has also brought a richer, more vivid, and interesting visual experience to audiences. The mainstream technologies in current video editing technology include the following aspects:
[0049] 1. Non-linear editing: The current mainstream video editing method, which can arbitrarily perform operations such as video editing, adjustment, and synthesis without damaging the original materials.
[0050] 2. Video special effects: Adding various visual effects to the video, such as transition effects, filters, text, animations, etc., can enhance the expressiveness and attractiveness of the video.
[0051] 3. Color correction: Adjust the color of the video to achieve a better visual effect. The current mainstream color correction software includes DaVinci Resolve, Adobe Premiere Pro, etc.
[0052] 4. 3D synthesis: Synthesize 3D models and real-shot materials to achieve a more realistic effect. The current mainstream 3D synthesis software includes Maya, 3ds Max, etc.
[0053] 5. AI Technology: Video editing effects such as automatic editing and intelligent color correction are achieved through AI video editing software. Currently, the mainstream AI video editing software includes Lumen5, Magisto, etc.
[0054] In related technologies, when users edit videos through general video editing software, they need to use multiple video editing software for video editing. However, the use of video editing software is relatively complex, and different video editing effects require users to use different video editing software. Moreover, all kinds of video editing software require users to have certain usage experience, with a high usage threshold, and beginners cannot quickly master it. As a result, users need to pay a certain learning cost to edit videos through various video editing software to obtain video works that meet their own needs. When users edit videos through AI technology, due to the intelligence level of intelligent AI, it is difficult to understand videos, that is, it cannot accurately understand the content in the videos (such as people, scenes, actions, etc.) in the videos, resulting in the inability to effectively edit the target editing subject according to the editor's prompts; moreover, the automation ability and video editing flexibility of related AI video editing technologies are relatively low: they cannot automatically identify the core parts, important parts, and insignificant parts in video frames, cannot perform selective editing according to the editor's prompts, and have low creativity; furthermore, the video editing and generation effects of related AI video editing technologies are relatively poor, with relatively low understanding ability, and cannot accurately execute the video editing instructions issued by users, resulting in the inability of the generated video effects to fit the original video and poor adaptation ability.
[0055] Furthermore, in related technologies, when users perform individual editing on a certain frame in the video to be edited, they do not consider the consistency of the video frame before and after in the overall video, and the clarity of the edited video frame will decrease, resulting in the inconsistency of the clarity and style of the obtained video frame after editing with the overall video, affecting the quality of the final video work.
[0056] In view of this, the present disclosure proposes a video processing method, which obtains video frames to be edited in a video to be processed, obtains text description information of the video frames, obtains multi-modal features of the video frames, and uses a large language model (LLM) to obtain guidance information for guiding the step-by-step editing of the video frames according to the text description information of the video frames and the image editing instructions of the text type input by the user. Then, according to the video frames to be edited, the above multi-modal features, and the above guidance information, a diffusion model is used to obtain the edited video frames, and then the edited video is obtained. Through the present disclosure, an interactive text-based video editing method is adopted, with a low technical threshold, enabling users to quickly get started with video content editing. By introducing a large language model for video editing, the content in the video, as well as the editing content and generated content, can be accurately understood, so as to accurately perform video editing. The multi-modal features of a single video frame are obtained, and image editing is performed according to the multi-modal features, improving the picture quality of the editing content and ensuring the quality of the generated video.
[0057] Figure 1 is a flowchart of a video processing method shown according to an exemplary embodiment. As Figure 1 shown, the method includes steps S101 to S104.
[0058] In step S101, the video to be processed is frame-divided to obtain multiple video frame images corresponding to the video to be processed, and the image to be processed is determined from the multiple video frame images based on the received selection instruction.
[0059] In step S102, an image editing instruction is received, the text description information of the image to be processed is determined, and the multi-modal features of the image to be processed are determined.
[0060] Among them, the image editing instruction is text information.
[0061] In step S103, the image to be processed is processed according to the image editing instruction, the text description information, and the multi-modal features to obtain a target image.
[0062] In step S104, the target image is inserted into the video to be processed to obtain a target video.
[0063] In the embodiment of the present disclosure, after obtaining the video to be processed for video editing, the video is frame-divided to obtain multiple video frame images corresponding to the video to be processed. It can be understood that the present disclosure performs frame division on the video to be processed based on a preset frame division principle, which has nothing to do with the frame rate of the video to be processed itself. In one example, the present disclosure adopts the principle of dividing 10 frames per second for frame division, so for videos with any frame rate (such as 24 frames per second, 60 frames per second), only 10 video frame images can be obtained per second.
[0064] In the embodiments of the present disclosure, an image is edited using an image editing instruction in the form of interactive text, and the text prompt provides coarse-grained adjustment content and guidance on actions for video editing. The user can edit the image to be processed in the form of text information according to personal needs, achieving image editing effects such as adding special effects to the image, deleting a certain shooting target in the image, and adjusting the overall color of the image.
[0065] In the embodiments of the present disclosure, when editing an image to be processed in the form of text interaction, it is necessary to obtain the text description content corresponding to the content of the image to be processed, and combine the text description content corresponding to the image to be processed with the image editing instruction in text form to gradually guide the image to be processed to transform towards the aspect that meets the user's needs during the subsequent image editing process. In one example, the present disclosure uses a Bootstrapping Language-Image Pre-training (BLIP) text encoder to process the image to be processed and obtain the text description content of the image to be processed (such as a person at the beach, the weather is sunny, and they are playing beach volleyball...).
[0066] In the embodiments of the present disclosure, when obtaining the video to be processed and receiving an instruction to perform video editing processing, multiple video frame images corresponding to the video to be processed are obtained through frame-by-frame processing, and the image frames that need to be edited are determined from the multiple video frame images. An image editing instruction in the form of text information is received, the text description information of the image frame is determined, and the multimodal features of the image frame are obtained. The image frame is edited according to the image editing instruction, the text description information, and the multimodal features to obtain the edited image frame, and the edited image frame is inserted into the video to be processed to complete the video editing. Through the present disclosure, the video frame is edited using interactive text, which facilitates the user to quickly get started with video content editing, and based on the content generation method of single-frame multimodality of the video, improves the image quality of the edited content and ensures the quality of the edited video.
[0067] It can be understood that the image processing method proposed in the present disclosure is a method for processing a single-frame image in a video to be processed. It can be understood that in actual image editing processing, the user will issue a video editing instruction for the entire video, that is, perform image editing processing on different groups of consecutive multiple video frame images. In this case, the video processing method proposed in the present disclosure can be executed for each selected video frame image one by one.
[0068] In an exemplary embodiment of the present disclosure, such as Figure 2As shown in the block diagram of the medium video processing method, the present disclosure performs video editing processing based on text interaction in the following manner: Input the video to be processed and text prompts (i.e., image editing instructions); perform data processing on the text prompt data and the video data corresponding to the video to be processed to obtain multiple groups of video frames in the video to be processed, and use each frame in the multiple groups of video frames as the image to be processed, and perform image editing processing one by one based on the text prompt data. After obtaining the multiple groups of video frames that have completed the editing process, finally perform video synthesis, insert the multiple groups of video frames that have completed the editing process into the original video, replace the corresponding multiple groups of video frames, and obtain the target video.
[0069] In the embodiments of the present disclosure, the multi-modal features of the image to be processed are a set of multiple features, including multiple editable contents in the image to be processed and the associated features of the image to be processed in the video to be processed. The following embodiments of the present disclosure illustrate the method for determining the multi-modal features.
[0070] Figure 3 It is a flowchart of a method for determining the multi-modal features of an image to be processed shown according to an exemplary embodiment. As Figure 3 shown, the method includes steps S201 to step S202.
[0071] In step S201, obtain the content segmentation image, depth segmentation image, and masking image corresponding to the image to be processed, and obtain the adjacent frame images of the image to be processed in the multi-frame video frame images.
[0072] In step S202, determine the multi-modal features of the image to be processed according to the content segmentation image, depth segmentation image, masking image, and adjacent frame images.
[0073] In the embodiments of the present disclosure, the content segmentation image, depth segmentation image, and masking image corresponding to the image to be processed reflect multiple editable contents in the image to be processed, and the adjacent frame images of the image to be processed can reflect the change trend of each pixel in the image to be processed.
[0074] In the embodiments of the present disclosure, the multi-modal features are obtained through the above-mentioned images corresponding to each image to be processed, and the image structure and image style guidance are provided according to the multi-modal features during the editing process of the image to be processed, so as to achieve fine-grained video spatial content control and diversified and stylized image content editing.
[0075] It can be understood that there will be multiple editable independent target subjects in the image to be processed, and the content segmentation image of the image to be processed can reflect the above-mentioned multiple editable independent target subjects. The following embodiments of the present disclosure illustrate the method for obtaining the content segmentation image.
[0076] Figure 4It is a flowchart of a method for obtaining a content segmentation image shown according to an exemplary embodiment. As Figure 4 shown, the method includes steps S301 to S303.
[0077] In step S301, a plurality of target objects included in the image to be processed are determined.
[0078] In step S302, the image to be processed is segmented according to the plurality of target objects, so that the image regions where the plurality of target objects are located are independent.
[0079] In step S303, the image to be processed after segmentation is determined as the content segmentation image.
[0080] In the embodiments of the present disclosure, after performing a frame splitting operation on the video to be processed and obtaining the image to be processed therein, the content and structure of the video at the corresponding time point can be reflected by the image to be processed. And according to the plurality of target objects (such as people, animals, and single objects, etc.) included in the image to be processed, image segmentation is performed on the current frame image, so that the image regions where each image object is located are independent, which is convenient for understanding the content in the video and the edited content and generated content (including characters, scenes, actions, etc.) during the subsequent image editing process, accurately identifying the target object specified by the image editing instruction, and performing separate editing processing on the target object to ensure the accuracy of video editing.
[0081] It can be understood that different contents in the image to be processed have different image depths. Based on the different image depths, different contents in the image can also be segmented. The following embodiments of the present disclosure illustrate the method for obtaining a depth segmentation map.
[0082] Figure 5 It is a flowchart of a method for obtaining a depth segmentation image shown according to an exemplary embodiment. As Figure 5 shown, the method includes steps S401 to S403.
[0083] In step S401, the depth image corresponding to the image to be processed is determined.
[0084] Among them, the gray value of the pixel in the depth image represents the depth information of the pixel;
[0085] In step S402, the depth image is segmented according to the texture features and structural features in the depth image, so that a plurality of regions with continuous gray levels in the depth image are independent.
[0086] In step S403, the depth image after segmentation is determined as the depth segmentation map.
[0087] In the embodiments of the present disclosure, the depth image corresponding to the image to be processed is a pure grayscale image. Different grayscale values in the depth image represent different depths. It can be understood that the region corresponding to each target object in the image to be processed is a region with continuous depth in the grayscale image, that is, a region with gradual grayscale change; while there are large depth changes at the edges of each target object in the image to be processed, which will be reflected as sudden and discontinuous grayscale changes in the corresponding depth image, that is, the texture features in the depth image. The present disclosure segments the depth image based on the texture features and structural features in the depth image, so that each independent structure corresponding region (region with continuous depth change) in the depth image is independent, and a depth segmentation map corresponding to the image to be processed is obtained. Further improve the accuracy of identifying the target object specified by the image editing instruction and ensure the accuracy of video editing.
[0088] In an exemplary embodiment of the present disclosure, a pixel difference network (PiDiNet) is used to obtain the depth map corresponding to the image to be processed, and the structural features and texture features in the depth image are identified through edge detection. Then, the depth image is segmented based on the structural features and texture features to obtain the depth segmentation map corresponding to the image to be processed.
[0089] In the embodiments of the present disclosure, the multiple image regions included in the content segmentation image and the multiple image regions included in the depth segmentation image respectively correspond to target objects that can be edited separately. The present disclosure can specifically obtain image features for each target object, which is convenient for the selection and editing of the target object to be edited in the subsequent image processing process. The following embodiments of the present disclosure illustrate the method for obtaining the masking image.
[0090] Figure 6 is a flowchart of a method for obtaining a masking image shown according to an exemplary embodiment. As Figure 6 shown, the method includes steps S501 to S502.
[0091] In step S501, for the multiple image regions included in the content segmentation image or the depth segmentation image, image masking processing is performed one by one.
[0092] In step S502, the multiple images obtained by performing the masking processing multiple times are determined as the masking images.
[0093] Among them, the image masking processing includes: retaining a single image region and masking other multiple image regions.
[0094] In the embodiments of the present disclosure, the obtained content segmentation image or depth segmentation image is further processed. For each of the multiple image regions included in the image, image masking processing is performed one by one, that is, a single image region is retained and other image regions are masked, and the masked image corresponding to each independent image region is obtained. The masked features are extended in the channel dimension, and the masked images corresponding to multiple independent image regions are determined as the masked image corresponding to the image to be processed. It can be understood that there is a certain correspondence between the multiple independent image regions in the content segmentation image and the multiple independent image regions in the depth segmentation image. Therefore, when performing image masking processing to obtain the masked image, only one of the content segmentation image and the depth segmentation image is required.
[0095] In the embodiments of the present disclosure, by segmentally masking the segmentation content (i.e., independent image regions) in the content segmentation image or the depth segmentation image, it is convenient for the editing and repair of the corresponding target subject in the image to be processed. It can be understood that when performing image masking processing, the user can manually add a mask based on personal needs and force the model to predict the covered area according to the observable information.
[0096] In the embodiments of the present disclosure, the adjacent frame images of the image to be processed can reflect the change trend of each pixel in the image to be processed, so that the subsequent editing and processing content can better fit the content style of the original video (i.e., the video to be processed). The following embodiments of the present disclosure illustrate the method for obtaining adjacent frame images.
[0097] Figure 7 is a flowchart of a method for obtaining adjacent frame images shown according to an exemplary embodiment. As Figure 7 shown, the method includes steps S601 to S602.
[0098] In step S601, according to the timing of multiple video frame images, multiple video frame images with a timing before the image to be processed are obtained, and multiple video frame images with a timing after the image to be processed are obtained.
[0099] In step S602, the multiple video frame images with a timing before the image to be processed and the multiple video frame images with a timing after the image to be processed are determined as adjacent frame images.
[0100] In the embodiments of the present disclosure, according to the preset acquisition quantity and the timing of multiple video frame images, multiple adjacent and continuous video frame images are obtained before the image to be processed, and multiple adjacent and continuous video frame images are obtained after the image to be processed. The number of multiple video frame images with a timing before the image to be processed is the same as the number of multiple video frame images with a timing after the image to be processed.
[0101] In the embodiments of the present disclosure, by obtaining adjacent frame images of the image to be processed, fine control along the time dimension can be achieved for the image to be processed in the subsequent image processing process. The image to be processed itself and its corresponding adjacent frame images are used as motion vectors distributed along the time axis, and the pixel-level motion between multiple adjacent frames can be encoded and displayed. By providing control details of pixel changes through a time series, precise editing can be achieved, ensuring the continuity of the target video.
[0102] In an exemplary embodiment of the present disclosure, after the frame splitting process of the image to be processed is completed, the 56th frame image of the video to be processed is selected as the image to be processed. The preset number of adjacent frame images to be collected is 6 frames. Then, the video frame images of the 53rd, 54th, and 55th frames before the image to be processed in terms of time sequence are obtained, and the video frame images of the 57th, 58th, and 59th frames after the image to be processed in terms of time sequence are obtained. The video frame images of the 53rd, 54th, 55th, 57th, 58th, and 59th frames are used as the adjacent frame images of the image to be processed.
[0103] In the embodiments of the present disclosure, the corresponding image features of the above-mentioned segmented image, depth segmented image, masked image, and adjacent frame images need to be used as conditions for guiding the image processing of the image to be processed, and the corresponding image features of the above-mentioned segmented image, depth segmented image, masked image, and adjacent frame images need to be subjected to unified fusion processing. The following embodiments of the present disclosure further illustrate the method for determining multi-modal features.
[0104] Figure 8 is a flowchart of a method for determining multi-modal features of an image to be processed shown according to an exemplary embodiment. As Figure 8 shown, the method includes steps S701 to S702.
[0105] In step S701, the same fusion feature encoder is used to perform unified feature encoding processing on the segmented image, depth segmented image, masked image, and adjacent frame images, and the image features corresponding to the segmented image, depth segmented image, masked image, and adjacent frame images are obtained respectively.
[0106] In step S702, the feature fusion module is used to perform feature fusion processing on the image features corresponding to the segmented image, depth segmented image, masked image, and adjacent frame images respectively, and the fusion features obtained by the processing of the feature fusion module are determined as multi-modal features.
[0107] In the embodiments of the present disclosure, when performing feature fusion processing on the corresponding image features of the above-mentioned segmented image, depth segmented image, masked image, and adjacent frame image, in order to ensure the effectiveness of different feature sharing, each image is uniformly feature encoded through the same fusion feature encoder (Late Fusion-Encoder, LF-Encoder), and each image is separately subjected to uniform convolution processing to obtain multiple post-convolution images with the same convolution structure (composed of 3 two-dimensional convolutions and 1 fully connected layer). After the features are uniformly encoded, the multiple post-convolution images are sent into a unified channel feature fusion (Concat-Fusion, C-Fusion) encoder, and the spatio-temporal relationships in different input features are captured by concatenation (Concat) in the channel dimension and the channel attention mechanism to generate an initialized feature map (i.e., multi-modal features), providing rich feature information for image processing operations, thereby enhancing the quality of the final output target image.
[0108] In the embodiments of the present disclosure, for the execution of the feature fusion processing flow, considering that multiple features (the corresponding image features of the segmented image, depth segmented image, masked image, and adjacent frame image) and features containing sequence conditions have rich and complex dependencies, which will affect the controllable guidance in the subsequent image processing flow. It is merged through a channel feature fusion module, namely a channel feature fusion (Concat-Fusion, C-Fusion) encoder, and the relationships of multiple features are simplified and enhanced to strengthen the temporal perception of the input conditions in the subsequent image processing flow.
[0109] In an exemplary embodiment of the present disclosure, the feature fusion processing flow is executed in the following manner: all features are composed of two convolutional kernels (Conv2D) and an average pooling layer (Average Pooling) to form a lightweight architecture to extract local spatial information and obtain a time series; the obtained time series is input into a model architecture (Transformer) for sequence conversion to perform temporal modeling. Through the above feature fusion processing flow, it is helpful to perform temporal and spatial modeling on multi-dimensional features, enhance the consistency between video frames of the target video, and simplify the feature dimension, which is beneficial to establishing efficient computing.
[0110] In the embodiments of the present disclosure, when editing the image to be processed in the form of text interaction, it is necessary to combine the text description content corresponding to the image to be processed and the image editing instruction in text form, and combine the image to be processed itself and the multi-modal features of the image to be processed to gradually guide the content transformation of the image to be processed in the subsequent image editing process to obtain the target image. The following embodiments of the present disclosure illustrate the method for obtaining the target image.
[0111] Figure 9 is a flowchart of a method for obtaining a target image shown according to an exemplary embodiment. AsFigure 9 As shown, the method includes steps S801 to S803.
[0112] In step S801, an image editing instruction and text description information are input into a large language model to obtain a semantic embedding layer.
[0113] Among them, the semantic embedding layer contains the image features of the target image.
[0114] In step S802, noise is added to the image to be processed, and multi-modal features are added to the image to be processed to obtain a multi-modal image.
[0115] In step S803, the semantic embedding layer and the multi-modal image are processed by a diffusion model and an encoder to obtain the target image.
[0116] In the embodiments of the present disclosure, after obtaining the image editing instruction in the form of interactive text and the text description information of the image to be processed, the large language model is used to process the editing instruction and the text description information to obtain a semantic embedding layer (Embedding) for providing coarse-grained adjustment content and action guidance for the image to be processed. In one example, a language model based on a large language model (such as ChatGLM, Chat General Language Model) is used to obtain the semantic embedding layer, or a text encoder model using a pre-training method based on contrastive text-image pairs (Contrastive Language-Image Pre-training, Clip) is used to obtain the semantic embedding layer, and the output model (Encoder) of the last set of model architectures (Transformer) modules for sequence conversion in the model is used as the semantic embedding layer of the text description. The present disclosure uses the image editing instruction and the text description information together as the prompt content for image editing to generate a semantic embedding layer, providing accurate guidance in the subsequent image processing process.
[0117] In the embodiments of the present disclosure, after obtaining the multi-modal features, noise is added to the image to be processed, and the multi-modal features are added to the image to be processed, providing initial noise for the image processing flow in the subsequent diffusion model and providing rich feature information, enabling the diffusion model to accurately understand the inherent content, editing content, and content to be generated (including people, scenes, actions, etc.) in the image to be processed, so as to accurately perform image processing based on the prompt of the semantic embedding layer.
[0118] It can be understood that the image output by the diffusion model already contains the content specified by the image editing instruction. However, there are differences in the length and width dimensions between the image output by the diffusion model and the video frames in the original video (the video to be processed). Therefore, it is necessary to restore the image size of the image output by the diffusion model through an encoder. The following embodiments of the present disclosure further illustrate the method for obtaining the target image.
[0119] Figure 10 is a flowchart of a method for obtaining a target image shown according to an exemplary embodiment. As Figure 10 shown, the method includes steps S901 to step S902.
[0120] In step S901, a semantic vector corresponding to the semantic embedding layer is obtained through the diffusion model, and noise is gradually removed from the multimodal image according to the semantic vector to obtain a feature conversion image.
[0121] In step S902, through the encoder, the feature conversion image is adjusted according to the image size of the image to be processed, and the feature conversion image after the size adjustment is determined as the target image.
[0122] In the embodiments of the present disclosure, after obtaining the semantic embedding layer for guiding image processing of the image to be processed and the multimodal image containing multimodal features, the multimodal image and the semantic embedding layer are input into the diffusion model (such as StableDiffusion, SD), and the natural language in the semantic embedding layer is converted into a mathematized semantic vector that can be understood by the machine. Then, in combination with the semantic vector, noise is gradually removed from the pure noise to generate a latent variable containing multimodal image information. Finally, the obtained latent variable is converted into a feature conversion image containing modal image information, and the feature conversion image is processed by a variational auto-encoder (VAE-Decoder) to restore the image size on the premise of ensuring image clarity and obtain the target image. The size and clarity of the target image are consistent with the size and clarity of the video frames in the original video (the video to be processed).
[0123] The following embodiments of the present disclosure illustrate the method for obtaining the target video.
[0124] Figure 11 is a flowchart of a method for obtaining a target video shown according to an exemplary embodiment. As Figure 11 shown, the method includes steps S1001 to step S1002.
[0125] In step S1001, the target image is obtained.
[0126] In step S1002, the target image replaces the video frame image corresponding to the processed image in the video to be processed, obtaining a target video.
[0127] In the embodiments of the present disclosure, after obtaining a target image with the same size and clarity as the video frames in the original video (video to be processed), the corresponding video frame images in multiple video frame images of the video to be processed are replaced with the target image, obtaining an edited video and completing video editing. It can be understood that when a user performs image editing on a video, the issued image editing instruction requires overall editing of the video. In this case, when obtaining the target image from multiple video frame images after frame division, based on the image selection instruction corresponding to the image editing instruction, multiple sets of consecutive video frame images can be obtained from the multiple video frame images as images to be processed, and multiple modalities are respectively performed on all images to be processed. Feature extraction, description information extraction, diffusion model processing, encoder processing, and other processing processes to obtain multiple target images corresponding to all images to be processed, and replace the corresponding multiple video frame images in all video frame images of the video to be processed with the multiple target images, obtaining an edited video and completing video editing.
[0128] The video processing method proposed by the present disclosure can be applied to various scenarios requiring video production and post-production, providing users with a richer, more vivid, and interesting visual experience. In an exemplary embodiment, the specific usage scenarios of the video processing method proposed by the present disclosure may include the following aspects:
[0129] 1. Short video production on social media and personal entertainment applications (such as the system video application on the terminal): The video processing method proposed by the present disclosure can realize the editing and generation of videos on the basis of the original video, and can be used to produce short videos on social media and submit the edited and generated video works on various social media applications. With the rise of short video platforms, more and more users start to use mobile terminals such as mobile phones and tablets to produce short videos. The video processing method proposed by the present disclosure can help users perform operations such as video editing, special effect production, and subtitle addition on terminals such as mobile phones and tablets, producing more vivid and interesting short videos. In terms of personal entertainment, the video processing method proposed by the present disclosure can help users produce personal entertainment videos closely related to the user's life on terminals such as mobile phones and tablets, such as Vlogs, travel videos, food videos, etc.
[0130] 2. In some jobs that require video editing, such as: Advertising production: The video processing method proposed in this disclosure can generate videos for advertising production according to the user's text description, including videos for aspects such as TV commercials, online ads, and social media ads. At the same time, for unsatisfactory details, segmented editing can be performed. Education and training: Video editing and generation technology can be used to produce education and training videos, such as online courses, training videos, etc. Film production: Video editing and generation technology can be used for post-production in film production, such as special effects production, scene synthesis, etc. Live streaming: Live streaming applications on mobile phones and tablets are becoming increasingly popular. Video editing and generation technology can help users perform real-time editing, special effects production, scene synthesis, etc. during live streaming, enhancing the viewing and interactivity of live streaming.
[0131] 3. Virtual Reality (VR) / Augmented Reality (AR) applications: The video processing method proposed in this disclosure can be used for scene synthesis and special effects production in VR / AR applications, enhancing the immersion of virtual reality and augmented reality.
[0132] In an exemplary embodiment of the present disclosure, as Figure 12 shown in the schematic diagram of the video processing method in, the following method is used to process the video to be processed to obtain the target video after processing:
[0133] Input the video to be processed, perform frame division processing on the video to be processed, and decompose the video to be processed frame by frame into multiple video frame images. Select the image to be processed from the multiple video frame images, add noise to the image to be processed, obtain the multi-modal conditions corresponding to the image to be processed, and add the multi-modal conditions to the image to be processed after adding noise to obtain a multi-modal image. Extract the basic content of the image to be processed through a visual language recognition model (such as Bootstrapping Language-Image Pre-training, BLIP), process the text prompt information sent by the user and the basic content of the image to be processed through a large language model, and obtain the semantic embedding layer (Embedding) for editing the image to be processed. Input the obtained multi-modal image and semantic embedding layer into a diffusion model (such as Stable Diffusion, SD), and convert the natural language in the semantic embedding layer into a mathematical semantic vector that can be understood by the machine. Then, in combination with the semantic vector, gradually remove the noise from the multi-modal image starting from pure noise to generate a latent variable containing the multi-modal image information. Finally, process the above feature-transformed image through a variational auto-encoder (VAE-Decoder), restore the image size to obtain the target image, and replace the corresponding video frame image in the multiple video frame images of the video to be processed with the target image to obtain the edited video and complete the video editing.
[0134] Among them, the multi-modal features corresponding to the image to be processed include the features of the image to be processed itself, the features of the depth segmentation map corresponding to the image to be processed, the features of the depth occlusion map corresponding to the image to be processed, and the temporal frame condition features corresponding to the image to be processed (e.g., the first three frames and the last three frames of the image to be processed). As Figure 12 shown in the schematic diagram of the image processing method in, the multi-modal features corresponding to the image to be processed are obtained in the following manner: the image to be processed is obtained by performing frame division on multiple frames of images included in the video to be processed; the depth map corresponding to the image to be processed is obtained through a Pixel Difference Networks (PiDiNet), and the depth segmentation map corresponding to the image to be processed is obtained by segmenting the depth map based on the structure and texture in the image to be processed; the obtained depth map is segmented and segmented for occlusion, and the masked features are extended in the channel dimension to obtain the depth occlusion map corresponding to the image to be processed; the adjacent frames (the first three frames and the last three frames) of the image to be processed are obtained as the temporal frame conditions corresponding to the image to be processed. The above-mentioned image to be processed, depth segmentation map, depth occlusion map, and adjacent frame images are uniformly feature-extracted through the same Late Fusion-Encoder (LF-Encoder). Multiple features obtained by the same feature extraction are input into a Concat-Fusion (C-Fusion) module to obtain the multi-modal features of the image to be processed.
[0135] In the embodiments of the present disclosure, an interactive text-based video editing method is adopted. Based on a large language model, text-based image editing instructions and text description information of the image to be processed are processed, and the image processing process is guided. It can accurately understand the inherent content, edited content, and generated content of the video, accurately perform video editing, and lower the technical threshold of video editing, enabling users to quickly get started with video content editing. After the video is frame-divided in the present disclosure, the image to be processed is frame-by-frame diffusion-edited through a diffusion model, combined with the content understanding of the large language model, and combined with the relevant frames before and after each frame of the image to be processed currently for frame-by-frame diffusion image processing, so that the quality and content style of the target image obtained after editing fit the original video. Based on the video single-frame multi-modal content generation method in the present disclosure, the multi-modal features corresponding to the image to be processed are obtained, and the multi-modal features participate in the image processing process, improving the accuracy of identifying the target subject specified by the image editing instruction, ensuring the accuracy of video editing, improving the image quality of the edited content, and making the generated video have lower losses in terms of quality, fusion degree, saturation, color difference, etc.
[0136] Based on the same concept, the embodiments of the present disclosure also provide a video processing device 100.
[0137] It can be understood that in order to implement the above functions, the video processing device 100 provided in the embodiments of the present disclosure includes the corresponding hardware structures and / or software modules for executing various functions. Combining the units and algorithm steps of the various examples disclosed in the embodiments of the present disclosure, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the technical solutions of the embodiments of the present disclosure.
[0138] Figure 13 is a block diagram of a video processing device 100 shown according to an exemplary embodiment. Referring to Figure 13 , the device includes a selection unit 101, a determination unit 102, a processing unit 103, and an insertion unit 104.
[0139] The selection unit 101 is configured to perform frame division processing on the video to be processed, obtain multiple video frame images corresponding to the video to be processed, and determine the image to be processed from the multiple video frame images based on the received selection instruction.
[0140] The determination unit 102 is configured to receive an image editing instruction, determine the text description information of the image to be processed, and determine the multimodal features of the image to be processed, where the image editing instruction is text information.
[0141] The processing unit 103 is configured to process the image to be processed according to the image editing instruction, the text description information, and the multimodal features to obtain a target image.
[0142] The insertion unit 104 is configured to insert the target image into the video to be processed to obtain a target video.
[0143] In one implementation, the determination unit 102 determines the multimodal features of the image to be processed in the following manner: obtaining the content segmentation image, depth segmentation image, and masking image corresponding to the image to be processed, and obtaining the adjacent frame images of the image to be processed from the multiple video frame images. Determine the multimodal features of the image to be processed according to the content segmentation image, depth segmentation image, masking image, and adjacent frame images.
[0144] In one implementation, the determination unit 102 obtains the content segmentation image corresponding to the image to be processed in the following manner: determining multiple target objects included in the image to be processed. Perform image segmentation on the image to be processed according to the multiple target objects to make the image regions where the multiple target objects are located independent. Determine the segmented image to be processed as the content segmentation image.
[0145] In one implementation, the determining unit 102 obtains the depth segmentation image corresponding to the image to be processed in the following manner: Determine the depth image corresponding to the image to be processed, where the gray value of a pixel in the depth image represents the depth information of the pixel. Segment the depth image according to the texture features and structural features in the depth image, so that multiple regions with continuous gray levels in the depth image are independent. The depth image after segmentation is determined as the depth segmentation map.
[0146] In one implementation, the determining unit 102 obtains the mask image corresponding to the image to be processed in the following manner: For multiple image regions included in the content segmentation image or the depth segmentation image, perform image masking processing one by one, and determine the multiple images obtained by performing the masking processing multiple times as the mask image. Among them, the image masking processing includes: retaining a single image region and masking other multiple image regions.
[0147] In one implementation, obtaining the adjacent frame images of the image to be processed from multiple video frame images includes: According to the time sequence of the multiple video frame images, obtain multiple video frame images whose time sequence is before the image to be processed, and obtain multiple video frame images whose time sequence is after the image to be processed. The multiple video frame images whose time sequence is before the image to be processed and the multiple video frame images whose time sequence is after the image to be processed are determined as the adjacent frame images.
[0148] In one implementation, the determining unit 102 determines the multi-modal features of the image to be processed according to the content segmentation image, the depth segmentation image, the mask image, and the adjacent frame images in the following manner: Use the same fusion feature encoder to perform unified feature encoding processing on the segmentation image, the depth segmentation image, the mask image, and the adjacent frame images, and obtain the image features corresponding to the segmentation image, the depth segmentation image, the mask image, and the adjacent frame images respectively. Perform feature fusion processing on the image features corresponding to the segmentation image, the depth segmentation image, the mask image, and the adjacent frame images through the feature fusion module, and determine the fusion features obtained by the processing of the feature fusion module as the multi-modal features.
[0149] In one implementation, the processing unit 103 processes the image to be processed according to the image editing instruction, the text description information, and the multi-modal features to obtain the target image in the following manner: Input the image editing instruction and the text description information into the large language model to obtain the semantic embedding layer, and the semantic embedding layer contains the image features of the target image. Perform noise addition processing on the image to be processed, and add the multi-modal features to the image to be processed to obtain the multi-modal image. Process the semantic embedding layer and the multi-modal image through the diffusion model and the encoder to obtain the target image.
[0150] In one implementation, the processing unit 103 processes the semantic embedding layer and the multimodal image through a diffusion model and an encoder in the following manner to obtain a target image: obtaining a semantic vector corresponding to the semantic embedding layer through the diffusion model, and gradually removing noise from the multimodal image according to the semantic vector to obtain a feature conversion image. Through the encoder, adjusting the feature conversion image according to the image size of the image to be processed, and determining the feature conversion image after the size adjustment as the target image.
[0151] In one implementation, the insertion unit 104 inserts the target image into the video to be processed in the following manner to obtain a target video: replacing the video frame image corresponding to the processed image in the video to be processed with the target image to obtain the target video.
[0152] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0153] Figure 14 It is a block diagram of a device 200 for video processing shown according to an exemplary embodiment. The device 200 may be provided as a terminal. For example, the device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0154] Referring to Figure 14 , the device 200 may include one or more of the following components: a processing component 202, a memory 204, a power component 206, a multimedia component 208, an audio component 210, an input / output (I / O) interface 212, a sensor component 214, and a communication component 216.
[0155] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 202 may include one or more modules to facilitate the interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate the interaction between the multimedia component 208 and the processing component 202.
[0156] The memory 204 is configured to store various types of data to support the operation of the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, and the like. The memory 204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0157] The power components 206 provide power to the various components of the device 200. The power components 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 200.
[0158] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0159] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.
[0160] The I / O interface 212 provides an interface between the processing component 202 and peripheral interface modules, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0161] The sensor assembly 214 includes one or more sensors for providing an assessment of various aspects of the status of the device 200. For example, the sensor assembly 214 can detect the on / off state of the device 200, the relative positioning of components, such as the display and keypad of the device 200. The sensor assembly 214 can also detect a change in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and the temperature change of the device 200. The sensor assembly 214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0162] The communication component 216 is configured to facilitate communication between the device 200 and other devices in a wired or wireless manner. The device 200 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0163] In an exemplary embodiment, the device 200 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0164] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 204 including instructions, is also provided. The above instructions can be executed by the processor 220 of the device 200 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0165] It can be understood that in this disclosure, "a plurality of" means two or more, and other quantifiers are similar. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The singular forms of "a", "the", and "said" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0166] It can be further understood that the terms "first", "second", etc. are used to describe various information, but this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other and do not represent a specific order or degree of importance. In fact, the expressions such as "first" and "second" can be used interchangeably. For example, without departing from the scope of this disclosure, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.
[0167] It can be further understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "front", "rear", "upper", "lower", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this embodiment and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation.
[0168] It can be further understood that unless otherwise specified, "connection" includes direct connection without other components between the two, and also includes indirect connection with other elements between the two.
[0169] It can be further understood that although the operations are described in a specific order in the drawings in the embodiments of this disclosure, it should not be understood as requiring these operations to be performed in the specific order shown or in a serial order, or requiring all the operations shown to obtain the desired result. In a specific environment, multitasking and parallel processing may be beneficial.
[0170] Those skilled in the art will readily think of other embodiments of this disclosure after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this solution, which follow the general principles of this disclosure and include common general knowledge or conventional technical means in this technical field not disclosed in this disclosure.
[0171] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A video processing method, characterized in that, comprising: Performing frame division on the video to be processed, obtaining multiple video frame images corresponding to the video to be processed, and determining the image to be processed from the multiple video frame images based on the received selection instruction; Receiving an image editing instruction, determining the text description information of the image to be processed, and determining the multimodal features of the image to be processed, where the image editing instruction is text information; Processing the image to be processed according to the image editing instruction, the text description information, and the multimodal features to obtain a target image; Inserting the target image into the video to be processed to obtain a target video.
2. The method according to claim 1, characterized in that, Determining the multimodal features of the image to be processed includes: Obtaining the content segmentation image, depth segmentation image, and masking image corresponding to the image to be processed, and obtaining the adjacent frame images of the image to be processed from the multiple video frame images; Determining the multimodal features of the image to be processed according to the content segmentation image, the depth segmentation image, the masking image, and the adjacent frame images.
3. The method according to claim 2, characterized in that, Obtaining the content segmentation image corresponding to the image to be processed includes: Determining multiple target objects included in the image to be processed; Performing image segmentation on the image to be processed according to the multiple target objects to make the image regions where the multiple target objects are located independent; Determining the processed image to be processed as the content segmentation image.
4. The method according to claim 2, characterized in that, Obtaining the depth segmentation image corresponding to the image to be processed includes: Determining the depth image corresponding to the image to be processed, where the gray value of the pixel in the depth image represents the depth information of the pixel; Performing segmentation on the depth image according to the texture features and structural features in the depth image to make multiple regions with continuous gray levels in the depth image independent; Determining the processed depth image as the depth segmentation map.
5. The method according to claim 2, characterized in that, Obtaining the masking image corresponding to the image to be processed includes: Performing image masking processing on each of the multiple image regions included in the content segmentation image or depth segmentation image, and determining the multiple images obtained by performing the masking processing multiple times as the masking image; wherein, the image masking processing includes: retaining a single image region and masking other multiple image regions.
6. The method according to claim 2, characterized in that, Obtaining the adjacent frame images of the image to be processed from the multiple video frame images includes: According to the time sequence of the multiple video frame images, obtaining multiple video frame images with a time sequence before the image to be processed, and obtaining multiple video frame images with a time sequence after the image to be processed; Determining the multiple video frame images with a time sequence before the image to be processed and the multiple video frame images with a time sequence after the image to be processed as the adjacent frame images.
7. The method according to claim 2, characterized in that, Segmenting the image, the depth segmentation image, the masking image, and the adjacent frame image according to the content, and determining the multi-modal features of the image to be processed, including: Using the same fusion feature encoder to perform unified feature encoding processing on the segmentation image, the depth segmentation image, the masking image, and the adjacent frame image, and obtaining the image features corresponding to the segmentation image, the depth segmentation image, the masking image, and the adjacent frame image respectively; Performing feature fusion processing on the image features corresponding to the segmentation image, the depth segmentation image, the masking image, and the adjacent frame image through a feature fusion module, and determining the fusion features obtained by the processing of the feature fusion module as the multi-modal features.
8. The method according to claim 1, wherein, Processing the image to be processed according to the image editing instruction, the text description information, and the multi-modal features to obtain a target image, including: Inputting the image editing instruction and the text description information into a large language model to obtain a semantic embedding layer, and the semantic embedding layer contains the image features of the target image; Performing noise addition processing on the image to be processed, and adding the multi-modal features to the image to be processed to obtain a multi-modal image; Processing the semantic embedding layer and the multi-modal image through a diffusion model and an encoder to obtain a target image.
9. The method according to claim 8, wherein, Processing the semantic embedding layer and the multi-modal image through a diffusion model and an encoder to obtain a target image, including: Obtaining a semantic vector corresponding to the semantic embedding layer through the diffusion model, and gradually removing noise from the multi-modal image according to the semantic vector to obtain a feature conversion image; Adjusting the feature conversion image according to the image size of the image to be processed through the encoder, and determining the feature conversion image after the size adjustment as the target image.
10. The method according to claim 8, wherein, Inserting the target image into the video to be processed to obtain a target video, including: Replacing the video frame image corresponding to the processed image in the video to be processed with the target image to obtain the target video.
11. A video processing device, wherein, including: A selection unit, configured to perform frame segmentation on a video to be processed, obtain multiple video frame images corresponding to the video to be processed, and determine an image to be processed from the multiple video frame images based on a received selection instruction; A determination unit, configured to receive an image editing instruction, determine text description information of the image to be processed, and determine multi-modal features of the image to be processed, where the image editing instruction is text information; A processing unit, configured to process the image to be processed according to the image editing instruction, the text description information, and the multi-modal features to obtain a target image; An insertion unit, configured to insert the target image into the video to be processed to obtain a target video.
12. The device according to claim 11, wherein, The determination unit determines the multi-modal features of the image to be processed in the following manner: Obtain the content segmentation image, depth segmentation image, and masking image corresponding to the image to be processed, and obtain the adjacent frame images of the image to be processed in the multiple video frame images; Determine the multi-modal features of the image to be processed according to the content segmentation image, the depth segmentation image, the masking image, and the adjacent frame images.
13. The apparatus according to claim 12, wherein, the determining unit obtains the content segmentation image corresponding to the image to be processed in the following manner: Determine multiple target objects included in the image to be processed; Perform image segmentation on the image to be processed according to the multiple target objects, so that the image regions where the multiple target objects are located are independent; Determine the image to be processed after segmentation as the content segmentation image.
14. The apparatus according to claim 12, wherein, the determining unit obtains the depth segmentation image corresponding to the image to be processed in the following manner: Determine the depth image corresponding to the image to be processed, and the gray value of the pixel in the depth image represents the depth information of the pixel; Segment the depth image according to the texture features and structural features in the depth image, so that multiple regions with continuous gray levels in the depth image are independent; Determine the depth image after segmentation as the depth segmentation map.
15. The apparatus according to claim 12, wherein, the determining unit obtains the masking image corresponding to the image to be processed in the following manner: Perform image masking processing on each of the multiple image regions included in the content segmentation image or the depth segmentation image, and determine the multiple images obtained by performing masking processing multiple times as the masking image; wherein, the image masking processing includes: retaining a single image region and masking other multiple image regions.
16. The apparatus according to claim 12, wherein, the obtaining of the adjacent frame images of the image to be processed in the multiple video frame images includes: According to the time sequence of the multiple video frame images, obtain multiple video frame images whose time sequence is before the image to be processed, and obtain multiple video frame images whose time sequence is after the image to be processed; Determine the multiple video frame images whose time sequence is before the image to be processed and the multiple video frame images whose time sequence is after the image to be processed as the adjacent frame images.
17. The apparatus according to claim 12, wherein, the determining unit determines the multi-modal features of the image to be processed according to the content segmentation image, the depth segmentation image, the masking image, and the adjacent frame images in the following manner: Use the same fusion feature encoder to perform unified feature encoding processing on the segmentation image, the depth segmentation image, the masking image, and the adjacent frame images, and obtain the image features corresponding to the segmentation image, the depth segmentation image, the masking image, and the adjacent frame images respectively; The feature fusion module performs feature fusion processing on the image features corresponding to the segmentation image, the depth segmentation image, the masking image, and the adjacent frame image respectively, and determines the fused features obtained by the processing of the feature fusion module as the multi-modal features.
18. The apparatus according to claim 11, wherein, the processing unit processes the image to be processed according to the image editing instruction, the text description information, and the multi-modal features in the following manner to obtain a target image: Input the image editing instruction and the text description information into a large language model to obtain a semantic embedding layer, and the semantic embedding layer contains the image features of the target image; Perform noise addition processing on the image to be processed, and add the multi-modal features to the image to be processed to obtain a multi-modal image; Process the semantic embedding layer and the multi-modal image through a diffusion model and an encoder to obtain a target image.
19. The apparatus according to claim 18, wherein, the processing unit processes the semantic embedding layer and the multi-modal image through a diffusion model and an encoder in the following manner to obtain a target image: Obtain the semantic vector corresponding to the semantic embedding layer through the diffusion model, and gradually remove the noise from the multi-modal image according to the semantic vector to obtain a feature conversion image; Through the encoder, adjust the feature conversion image according to the image size of the image to be processed, and determine the feature conversion image after the size adjustment as the target image.
20. The apparatus according to claim 18, wherein, the insertion unit inserts the target image into the video to be processed in the following manner to obtain a target video: Replace the video frame image corresponding to the processed image in the video to be processed with the target image to obtain the target video.
21. A video processing apparatus, wherein, comprises: a processor: a memory for storing instructions executable by the processor; wherein, the processor is configured to: execute the video processing method according to any one of claims 1 to 10.
22. A storage medium, wherein, instructions are stored in the storage medium, and when the instructions in the storage medium are executed by a processor, the processor is enabled to execute the video processing method according to any one of claims 1 to 10.
Citation Information
Cited By
Intelligent video dynamic editing method based on diffusion model
CN120897093A