Video processing method and system, and model training method and system

By adopting a target model with a full 3D attention mechanism in video dressing technology, the video features and time-space dependence relationships are captured, and the problem of unsmooth transition between adjacent frames in video dressing is solved, improving the visual effect and user experience of the video.

CN120163716APending Publication Date: 2025-06-17ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510238452.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, during the video dressing process, the transition between adjacent frames is not smooth, resulting in poor video visual effects and affecting the user's virtual dressing experience.

Method used

By generating video features and conditional features, and using pre-trained target models, a full 3D attention mechanism is used to capture space-time dependencies to generate target videos with smooth transitions.

Benefits of technology

It realizes a smooth transition between adjacent frames, improves the visual effect of the target video, and improves the user's virtual dressing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163716A_ABST
    Figure CN120163716A_ABST
Patent Text Reader

Abstract

The invention provides a video processing method and system and a model training method and system. The video processing method comprises the following steps: generating a video feature based on an original video, wherein the original video comprises a plurality of video frames displaying original objects; replacement requirement information corresponding to the original video is obtained, condition features are generated based on the replacement requirement information, and the replacement requirement information represents that the original object in the original video is replaced with the target object. And inputting the video features and the condition features into a pre-trained target model to generate a target video through the target model, the target video being a video obtained by replacing an original object in the original video with a target object. Wherein the target model can be trained to capture the space-time dependency relationship from the video features by adopting a full 3D attention mechanism through a model training method, and the target video is generated by taking the space-time dependency relationship and the condition features as constraints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and particularly to a video processing method and system, and a model training method and system. Background Art

[0002] With the development of artificial intelligence, virtual clothing change technology has been widely applied in fields such as e-commerce, entertainment, and social media. Traditional virtual clothing change technology mainly focuses on static images. However, with the rise of short video and live broadcast platforms, virtual clothing change technology for videos has become increasingly important.

[0003] Currently, for video clothing change, a frame-by-frame clothing change method is usually adopted. That is, for each video frame in the video, the clothing change technology for static images is respectively used for clothing change processing, and then the images after clothing change are combined into a video after clothing change according to the time sequence of each video frame.

[0004] However, for the video generated by the above solution, the transition between adjacent frames is not smooth, resulting in poor visual effects of the video, thus affecting the user's virtual clothing change experience.

[0005] The content in the background art section is only the information known to the inventor personally, and does not represent that the above information has entered the public domain before the filing date of this disclosure, nor does it represent that it can become the prior art of this disclosure. Summary of the Invention

[0006] This specification provides a video processing method and system, and a model training method and system, which can use a target model to perform object replacement on an original video to obtain a target video, and in the replacement process, not only use the replacement requirement information as a condition, but also use the spatio-temporal dependence relationship in the original video as a condition, so that the transition between adjacent frames in the target video is smooth, and the visual effect of the target video is improved.

[0007] In a first aspect, this specification provides a video processing method, including: obtaining an original video, and generating video features based on the original video, where the original video includes a plurality of video frames showing an original object; obtaining replacement requirement information corresponding to the original video, and generating conditional features based on the replacement requirement information, where the replacement requirement information represents replacing the original object in the original video with a target object; and inputting the video features and the conditional features into a pre-trained target model to generate a target video through the target model, where the target video is a video obtained by replacing the original object in the original video with the target object, and where the target model is trained to: capture spatio-temporal dependence relationships from the video features using a full 3D attention mechanism, and generate the target video with the spatio-temporal dependence relationships and the conditional features as constraints.

[0008] In some embodiments, generating video features based on the original video includes: extracting first video features from the original video; generating a masked video based on the original video and extracting second video features from the masked video, where the masked video includes masked images corresponding to the respective multiple video frames, and the masked images represent the position information of the original object; and generating the video features based on the first video features and the second video features.

[0009] In some embodiments, generating the masked video based on the original video includes: determining the position information of the original object in the multiple video frames; generating masked images corresponding to the respective multiple video frames based on the position information of the original object in the multiple video frames; and generating the masked video based on the masked images corresponding to the respective multiple video frames.

[0010] In some embodiments, determining the position information of the original object in the multiple video frames includes: obtaining the object type of the original object; and using a pre-trained image recognition model to perform object recognition on the multiple video frames based on the object type to obtain the position information of the original object in the multiple video frames.

[0011] In some embodiments, determining the position information of the original object in the multiple video frames includes: obtaining the position information of the original object in a reference frame, where the reference frame is one of the multiple video frames; and using a pre-trained tracking model to perform tracking recognition on the multiple video frames based on the position information of the original object in the reference frame to obtain the position information of the original object in the multiple video frames.

[0012] In some embodiments, the method further includes: generating a depth video based on the original video and extracting third video features from the depth video, where the depth video includes depth images corresponding to the respective multiple video frames, and the depth images represent the depth information of each pixel point in the video frames; generating the video features based on the first video features and the second video features includes: generating the video features based on the first video features, the second video features, and the third video features.

[0013] In some embodiments, generating the depth video based on the original video includes: obtaining the depth images corresponding to the respective multiple video frames through a pre-trained pose recognition model; and generating the depth video based on the depth images corresponding to the respective multiple video frames.

[0014] In some embodiments, the replacement requirement information includes at least one of a reference text or a reference image, wherein the reference text at least describes the appearance information of the target object, and the reference image at least shows the appearance of the target object; the conditional feature includes at least one of a text embedding feature generated based on the reference text or an image embedding feature generated based on the reference image.

[0015] In some embodiments, obtaining the replacement requirement information corresponding to the original video includes: obtaining a processing instruction input by a user for the original video, and parsing at least one of the reference text or the reference image from the processing instruction.

[0016] In some embodiments, the image embedding feature is generated in the following manner: extracting the high-frequency feature and the low-frequency feature of the reference image, where the high-frequency feature is used to characterize the detailed information of the reference image, and the low-frequency feature is used to characterize the global information of the reference image; and fusing the high-frequency feature and the low-frequency feature to generate the image embedding feature.

[0017] In a second aspect, this specification provides a method for training a model, including: obtaining a training sample set, where each training sample in the training sample set includes: a first video, a second video, and replacement requirement information, the first video includes multiple video frames showing an original object, the replacement requirement information represents replacing the original object in the first video with a target object, and the second video is a video obtained by previously replacing the original object in the first video with the target object; and performing multiple rounds of iterative training on a target model using the training sample set, where the process of each round of iterative training includes: generating a video feature based on the first video, and generating a conditional feature based on the replacement requirement information, inputting the video feature and the conditional feature into the target model, so that the target model uses a full 3D attention mechanism to capture spatio-temporal dependencies from the video feature, and using the spatio-temporal dependencies and the conditional feature as constraints to generate a predicted video, and adjusting the parameters of the target model with the goal of minimizing the difference between the predicted video and the second video.

[0018] In some embodiments, the training sample set includes: a general sample subset, where the original objects replaced in the training samples in the general sample subset cover multiple object types, and a dedicated sample subset, where the original objects replaced in the training samples in the dedicated sample subset are all of the target type.

[0019] Performing multiple rounds of iterative training on the target model using the training sample set includes: performing multiple rounds of iterative training on the target model using the general sample subset to enable the target model to have the ability to replace various types of objects, and performing multiple rounds of iterative training on the target model using the dedicated sample subset to enable the target model to have a higher replacement ability for the target type of object than for other types of objects.

[0020] In some embodiments, with the goal of minimizing the difference between the predicted video and the second video, adjusting the parameters of the target model includes: determining the prediction loss corresponding to the predicted video and adjusting the parameters of the target model with the goal of minimizing the prediction loss, where the prediction loss includes a global loss and at least one of a detail loss or a boundary fusion loss. The detail loss is used to characterize the gap between the detail features of the target object and the conditional features in the predicted video, the global loss is used to characterize the difference between the predicted video and the second video, and the boundary fusion loss is used to characterize the degree of fusion distortion between the edge of the target object and the surrounding picture in the predicted video.

[0021] In some embodiments, generating video features based on the first video includes: extracting first video features from the first video; generating a masked video based on the first video and extracting second video features from the masked video, where the masked video includes masked images corresponding to each video frame of the first video, and the masked images represent the position information of the original object; and generating the video features based on the first video features and the second video features.

[0022] In some embodiments, the method further includes: generating a depth video based on the first video and extracting third video features from the depth video, where the depth video includes depth images corresponding to each video frame in the first video, and the depth images represent the depth information of each pixel point in the video frame; generating the video features based on the first video features and the second video features includes: generating the video features based on the first video features, the second video features, and the third video features.

[0023] In some embodiments, the replacement requirement information includes at least one of a reference text or a reference image, where the reference text at least describes the appearance information of the target object, and the reference image at least shows the appearance of the target object; the conditional features include at least one of a text embedding feature generated based on the reference text or an image embedding feature generated based on the reference image.

[0024] In a third aspect, this specification provides a video processing system, including: at least one storage medium storing at least one instruction set for video processing; and at least one processor communicatively connected to the at least one storage medium, wherein when the video processing system runs, the at least one processor reads the at least one instruction set and implements the method according to any one of the first aspects as indicated by the at least one instruction set.

[0025] In a fourth aspect, this specification provides a model training system, including: at least one storage medium storing at least one instruction set for training a model; and at least one processor communicatively connected to the at least one storage medium, wherein when the model training system runs, the at least one processor reads the at least one instruction set and implements the method according to any one of the second aspects as indicated by the at least one instruction set.

[0026] In a fifth aspect, this specification further provides a computer-readable non-volatile storage medium, wherein at least one instruction set is stored in the computer-readable non-volatile storage medium, and when the at least one instruction set is executed by at least one processor, the method provided in any one of the first aspect or the second aspect is implemented.

[0027] Some of the other functions of the video processing method and system, and the model training method and system provided in this specification will be listed in the following description. The creative aspects of the video processing method and system, and the model training method and system provided in this specification can be fully explained through practice or use of the methods, devices, and combinations described in the following detailed examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] To more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1A Shows a schematic diagram of a video clothing change scenario;

[0030] Figure 1B Shows a schematic diagram of a video processing scenario provided according to an embodiment of this specification;

[0031] Figure 2 Shows a hardware structure diagram of a computing system provided according to an embodiment of this specification;

[0032] Figure 3Shows a flowchart of a video processing method provided according to an embodiment of this specification;

[0033] Figure 4 Shows a schematic flowchart of generating a target video provided according to an embodiment of this specification;

[0034] Figure 5 Shows a flowchart of a training method of a model provided according to an embodiment of this specification; and

[0035] Figure 6 Shows a schematic diagram of training a target model provided according to an embodiment of this specification. Detailed implementation manners

[0036] The following description provides specific application scenarios and requirements of this specification, aiming to enable those skilled in the art to manufacture and use the content in this specification. For those skilled in the art, various partial modifications to the disclosed embodiments are obvious, and without departing from the spirit and scope of this specification, the general principles defined here can be applied to other embodiments and applications. Therefore, this specification is not limited to the shown embodiments, but has the broadest scope consistent with the claims.

[0037] The terms used here are only for the purpose of describing specific example embodiments and are not restrictive. For example, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" used here may also include the plural forms. When used in this specification, the terms "comprise", "include" and / or "contain" mean that the associated integers, steps, operations, elements and / or components exist, but do not exclude the existence of one or more other features, integers, steps, operations, elements, components and / or groups, or the addition of other features, integers, steps, operations, elements, components and / or groups in this system / method.

[0038] Considering the following description, these features and other features of this specification, as well as the operations and functions of the related elements of the structure, and the economy of the combination and manufacture of the components can be significantly improved. Referring to the accompanying drawings, all of these form a part of this specification. However, it should be clearly understood that the drawings are only for the purpose of illustration and description and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.

[0039] The flowcharts used in this specification show the operations implemented by the system according to some embodiments in this specification. It should be clearly understood that the operations in the flowchart may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.

[0040] The application scenarios of this specification will be introduced below.

[0041] The technical solution provided in this specification is applicable to the scenario of video content replacement. For example, the technical solution provided in this specification can be applied to the video clothing replacement scenario, that is, replacing the clothing of the person in the original video to obtain the target video. Figure 1A A schematic diagram of a video clothing replacement scenario is shown. As Figure 1A shown, the person in the original video is wearing a black long-sleeved top. After performing video clothing replacement processing on the original video, the target video is obtained. The difference between the target video and the original video is that the top of the person is replaced with a floral short-sleeved top.

[0042] It should be noted that the above video clothing replacement scenario is only one of the multiple usage scenarios applicable to this specification. The technical solution provided in this specification can not only be used to replace the clothing in the original video, but also can be applied to replace other objects in the original video. The objects to be replaced may include but are not limited to: clothing appearing in the video (such as tops, pants, shoes, etc.), body parts of a person (such as a face, legs, etc.), plants, animals, or other objects (such as vehicles, airplanes, tables, etc.), and so on. For the sake of easy understanding, when giving examples in the following text of the specification, the object to be replaced is taken as the top for illustration. Those skilled in the art should understand that when the object to be replaced is other objects, the implementation methods and technical effects are similar.

[0043] This specification provides a video processing method, which can be executed by a video processing system. The video processing system can obtain the original video and generate video features based on the original video. The original video includes multiple video frames showing the original object. The video processing system can also obtain the replacement requirement information corresponding to the original video and generate conditional features based on the replacement requirement information. The replacement requirement information represents replacing the original object in the original video with a target object. Furthermore, the video processing system can input the video features and the conditional features into a pre-trained target model to generate a target video through the target model. The target video is a video obtained by replacing the original object in the original video with the target object. Among them, the target model is trained to: capture spatio-temporal dependence relationships from the video features by using a full 3D attention mechanism, and generate the target video with the spatio-temporal dependence relationships and the conditional features as constraints.

[0044] In the video processing method provided in this specification, the target model is trained to capture spatio-temporal dependencies from the video features using a full 3D attention mechanism, and generate the target video with the spatio-temporal dependencies and the conditional features as constraints. Since the pre-trained target model has the ability to capture spatio-temporal dependencies from the video features using a full 3D attention mechanism, this enables the target model to refer to the spatio-temporal dependencies of the original video when generating the target video, thereby ensuring a high spatio-temporal consistency between consecutive frames. The target video generated in this way can ensure smooth transitions between consecutive frames and enhance the visual effect of the target video.

[0045] Figure 1B FIG. shows a schematic diagram of a video processing scenario provided according to an embodiment of this specification. As Figure 1B shown, the scenario 100 may include a model training system 11 and a video processing system 12.

[0046] In the model training phase, the model training system 11 can be used to train the target model. In some embodiments, the model training system 11 can perform multiple rounds of iterative training on the target model according to a training sample set, adjust the parameters of the target model, and train the target model to capture spatio-temporal dependencies from video features using a full 3D attention mechanism, and generate the target video with the spatio-temporal dependencies and conditional features as constraints. At this time, the model training system 11 can store data or instructions for performing the model training method described in this specification, and can execute or be used to execute the data or instructions. In some embodiments, the model training system 11 can include a hardware device with data information processing capabilities and the necessary programs for driving the hardware device to work.

[0047] In the video replacement phase, the video processing system 12 can be used to execute the video processing method. In some embodiments, the video processing system 12 can obtain the original video and the replacement requirement information corresponding to the original video, extract video features based on the original video, and generate conditional features based on the replacement requirement information, and then input the video features and conditional features into the pre-trained target model obtained in the model training phase, and generate the target video through the target model. At this time, the video processing system 12 can store data or instructions for performing the video processing method described in this specification, and can execute or be used to execute the data or instructions. In some embodiments, the video processing system 12 can include a hardware device with data information processing capabilities and the necessary programs for driving the hardware device to work.

[0048] In some embodiments, the model training system 11 and the video processing system 12 may correspond to a single computing device or a computing cluster composed of multiple computing devices. For example, the model training system 11 may be a server or a server cluster, and the video processing system 12 may be a mobile device, a tablet computer, a laptop computer, a server, or the like, or any combination thereof.

[0049] It should be noted that the above-mentioned model training system 11 and video processing system 12 may correspond to the same physical device or different physical devices, and this specification does not limit this.

[0050] Figure 2 The hardware structure diagram of a computing system 200 provided according to an embodiment of this specification is shown. The computing system 200 can be used as Figure 1B the model training system 11 or the video processing system 12 in, and execute the video processing method or the model training method described in this specification.

[0051] As Figure 2 shown, the computing system 200 may include at least one storage medium 230 and at least one processor 220. In some embodiments, the computing system 200 may further include a communication port 250 and an internal communication bus 210. The computing system 200 may further include I / O components 260.

[0052] The internal communication bus 210 can connect different system components. For example, the internal communication bus 210 can connect the storage medium 230, the processor 220, the communication port 250, and the I / O components 260, etc.

[0053] The I / O components 260 support input / output between the computing system 200 and other components.

[0054] The communication port 250 is used for data communication between the computing system 200 and the outside world. For example, the communication port 250 can be used for data communication between the computing system 200 and the network. The communication port 250 can be a wired communication port or a wireless communication port.

[0055] The storage medium 230 may include a data storage device. The data storage device may be a non-transitory storage medium or a transitory storage medium. For example, the data storage device may include one or more of a magnetic disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 235. The storage medium 230 further includes at least one instruction set stored in the data storage device. The instruction set may include computer program code, and the computer program code may include programs, routines, objects, components, data structures, processes, modules, and so on.

[0056] At least one processor 220 may be communicatively connected to at least one storage medium 230. When the computing system 200 is running, the at least one processor 220 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the video processing method or the model training method provided in this specification. The processor 220 may execute the steps included in the video processing method or the model training method. The processor 220 may be in the form of one or more processors. In some embodiments, the processor 220 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of executing one or more functions, etc., or any combination thereof.

[0057] For illustrative purposes only, only one processor 220 is shown in the computing system 200 in the drawings. However, it should be noted that the computing system 200 in this specification may also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification may be executed by one processor or jointly executed by multiple processors. For example, if it is described in this specification that the processor 220 of the computing system 200 executes step A and step B, it should be understood that step A and step B may also be jointly or separately executed by two different processors 220 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).

[0058] Figure 3 A flowchart showing a video processing method provided according to an embodiment of this specification is shown. As before, the video processing system 12 may execute the video processing method of this specification.

[0059] As Figure 3 shown, the video processing method may include S310 - S330.

[0060] S310: Obtain the original video and generate video features based on the original video, where the original video includes multiple video frames showing the original object.

[0061] In some embodiments, the original video can be pre-shot or captured in real time. The acquisition method of the original video can be determined according to the actual application scenario. For example, in some application scenarios, the video processing system can receive the pre-shot original video uploaded by the user and generate a target video corresponding to the original video for the user to download. In other application scenarios, the video processing system can be used as a live broadcast system, execute the video processing method based on the received live video, and display the target video during the live broadcast.

[0062] At least some video frames in the original video show an original object. The original object can be any object, including but not limited to body parts of a person, clothing of a person, objects, animals, plants, etc. The video processing method provided in this specification is used to replace the original object in the original video with a target object. The replaced target object can be of the same type as the original object or different, and this specification does not make any limitations in this regard.

[0063] Figure 4 A schematic flowchart of generating a target video according to an embodiment of this specification is shown.

[0064] In some embodiments, referring to Figure 4 , the video processing system can perform feature extraction on the original video to obtain a first video feature. For example, the video processing system can perform feature extraction on the original video through the encoder part of a pre-trained video variational autoencoder to obtain the first video feature. The video variational autoencoder includes an encoder part and a decoder part. The encoder part is used to compress the spatio-temporal information of the input video (original video) into a low-dimensional latent space through a neural network to obtain the probability distribution of the original video mapped to the latent space (i.e., the first video feature), and the decoder part is used to sample from the latent space and generate output data (i.e., the target video). Among them, the latent space is a low-dimensional space used to represent data in machine learning. The latent space usually has a lower dimension than the original data space to facilitate the analysis and understanding of the data in the latent space.

[0065] In some embodiments, the first video feature can characterize the spatial feature and the temporal feature of the original video. Among them, the spatial feature of the original video reflects the spatial relationship between different pixels within a video frame, and the temporal feature of the original video reflects the feature of the dynamic change of consecutive video frames.

[0066] In some embodiments, the video processing system can directly use the first video feature as the video feature in S310.

[0067] In some embodiments, the video processing system may also generate a masked video based on the original video, and perform feature extraction on the masked video to obtain second video features. Among them, the masked video includes masked images corresponding to multiple video frames respectively, and the masked images represent the position information of the original object. In this case, the video processing system may also generate the video features in S310 based on the first video features and the second video features.

[0068] In some embodiments, the video processing system may adopt the following method when generating the masked video: The video processing system first determines the position information of the original object in multiple video frames, and generates masked images corresponding to multiple video frames respectively based on the position information of the original object in multiple video frames. Furthermore, the video processing system generates a masked video based on the masked images corresponding to multiple video frames respectively.

[0069] The position information of the original object in the video frame represents the area range occupied by the original object in the video frame. For example, the position information of the original object in the video frame may include the coordinates of the original object. After the video processing system determines the position information of the original object in multiple video frames, it may generate a binary image with the same size as the video frame as the masked image. In the masked image, the position information of the original object may be represented by pixel values. For example, the area corresponding to the original object may be white (marking its pixel value as 1), and the background area (i.e., the area where the original object does not exist) may be black (marking its pixel value as 0). In this way, the range and coordinates corresponding to the white area in the masked image are the position information of the original object. The video processing system sorts the masked images according to the time sequence of the corresponding video frames, and then can generate a masked video.

[0070] In some embodiments, the video processing system may adopt different methods to determine the position information of the original object in multiple video frames according to whether there is a reference frame when acquiring the original video. The reference frame is one of the multiple video frames of the original video. When acquiring the reference frame, the pre-marked position information of the original object in the reference frame can also be acquired.

[0071] In some embodiments, when the video processing system acquires the original video, it can also acquire a reference frame and the position information of the original object in the reference frame. As an example, the reference frame can be uploaded together with the original video by the user. The reference frame can include the position information of the original object pre-marked by the user. The video processing system can directly use the position information of the original object pre-marked by the user as the position information of the original object in the reference frame. Alternatively, the reference frame can also be the first video frame of the original video or a video frame that meets a preset condition (such as a video frame that meets a preset clarity score or a preset brightness score). In this case, the video processing system can, based on the method in the above example, acquire the position information of the original object in the reference frame through a pre-trained image recognition model (such as Mask R-CNN). Then, the video processing system can use the pre-trained tracking model to perform tracking and recognition on multiple video frames based on the position information of the original object in the reference frame, and obtain the position information of the original object in multiple video frames.

[0072] As an example, the pre-trained tracking model can be an improved Siamese Region Proposal Network++ (SiamRPN++) model. SiamRPN++ includes two branches: a template branch and a search branch. These two branches perform feature extraction through a Siamese network. The template branch is used to extract features from the position information of the original object in the reference frame, while the search branch is used to extract features from video frames other than the reference frame. Then, SiamRPN++ performs feature matching and target localization in the features extracted by the search branch based on the features extracted by the template branch through a Region Proposal Network (RPN) to achieve tracking and recognition, and obtain the position information of the original object in multiple video frames.

[0073] In some embodiments, if the user does not specify a reference frame when uploading the original video, the video processing system can also use other methods to acquire the object type of the original object. For example, the object type of the original object can be specified by the user. For example, when uploading the original video, the user can provide the object type of the original object in a text description or other ways to instruct the video processing system to process the original video according to the object type provided by the user. Another example is that the object type of the original object can also be determined according to the application program that calls the video processing system. For example, if the current application program that calls the video processing system is a video clothing-changing application, the object type of the original object can be determined as clothing.

[0074] Then, the video processing system can use a pre-trained image recognition model to perform object recognition on multiple video frames based on the object type, and obtain the position information of the original object in the multiple video frames. For example, if the object type is clothing, each video frame can be detected by a pre-trained object detection model to obtain the position information of the clothing in each video frame. As an example, the pre-trained image recognition model can be a Mask Region-based Convolutional Neural Network (Mask R-CNN). Mask R-CNN is an image segmentation algorithm. When the object type is clothing, Mask R-CNN can generate a pixel-level segmentation mask for each piece of clothing detected in the video frame, that is, the specific shape and position of the clothing in the video frame can be obtained.

[0075] In some embodiments, when the video processing system extracts the second video feature from the masked video, it can also be implemented by using the encoder part of a pre-trained video variational autoencoder. That is, the video processing system can extract the spatial feature and the temporal feature of the masked video as the second video feature through the encoder part. The specific implementation method is similar to the method of extracting the first video feature described above and will not be elaborated here.

[0076] In some embodiments, after obtaining the first video feature and the second video feature, the video processing system can also concatenate the first video feature and the second video feature according to the channel dimension of the features to generate a video feature. For example, assume that the first video feature includes 3 channels, and the length of each channel is 1024, then the channel dimension of the first video feature can be denoted as 3*1024. The second video feature includes 2 channels, and the length of each channel is 1024, then the channel dimension of the second video feature can be denoted as 2*1024. Then the video feature generated by the video processing system by concatenating the first video feature and the second video feature according to the channel dimension includes 5 channels, and the length of each channel is 1024, denoted as 5*1024.

[0077] When the object type of the original object is clothing, since the clothing is worn on the human body, therefore, when replacing the original video, the depth information of the human body in the video frame can also be referred to to make the replacement result more realistic. Therefore, in some embodiments, the video processing system can also generate a depth video based on the original video, and extract the third video feature from the depth video, where the depth video includes depth images corresponding to multiple video frames respectively, and the depth image represents the depth information of each pixel point in the video frame. Correspondingly, the video processing system can generate a video feature based on the first video feature, the second video feature and the third video feature.

[0078] In some embodiments, the video processing system may process each video frame of the original video through a pre-trained pose recognition model to obtain depth images corresponding to the respective video frames. Then, the video processing system sorts the depth images according to the time sequence of the corresponding video frames to generate a depth video.

[0079] Among them, the pre-trained pose recognition model may be a Dense Pose model. The Dense Pose model can predict the position of each pixel on the three-dimensional human model from the video frame and map the human surface in the video frame to a standardized three-dimensional grid, thereby obtaining the three-dimensional coordinates of each pixel. Furthermore, the Dense Pose model can calculate the distance between each pixel in the video frame and the shooting device (i.e., depth information, also known as depth value) based on the three-dimensional coordinates of each pixel.

[0080] In some embodiments, the video processing system may generate a corresponding depth image based on the depth information of each pixel point in each video frame. Among them, the depth image may be a grayscale image or a red, green, blue (RGB) image.

[0081] As an example, when the depth image is a grayscale image, the video processing system may normalize the depth value to the range of [0, 255] and generate a grayscale image according to the depth value corresponding to each pixel. In the grayscale image, each grayscale value corresponds to a depth value. The smaller the depth value (i.e., the closer the pixel in the image is to the shooting device), the brighter grayscale value can be used to represent it, and the larger the depth value (i.e., the farther the pixel in the image is from the shooting device), the darker grayscale value can be used to represent it. As an example, when the depth image is an RGB image, the video processing system may map the depth value to the red, green, and blue channels so that different depth values are mapped to different colors. For example, the smaller the depth value, the warmer color can be used to represent it (such as red, yellow), and the larger the depth value, the colder color (such as blue, green) can be used to represent it.

[0082] In some embodiments, when the video processing system extracts the third video feature from the depth video, it can also be implemented by using the encoder part of the pre-trained video variational autoencoder. That is, the video processing system can extract the spatial feature of the depth video and the temporal feature of the depth video as the third video feature through the encoder part. The specific implementation method is similar to the method of extracting the first video feature described above and will not be elaborated here.

[0083] In some embodiments, the method used by the video processing system to generate video features based on the first video feature, the second video feature, and the third video feature is similar to that when the video processing system generates video features based on the first video feature and the second video feature, and will not be elaborated here either.

[0084] S320: Obtain the replacement requirement information corresponding to the original video, and generate conditional features based on the replacement requirement information. The replacement requirement information represents replacing the original object in the original video with a target object.

[0085] In some embodiments, the replacement requirement information includes at least one of a reference text or a reference image. The video processing system can obtain the processing instruction input by the user for the original video, and parse at least one of the reference text or the reference image from the processing instruction. Among them, the reference text at least describes the appearance information of the target object, and the reference image at least shows the appearance of the target object.

[0086] In some embodiments, the user can describe the processing instruction for the original video through text. As an example, the processing instruction can completely describe the user's request. For example, processing instruction 1 can be "Please replace the upper garment of the person in the original video with a floral short-sleeved shirt". Or, the processing instruction can also only provide the appearance information of the target object. For example, processing instruction 2 can be "floral short-sleeved shirt". In some other examples, the processing instruction can also indicate to use the uploaded image as a reference image to process the original video. For example, processing instruction 3 can be "Please replace the upper garment of the person in the original video with the upper garment in the reference image". Or, the processing instruction can also not include text content. For example, processing instruction 4 can be the uploaded reference image.

[0087] In some embodiments, the video processing system can parse the reference text "floral short-sleeved shirt" according to processing instruction 1 and processing instruction 2; the video processing system can parse the reference image according to processing instruction 3 and processing instruction 4.

[0088] In some embodiments, the conditional features include at least one of a text embedding feature generated based on the reference text or an image embedding feature generated based on the reference image. The text embedding feature or the image embedding feature is used to represent the appearance features of the target object.

[0089] In some embodiments, when the video processing system parses the reference text, it can generate a text embedding feature based on the reference text. Or, the video processing system can also obtain the corresponding reference image based on the parsed reference text, and then generate an image embedding feature based on the reference image. For example, when the reference text is "floral short-sleeved shirt", the video processing system can generate a text embedding feature based on the text description of "floral short-sleeved shirt". Or, the video processing system can obtain a reference image including "floral short-sleeved shirt" from the database according to the description of "floral short-sleeved shirt", and generate an image embedding feature based on the reference image.

[0090] In some embodiments, the text embedding feature can be generated in any of the following ways:

[0091] Method 1: The video processing system can input the reference text into a pre-trained inference large model and instruct the large model to generate corresponding text embedding features based on the reference text.

[0092] Method 2: The video processing system can obtain the feature data corresponding to the reference text from the database according to the reference text. For example, multiple groups of feature data can be stored in the database, and each group of feature data is used to characterize a text embedding feature. The video processing system obtains the relevance between the reference text and each group of feature data, and uses the group of feature data with the highest relevance as the text embedding feature.

[0093] In some embodiments, the image embedding features are generated in the following manner: The video processing system extracts the high-frequency features and low-frequency features of the reference image. The high-frequency features are used to characterize the detailed information of the reference image, and the low-frequency features are used to characterize the global information of the reference image. The video processing system fuses the high-frequency features and low-frequency features to generate the image embedding features.

[0094] As an example, the video processing system can extract the low-frequency features (i.e., global information) of the reference image through the encoder part of a pre-trained video variational autoencoder, and extract the high-frequency features (i.e., detailed information) of the reference image through a pre-trained Knowledge Distillation with No Labels V2 (DiNOv2). Then, the video processing system can concatenate the low-frequency features and high-frequency features of the reference image, and learn the concatenated low-frequency features and high-frequency features through a Multilayer Perceptron (MLP) to obtain the image embedding features. Among them, the method used by the video processing system to concatenate the low-frequency features and high-frequency features of the reference image is similar to that when the video processing system generates video features based on the first video feature and the second video feature, which will not be elaborated here.

[0095] In this embodiment, by parsing at least one of the reference text or the reference image from the processing instruction, and obtaining the text embedding features or the image embedding features based on the parsed reference text or reference image to characterize the appearance features of the target object, the multi-modal (text or image) information can be effectively converted into conditional features. Furthermore, when the conditional features are used for the replacement task, the appearance features of the target object can be provided to optimize the appearance and details of the replaced target object, making the replacement effect more realistic.

[0096] S330: Input the video features and conditional features into a pre-trained target model to generate a target video through the target model. The target video is a video obtained by replacing the original object in the original video with a target object. Among them, the target model is trained to: capture spatio-temporal dependencies from the video features using a full 3D attention mechanism, and generate the target video with the spatio-temporal dependencies and conditional features as constraints.

[0097] Among them, the target model is a pre-trained model with the ability to replace objects of a target type. Among them, the object types supported by the target model for replacement are the same as the object type of the original object. For example, if the object type of the original object is clothing, the target model is a model for video clothing replacement.

[0098] For the training method of the target model, please refer to the relevant description later, and it will not be elaborated here for the time being. After the video processing system inputs the video features and conditional features into the target model, the target model can capture spatio-temporal dependencies from the video features using a full 3D attention mechanism, and generate the target video with the spatio-temporal dependencies and conditional features as constraints.

[0099] Among them, the full 3D attention mechanism means that when processing video features, the target model can simultaneously focus on the features in the three dimensions of height, width, and time sequence in the video features. By comprehensively considering the features in these three dimensions, the target model can capture spatio-temporal dependencies from the video features, so as to achieve accurate content replacement. For example, the full 3D attention mechanism can calculate the corresponding spatial attention weights and temporal attention weights according to the video features, and then combine the spatial attention and temporal attention to obtain a comprehensive attention weight (i.e., spatio-temporal dependency). The comprehensive attention weight is used to represent the importance of the objects in the video at different positions and different time points. The target model can adjust the processing method for each video frame according to the comprehensive attention weight to achieve accurate video content replacement.

[0100] In some embodiments, the target model may adopt a diffusion model. The diffusion model includes two processes when generating a video, namely, the diffusion process and the denoising process. Among them, in the diffusion process, noise is gradually added to the data distribution of the video features, so that the data distribution of the video features is transformed into a simple distribution (such as a Gaussian distribution). The denoising process refers to gradually removing noise from the Gaussian distribution to restore the data distribution. The diffusion model can learn the noise change law of the data distribution of the video features of the original video based on a full 3D spatio-temporal attention mechanism. Then, it gradually removes noise from the Gaussian noise to generate a target video that meets the conditions, where these conditions may include spatio-temporal dependencies and conditional features. The diffusion model can inject spatio-temporal dependencies and conditional features into the full 3D spatio-temporal attention mechanism through conditional embedding (for example, spatio-temporal dependencies and conditional features can be encoded as embedding vectors and combined with the video features), so that the full 3D spatio-temporal attention mechanism can perceive spatio-temporal dependencies and conditional features.

[0101] In this way, during the video generation process, the diffusion model can use spatio-temporal dependencies and conditional features as constraints to guide the video generation. In this way, the generated target video can, on the one hand, meet the replacement requirements, and on the other hand, ensure a high spatio-temporal consistency between consecutive video frames, thus avoiding the problem of insufficiently smooth transitions between consecutive video frames.

[0102] The following Figure 4 briefly describes the structure of the target model and its video generation process.

[0103] As Figure 4 shown, the target model may include multiple sequentially connected full 3D attention conversion layers, and each full 3D attention conversion layer is used to implement a diffusion conversion process based on full 3D attention. The video processing system inputs the video features into the first full 3D attention conversion layer and processes them through each layer in turn until the last layer outputs the target video. The video processing system inputs the conditional features into each full 3D attention conversion layer respectively for self-attention control.

[0104] In the technical solution provided in this specification, the target model can perform content replacement processing in units / granularities of multiple consecutive video frames at a time. That is to say, the target model does not perform content replacement processing in units / granularities of individual video frames, but in units / granularities of videos (multiple consecutive video frames). This enables the target model to use the spatio-temporal dependencies in the original video as constraints during the video generation process.

[0105] In some cases, when the number of video frames contained in the original video is large, due to the limited processing performance of the target model, if the target model processes the original video as a whole, the processing efficiency will be reduced. Therefore, when the number of video frames contained in the original video is large, the video processing system can segment the original video according to the processing performance and the resolution of the original video, and process a part of the video frames of the original video each time to improve the overall processing speed.

[0106] In some embodiments, the video processing system can split the original video into M sub-videos. Each sub-video includes N consecutive video frames. Both M and N are integers greater than 1, and the value of N can be determined based on the processing performance of the target model and the resolution of the original video. The video processing system can sequentially perform content replacement processing on the M sub-videos respectively to obtain the replaced sub-videos corresponding to the M sub-videos. Further, the replaced sub-videos are spliced to obtain the target video. Among them, the processing process of each sub-video by the video processing system is similar to the description above and will not be elaborated here.

[0107] In summary, in the video processing method provided in this specification, the video processing system generates video features based on the original video, generates conditional features based on the replacement requirement information, and then inputs the video features and the conditional features into a pre-trained target model. The target video is generated by the target model. Since the pre-trained target model has the ability to capture spatio-temporal dependencies from the video features by using a full 3D attention mechanism, this enables the target model to refer to the spatio-temporal dependencies of the original video when generating the target video, thereby improving the spatio-temporal consistency between consecutive frames. The target video generated in this way can ensure smooth transitions between consecutive frames and has a good visual effect.

[0108] The above describes the process of the video processing system using the target model to replace the original object in the original video. The following will be combined with Figure 5 to illustrate the training process of the target model.

[0109] Figure 5 FIG. shows a flowchart of a method for training a model according to an embodiment of this specification. As before, the model training system 11 can execute the method for training the model of this specification.

[0110] As Figure 5 shown, the method for training the model may include:

[0111] S510: Obtain a training sample set. Each training sample in the training sample set includes: a first video, a second video, and replacement requirement information. The first video includes multiple video frames showing an original object. The replacement requirement information represents replacing the original object in the first video with a target object. The second video is a video obtained by previously replacing the original object in the first video with the target object.

[0112] Among them, the first video in each training sample can also be called the video before replacement, and the second video can also be called the video after replacement.

[0113] In some embodiments, the training sample set includes: a general sample subset, where the original objects replaced in the training samples in the general sample subset cover multiple object types, and a dedicated sample subset, where the original objects replaced in the training samples in the dedicated sample subset are all of the target type.

[0114] Among them, when the model training system trains the target model based on the general sample subset, it can improve the generalization ability of the target model and achieve the replacement of multiple types of objects. As an example, the original objects in the general sample subset can include people, animals, vehicles, items, clothing, etc. When the model training system trains the target model based on the dedicated sample subset, it can strengthen the processing ability of the original objects of the target type. As an example, when the target model is used for video clothing replacement, the original objects of the target type are clothing, and the original objects replaced in the training samples in the dedicated sample subset are all clothing.

[0115] S520: Perform multiple rounds of iterative training on the target model using the training sample set. Among them, the process of each round of iterative training includes: generating video features based on the first video, generating conditional features based on the replacement requirement information, inputting the video features and the conditional features into the target model, so that the target model uses a full 3D attention mechanism to capture spatio-temporal dependencies from the video features, and using the spatio-temporal dependencies and the conditional features as constraints to generate a predicted video, and aiming to minimize the difference between the predicted video and the second video, adjusting the parameters of the target model.

[0116] Figure 6 Shows a schematic diagram of training a target model according to an embodiment of the present specification.

[0117] In some embodiments, referring to Figure 6 , when the training sample set includes a general sample subset and a dedicated sample subset, the model training system performing multiple rounds of iterative training on the target model using the training sample set can include two stages:

[0118] The first stage: The model training system performs multiple rounds of iterative training using the target model so that the target model has the ability to replace multiple types of objects (i.e., a general target model).

[0119] Second stage: The model training system performs multiple rounds of iterative training on the target model using a dedicated sample subset, so that the target model has a higher replacement ability for objects of the target type than for objects of other types (i.e., the dedicated target model).

[0120] Among them, in the first stage, the target model to be trained can be a base model with the ability to capture spatio-temporal dependencies from video features using a full 3D attention mechanism. As an example, the base model can be a Conditional Diffusion Transformer using a full 3D attention mechanism. The model training system uses a general sample subset to train the base model, and a target model with the ability to replace various types of objects can be obtained. In the second stage, the target model to be trained can be the target model obtained from the training in the first stage and having the ability to replace various types of objects. The model training system uses a dedicated sample subset to train the target model with the ability to replace various types of objects, and a target model with a higher replacement ability for objects of the target type than for objects of other types can be obtained.

[0121] In the first stage of model training, the model training system can adopt the following method to train the base model through a general sample subset:

[0122] When the model training system trains the target model to be trained, it can extract features from the first video to obtain the first video feature; generate a masked video based on the first video, and extract features from the masked video to obtain the second video feature, where the masked video includes masked images corresponding to each video frame of the first video, and the masked images represent the position information of the original object; and generate a video feature based on the first video feature and the second video feature.

[0123] Among them, the method used by the model training system to extract features from the first video to obtain the first video feature is similar to the method used by the video processing system to extract features from the original video to obtain the first video feature; the method used by the model training system to generate a masked video based on the first video and extract features from the masked video to obtain the second video feature is similar to the method used by the video processing system to generate a masked video based on the original video and extract features from the masked video to obtain the second video feature; the method used by the model training system to generate a video feature based on the first video feature and the second video feature is similar to the method used by the video processing system to generate a video feature based on the first video feature and the second video feature, which will not be elaborated here.

[0124] In some embodiments, the model training system may also generate a depth video based on the first video, and extract features from the depth video to obtain third video features. The depth video includes depth images corresponding to each video frame in the first video, and the depth images represent the depth information of each pixel point in the video frame. Generating video features based on the first video feature and the second video feature includes: generating video features based on the first video feature, the second video feature, and the third video feature.

[0125] Among them, the method by which the model training system generates a depth video based on the first video and extracts features from the depth video to obtain third video features is similar to the method by which the video processing system generates a depth video based on the original video and extracts features from the depth video to obtain third video features, and will not be elaborated here.

[0126] In some embodiments, the replacement requirement information includes at least one of a reference text or a reference image. The reference text at least describes the appearance information of the target object, and the reference image at least shows the appearance of the target object. The conditional features include: at least one of a text embedding feature generated based on the reference text or an image embedding feature generated based on the reference image. Among them, the method by which the model training system generates conditional features based on the replacement requirement information is similar to the method by which the video processing system generates conditional features based on the replacement requirement information, and will not be elaborated here.

[0127] In some embodiments, the model training system adjusts the parameters of the target model with the goal of minimizing the difference between the predicted video and the second video, including: determining the prediction loss corresponding to the predicted video, and adjusting the parameters of the target model with the goal of minimizing the prediction loss.

[0128] In some embodiments, the prediction loss includes a global loss, and the global loss is used to characterize the difference between the predicted video and the second video.

[0129] In some embodiments, in addition to the global loss, the prediction loss further includes at least one of a detail loss or a boundary fusion loss. The detail loss is used to characterize the gap between the detail features of the target object in the predicted video and the conditional features. The boundary fusion loss is used to characterize the degree of fusion distortion between the edge of the target object in the predicted video and the surrounding picture.

[0130] In some embodiments, the global loss can measure the overall difference between the predicted video and the second video, and usually uses a pixel-level or feature-level error metric. The smaller the global loss, the closer the predicted video is to the second video. As an example, the model training system can use the L2 loss as the global loss to compare the pixel value differences of each frame between the predicted video and the second video. The L2 loss can be expressed by the following formula:

[0131]

[0132] Among them, the predicted video and the second video include the same number of video frames, and T is the total number of video frames. PredictedFrame t represents the t-th frame in the predicted video, and RealFrame t represents the t-th frame in the second video.

[0133] In some embodiments, the detail loss can measure the difference between the detail features (such as texture, color, material, etc.) of the target object in the predicted video and the detail features of the target object in the second video. The smaller the detail loss, the closer the detail features of the target object in the predicted video are to the detail features of the target object in the second video. As an example, the model training system can use the texture loss as the detail loss, and the texture loss can compare the texture differences of the corresponding video frames in the predicted video and the second video under different scales or filters. Alternatively, the model training system can use the Structural Similarity Index (SSIM) loss as the detail loss, and the structural similarity loss can measure the consistency of the corresponding video frames in the predicted video and the second video in dimensions such as structure, texture, and illumination. The Structural Similarity Index (SSIM) loss can be expressed by the following formula:

[0134] SSIMLoss = 1 - SSIM(I pred - I real )

[0135] where, I pred represents the video frame of the predicted video, and I real represents the video frame of the second video, and SSIM is an index used to measure the similarity of image structure and texture.

[0136] In some embodiments, the boundary fusion loss can measure the boundary fusion between the edges of the target object in the predicted video and the surrounding video frames. The smaller the boundary fusion loss, the more natural the transition between the edges of the target object and the surrounding video frames. As an example, the model training system can use any one of the Gradient Loss, Boundary Similarity Loss, or Mask Loss as the boundary fusion loss. Among them, the Gradient Loss can measure the fusion effect between the target object and the surrounding video frames by calculating the gradient change of the edges of the target object in the predicted video. The smaller the gradient change of the edges of the target object (i.e., the more natural the edges of the target object and the surrounding video frames), the smaller the Gradient Loss; the Boundary Similarity Loss can extract the boundaries of the target object through a segmentation network and calculate the similarity between the boundaries of the target object in the predicted video and the boundaries of the target object in the second video. The smaller the similarity between the two, the smaller the Boundary Similarity Loss; the Mask Loss can calculate the cross-entropy loss between the target object and the surrounding video frames based on the mask image in the mask video. The smaller the cross-entropy loss, the smaller the Mask Loss.

[0137] In some embodiments, when the prediction loss includes the global loss, the detail loss, and the boundary fusion loss, the prediction loss (L total ) can be represented by the following formula:

[0138] L total = λ1·L global + λ2·L detail + λ3·L boundaty

[0139] where L global represents the global loss, L detail represents the detail loss, L boundaty represents the boundary fusion loss, and λ1, λ2, and λ3 respectively represent the weights of the corresponding losses.

[0140] In some embodiments, the model training system can adjust the parameters of the target model according to the prediction loss in each round of iterative training and use the adjusted target model for the next round of iterative training until the target loss reaches the expectation or the preset number of iterative rounds to obtain a target model with the ability to replace various types of objects.

[0141] In some embodiments, the method for the model training system to use dedicated samples for the second-stage training is similar to the method for using a subset of general samples for the first-stage training, which will not be elaborated here.

[0142] It should be noted that in different stages, the training direction of the target model can be finely tuned by adjusting the weights of each loss in the prediction loss to achieve a better generation effect.

[0143] In this embodiment, through a stage-by-stage training design, the target model can start from processing general object replacement tasks, gradually focus on clothing replacement tasks, and finally achieve a realistic dynamic clothing replacement effect. In this process, by setting the global loss, detail loss, and boundary fusion loss, it can be ensured that when training the target model, the overall performance of the video, the details of the target object, and the natural transition of the boundary are optimized simultaneously. By setting the weights corresponding to different losses, when training the target model, the key points of model optimization can be controlled, the generation quality in all aspects can be balanced, so that the final predicted video can minimize the difference from the second video, retain fine details, ensure the seamless fusion of the object and the background, further optimize the appearance and details of the replaced target object, and make the replacement effect more realistic.

[0144] In summary, in the model training method provided in this specification, the model training system uses the training sample set to train the target model, and trains the target model to: capture spatio-temporal dependencies from the video features using a full 3D attention mechanism, and generate the target video with the spatio-temporal dependencies and the conditional features as constraints. Since the trained target model has the ability to capture spatio-temporal dependencies from the video features using a full 3D attention mechanism, this enables the target model to refer to the spatio-temporal dependencies of the original video when generating the target video, thereby improving the spatio-temporal consistency between consecutive frames. That is to say, the target video generated by the target model can ensure smooth transitions between consecutive frames and has a high visual effect.

[0145] On the other hand, this specification provides a computer-readable non-transitory storage medium storing at least one instruction set for video processing or model training. When the at least one instruction set is executed by a processor, the at least one instruction set instructs the processor to perform the steps of the video processing method or the model training method described in this specification. In some possible implementation manners, various aspects of this specification can also be implemented in the form of a program product, which includes program code. When the program product runs on a computing system 200, the program code is used to cause the computing system 200 to perform the steps of the video processing method or the model training method described in this specification. The program product for implementing the above method can be a portable compact disc read-only memory (CD-ROM) including program code and can run on the computing system 200. However, the program product of this specification is not limited thereto. In this specification, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system. The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program used by or combined with an instruction execution system, apparatus, or device. The program code included on the readable storage medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the computing system 200, partially on the computing system 200, executed as an independent software package, partially on the computing system 200 and partially on a remote computing device, or entirely on a remote computing device.

[0146] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require a particular order or a sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0147] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented only by way of example and may not be limiting. Although not explicitly stated herein, those skilled in the art will understand that this specification is intended to encompass various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.

[0148] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with that embodiment may be included in at least one embodiment of this specification. Thus, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.

[0149] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, figure, or its description. However, this does not mean that the combination of these features is necessary, and it is entirely possible for those skilled in the art, when reading this specification, to mark out some of these devices as separate embodiments for understanding. That is to say, the embodiments in this specification can also be understood as the integration of multiple sub - embodiments. And the content of each sub - embodiment is also valid when it has fewer features than all the features of a single foregoing disclosed embodiment.

[0150] Each patent, patent application, publication of patent application, and other materials cited herein, such as articles, books, specifications, publications, documents, items, etc., except those inconsistent with or conflicting with this document, or those having a limiting effect on the broadest scope of the claims, may be incorporated herein by reference and used for all purposes now or hereafter associated with this document. Further, in the event of any inconsistency or conflict between the description, definition, and / or use of relevant terms in any material and those in this document, the terms in this document shall prevail.

[0151] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.

Claims

1. A video processing method, comprising: Obtaining an original video, and generating video features based on the original video, wherein the original video includes a plurality of video frames showing an original object; Obtaining replacement requirement information corresponding to the original video, and generating a conditional feature based on the replacement requirement information, wherein the replacement requirement information represents replacing the original object in the original video with a target object; as well as The video features and the conditional features are input into a pre-trained target model to generate a target video through the target model, wherein the target video is a video obtained by replacing the original object in the original video with the target object. The target model is trained to capture spatiotemporal dependencies from the video features using a full 3D attention mechanism, and generate the target video using the spatiotemporal dependencies and the conditional features as constraints.

2. The method according to claim 1, wherein: Generating video features based on the original video includes: Extracting features from the original video to obtain a first video feature; generating a mask video based on the original video, and performing feature extraction on the mask video to obtain a second video feature, wherein the mask video includes mask images corresponding to each of the plurality of video frames, and the mask images represent position information of the original object; and The video feature is generated based on the first video feature and the second video feature.

3. The method according to claim 2, wherein: The step of generating a mask video based on the original video comprises: Determining position information of the original object in the plurality of video frames; Based on the position information of the original object in the multiple video frames, generating mask images corresponding to each of the multiple video frames; and The mask video is generated based on the mask images corresponding to each of the plurality of video frames.

4. The method according to claim 3, wherein: The determining the position information of the original object in the plurality of video frames comprises: Obtaining the object type of the original object; and The pre-trained image recognition model is used to perform object recognition on the multiple video frames based on the object type to obtain position information of the original object in the multiple video frames.

5. The method according to claim 3, wherein: Determining the position information of the original object in the plurality of video frames comprises: Acquire position information of the original object in a reference frame, where the reference frame is one of the multiple video frames; and The plurality of video frames are tracked and identified based on the position information of the original object in the reference frame using a pre-trained tracking model to obtain the position information of the original object in the plurality of video frames.

6. The method according to claim 2, wherein: The method further comprises: Generate a depth video based on the original video, and extract features from the depth video to obtain a third video feature, wherein the depth video includes depth images corresponding to each of the multiple video frames, and the depth images represent depth information of each pixel in the video frame; Generating the video feature based on the first video feature and the second video feature includes: The video feature is generated based on the first video feature, the second video feature, and the third video feature.

7. The method according to claim 6, wherein: The generating a depth video based on the original video comprises: Obtaining depth images corresponding to each of the plurality of video frames through a pre-trained gesture recognition model; and The depth video is generated based on the depth images corresponding to each of the multiple video frames.

8. The method according to claim 1, wherein: The replacement requirement information includes at least one of a reference text or a reference image, wherein the reference text at least describes the appearance information of the target object, and the reference image at least shows the appearance of the target object; The conditional feature includes: at least one of a text embedding feature generated based on the reference text, or an image embedding feature generated based on the reference image.

9. The method according to claim 8, wherein: Obtaining replacement requirement information corresponding to the original video, including: A processing instruction input by a user for the original video is obtained, and at least one of the reference text or the reference image is obtained by parsing the processing instruction.

10. The method according to claim 8, wherein: The image embedding feature is generated in the following way: extracting high-frequency features and low-frequency features of the reference image, wherein the high-frequency features are used to characterize detail information of the reference image, and the low-frequency features are used to characterize global information of the reference image; and The high-frequency features and the low-frequency features are fused to generate the image embedding features.

11. A model training method, comprising: Acquire a training sample set, each training sample in the training sample set includes: a first video, a second video, and replacement requirement information, the first video includes a plurality of video frames showing an original object, the replacement requirement information indicates that the original object in the first video is replaced with a target object, and the second video is a video obtained by replacing the original object in the first video with the target object in advance; and The target model is trained multiple times using the training sample set, wherein each round of iterative training includes: generating video features based on the first video and generating conditional features based on the replacement requirement information, inputting the video features and the conditional features into the target model, so as to capture spatiotemporal dependencies from the video features by the target model using a full 3D attention mechanism, and generating a predicted video using the spatiotemporal dependencies and the conditional features as constraints, and The parameters of the target model are adjusted with the goal of minimizing the difference between the predicted video and the second video.

12. The method according to claim 11, wherein: The training sample set includes: A general sample subset, wherein the original objects replaced by each training sample in the general sample subset cover multiple object types, and A dedicated sample subset, in which the original objects replaced by the training samples in each dedicated sample subset are all target types; The step of performing multiple rounds of iterative training on the target model using the training sample set includes: Performing multiple rounds of iterative training on the target model using the common sample subset, so that the target model has the ability to replace multiple types of objects, and The target model is trained for multiple rounds using the dedicated sample subset, so that the target model has a higher replacement capability for the target type of object than for other types of objects.

13. The method according to claim 11, wherein: Adjusting the parameters of the target model with the goal of minimizing the difference between the predicted video and the second video includes: Determine a prediction loss corresponding to the predicted video, and adjust the parameters of the target model with the goal of minimizing the prediction loss, wherein the prediction loss includes a global loss and at least one of a detail loss or a boundary fusion loss, The detail loss is used to characterize the gap between the detail features of the target object in the predicted video and the conditional features. The global loss is used to characterize the difference between the predicted video and the second video, and The boundary fusion loss is used to characterize the degree of fusion distortion between the edge of the target object in the predicted video and the surrounding images.

14. The method according to claim 11, wherein: Generating a video feature based on the first video includes: Extracting features from the first video to obtain first video features; generating a mask video based on the first video, and performing feature extraction on the mask video to obtain second video features, wherein the mask video includes a mask image corresponding to each video frame of the first video, and the mask image represents position information of the original object; and The video feature is generated based on the first video feature and the second video feature.

15. The method according to claim 14, wherein: The method further comprises: Generate a depth video based on the first video, and extract features from the depth video to obtain a third video feature, wherein the depth video includes a depth image corresponding to each video frame in the first video, and the depth image represents the depth information of each pixel in the video frame; Generating the video feature based on the first video feature and the second video feature includes: The video feature is generated based on the first video feature, the second video feature, and the third video feature.

16. The method according to claim 11, wherein: The replacement requirement information includes at least one of a reference text or a reference image, wherein the reference text at least describes the appearance information of the target object, and the reference image at least shows the appearance of the target object; The conditional feature includes: at least one of a text embedding feature generated based on the reference text, or an image embedding feature generated based on the reference image.

17. A video processing system, comprising: at least one storage medium storing at least one instruction set for video processing; as well as At least one processor is communicatively connected to the at least one storage medium, wherein when the video processing system is running, the at least one processor reads the at least one instruction set and implements the method as described in any one of claims 1-10 according to the instructions of the at least one instruction set.

18. A model training system, comprising: At least one storage medium storing at least one instruction set for training a model; as well as At least one processor is communicatively connected to the at least one storage medium, wherein when the training system of the model is running, the at least one processor reads the at least one instruction set and implements the method as described in any one of claims 11-16 according to the instructions of the at least one instruction set.

Citation Information

Cited By

  • Video generation method and device, electronic equipment, medium and product

    CN121218003A