Video generation method and apparatus, electronic device, and computer storage medium
Patent Information
- Application Number
- CN202610838190.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]然而,在虚拟试衣的场景中,目前的视频生成技术往往生成的是处于非交互场景下的虚拟试衣视频,源视频中的模特只通过肢体摆动等简单动作进行服装展示,缺乏与衣服之间的交互动作,从而使得所生成的视频不能反映出试穿的真实效果
[0009] This invention provides a computer program product, including: a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the method described above.
Smart Images

Figure CN122802758A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video technology, and in particular to a video generation method, apparatus, electronic device, and computer storage medium. Background Technology
[0002] Video generation refers to the use of computer vision algorithm models to generate new video sequences. When video generation technology is applied to the application scenario of virtual try-on, it can generate videos of a target person wearing specified clothing in a video sequence, so that consumers can intuitively know the effect of trying on clothing without actually trying it on.
[0003] However, in the context of virtual try-on, current video generation technology often produces virtual try-on videos in non-interactive scenarios. The models in the source videos only display clothing through simple movements such as body swaying, lacking interactive movements with the clothes, thus making the generated videos unable to reflect the real effect of trying on the clothes. Summary of the Invention
[0004] This application provides a video generation method, apparatus, electronic device, and computer storage medium capable of generating videos including interactive actions.
[0005] This invention provides a video generation method, including: Acquire a reference video and a target object image. The reference video includes a person and a preset object. There is an interactive action between a part of the person's body and the preset object. The target object image includes a target object that is different from the preset object. Based on the reference video and the target object image, a context is determined, wherein the context is used to identify the region in the reference video where the preset object is replaced with the target object; Based on the reference video, a 3D partial body rendering image corresponding to the interactive action is determined; Based on the target object image, context, and the 3D body partial rendering map, a target video is generated. The target video includes the person and the target object, and there is an interactive action between the person's body part and the target object.
[0006] This invention provides a video generation apparatus, comprising: The first acquisition module is used to acquire a reference video and a target object image. The reference video includes a person and a preset object. There is an interactive action between a part of the person's body and the preset object. The target object image includes a target object that is different from the preset object. The first determining module is used to determine a context based on the reference video and the target object image, wherein the context is used to identify the region where the preset object is replaced with the target object; The first determining module is further configured to determine a 3D body partial rendering image corresponding to the interactive action based on the reference video; The first processing module is used to generate a target video based on the target object image, the context, and the 3D body partial rendering image. The target video includes the person and the target object, and there is an interactive action between the person's body partial and the target object.
[0007] This invention provides an electronic device, including: a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the above-described method.
[0008] This invention provides a computer storage medium for storing a computer program that enables a computer to implement the above-described method when executed.
[0009] This invention provides a computer program product, including: a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the method described above.
[0010] The video generation method, apparatus, electronic device, and computer storage medium provided in this embodiment acquire a reference video and a target object image, determine the context based on the reference video and the target object image, and determine a 3D body partial rendering map corresponding to the interactive action based on the reference video. Then, a target video is generated based on the target object image, the context, and the 3D body partial rendering map. This effectively enables the generation of realistic target videos based on user-input reference videos and target object images. Furthermore, since the 3D body partial rendering map is introduced as spatial prior information during the target video generation process, accurate gesture shape and depth information can be obtained. This solves the serious clipping problem that exists when interacting between body parts and the target object during video generation, thus ensuring the realism of complex interactive actions in the target video to a certain extent, further improving the quality and effect of target video generation, and ensuring the practicality of the method. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A schematic diagram of a video generation method provided for an exemplary embodiment of this application; Figure 2 A flowchart illustrating a video generation method provided for an exemplary embodiment of this application; Figure 3 A schematic diagram illustrating the process of generating a target video based on the target object image, context, and the 3D body partial rendering image, provided as an exemplary embodiment of this application; Figure 4 A schematic diagram illustrating the process of generating the target video based on the target object image, context, the 3D body partial rendering image, global descriptive text, and the action category label, provided as an exemplary embodiment of this application; Figure 5 A flowchart illustrating the process of using the temporal cross-attention layer to process the first embedded feature and the action category label to obtain the second embedded representation, which is an exemplary embodiment of this application; Figure 6 A schematic diagram of the principle of a video generation method provided in an exemplary application embodiment of this application. Figure 1 ; Figure 7 A schematic diagram of the principle of a video generation method provided in an exemplary application embodiment of this application. Figure 2 ; Figure 8 A schematic diagram of the principle of a video generation method provided in an exemplary application embodiment of this application. Figure 2 ; Figure 9 A schematic diagram of the structure of a video generation apparatus provided for an exemplary embodiment of this application; Figure 10 A schematic diagram of the structure of an electronic device provided for an exemplary embodiment of this application. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.
[0014] Additionally, it should be noted that when user interaction operations or triggering operations are involved in the embodiments of this application, these operations include, but are not limited to, various interaction methods such as touch operations, gesture operations, voice operations, head movement operations, and eye movement operations. Touch operations include, but are not limited to, click operations, double-click operations, long-press operations, swipe operations, pinch operations, or mouse hover operations. Swipe operations include, but are not limited to, straight-line swipes and curved-line swipes.
[0015] To facilitate understanding of the video generation method, apparatus, electronic device, and computer storage medium provided in the embodiments of this application, the relevant technologies are briefly described below: Video virtual try-on refers to putting specified clothing on a target person in a video sequence. It needs to preserve the appearance of the clothing and the person's movements. This not only provides consumers with a special clothing try-on experience, but also allows them to explore clothing options without actually trying them on, which has attracted widespread attention from the fashion industry and consumers.
[0016] Existing video virtual try-on methods can only handle non-interactive scenarios, where models in the generated videos only demonstrate clothing through simple movements like body swaying, lacking any interactive interaction with the garments. In real-world live-streaming e-commerce scenarios, hosts typically interact with clothing (such as unzipping a coat) to showcase its material, elasticity, or details, but current video virtual try-on methods lack the ability to handle such interactive scenarios.
[0017] Related technology 1 proposes a video virtual try-on architecture based on a deep neural network model (e.g., Transformer) with a two-stream self-attention mechanism. It also proposes an occlusion-resistant clothing distortion module and a smoothing module based on ridge regression and optical flow correction. The key ideas of this scheme include: 1) using a distortion module based on Thin Plate Spline (TPS) to predict and mask occlusion areas; 2) correcting optical flow based on ridge regression and optical flow correction methods to obtain more accurate clothing distortion results; and 3) using a two-stream Transformer to fuse the distorted clothing and the person's pose, as well as the clothing-independent background image and the person's pose, respectively.
[0018] However, the above implementation method has the following drawbacks: 1) It can only support virtual try-on in videos with small movements, and cannot cope with the complex relative movement relationship between the camera and the person in real situations, let alone generate interactive actions; 2) It can only support try-on of close-fitting clothes (such as t-shirts, shirts, etc.) with simple textures and patterns, and cannot cope with clothing with complex textures and patterns, and cannot adapt to the diverse types of clothing in real scenes; 3) The model consists of multiple processes connected in series, the implementation plan is complex and cumbersome, and the details of the generated results are not realistic enough.
[0019] Related technology 2 provides a video virtual try-on framework based on the Video Diffusion Transformer (DiT) model, aiming to solve the problems of spatiotemporal consistency and clothing detail preservation in generated videos. Its implementation principle is as follows: 1) The traditional U-Net is replaced with the DiT architecture, and a fully self-attention mechanism is combined to jointly model the spatiotemporal consistency of the video; 2) A coarse-to-fine clothing feature preservation strategy is designed, where the coarse-grained strategy fuses clothing features in the embedding stage, and the fine-grained strategy introduces various clothing-based conditions such as semantics, texture, and contour lines in the denoising stage; 3) A perceptual mask loss is introduced to further optimize the generation fidelity of clothing regions.
[0020] However, the above solution has the following drawbacks: 1) Lack of modeling capability for human-clothing interaction, unable to generate physical contact actions: Modeling only for routine movements of the human body (such as walking, turning, and other non-interactive actions), failing to incorporate prior spatial depth information of the body parts (such as 3D body structure). When faced with real-world interactive scenarios common in e-commerce displays (such as pulling on clothing corners, rolling up sleeves, and zipping up zippers), the model cannot distinguish the occlusion and pulling relationship between body parts and clothing, easily resulting in severe clipping between body parts and clothing, and failing to generate reasonable interactive actions; 2) Lack of action-level temporal semantic control: The introduced semantic and texture conditions are static features at the global level, which makes the model unable to perceive "when the specific interactive action occurs and when it ends", thus failing to accurately drive the generation of interactive effects with physical deformation within a specific time period. 3) Insufficient supervision of complex interaction features: Although perceptual mask loss is introduced to optimize the clothing region, this loss function does not differentiate between temporal dimensions for interactive actions. In real videos, frames containing complex interactive actions are usually very sparse. Using a conventional loss function will cause the features of these key interactive action video frames to be submerged by a massive number of non-interactive frames during training, ultimately leading the model to tend to generate fitting results without variation.
[0021] To address the aforementioned problems, this application provides a video generation method, apparatus, electronic device, and computer storage medium. For details, please refer to the appendix. Figure 1 As shown, the execution entity of this video generation method can be a video generation device 200, which can be implemented as a local server, a cloud server, etc. When the video generation device 200 is implemented as a cloud server, the video generation method can be executed in the cloud. Several computing processes (cloud servers) can be deployed in the cloud, each with processing resources such as computing and storage. In the cloud, multiple computing processes can be organized to provide a certain service; of course, a single computing process can also provide one or more services. The cloud can provide this service by providing a service interface, which users call to use the corresponding service. Service interfaces include Software Development Kits (SDKs), Application Programming Interfaces (APIs), etc.
[0022] The video generation device 200 can communicate with the requesting terminal 100, which is used by users to trigger video generation operations. The requesting terminal 100 can be any computing device with a certain information interaction capability; specifically, it can be a mobile phone, a personal computer (PC), a tablet computer, a setting application, etc. Furthermore, the basic structure of the requesting terminal 100 can include at least one processor. The number of processors depends on the configuration and type of the requesting terminal 100. The requesting terminal 100 may also include memory, which can be volatile, such as Random Access Memory (RAM), or non-volatile, such as Read-Only Memory (ROM), flash memory, or both. The memory typically stores the operating system (OS), one or more applications, and may also store program data. In addition to the processing unit and memory, the requesting terminal 100 also includes some basic configurations, such as a network interface card (NIC) chip, an I / O bus, a display component, and some peripheral devices. Optionally, some peripheral devices may include, for example, a keyboard, mouse, stylus, printer, etc. Other peripheral devices are well known in the art and will not be described in detail here.
[0023] A video generation device 200 refers to a device capable of performing video generation operations in a network virtual environment, typically referring to a device that utilizes a network for information planning and video generation operations. This video generation device 200 can be a network model used to implement video generation operations. In physical implementation, the video generation device 200 can be any device capable of providing computing services and performing corresponding video generation operations, such as a processor, server, etc. The video generation device 200 mainly consists of a processor, hard disk, memory, system bus, etc., and its architecture is similar to that of a general-purpose computer.
[0024] In this embodiment described above, the video generation device 200 and the requesting terminal 100 are connected via a network, which can be a wireless or wired network connection. If the video generation device 200 and the requesting terminal 100 are connected via a communication connection, the mobile network standard can be any one of 2G (Global System for Mobile Communications GSM), 2.5G (General Packet Radio Service GPRS), 3G (Wideband Code Division Multiple Access (WCDMA), Time Division Synchronous Code Division Multiple Access (TD-SCDMA), 4G (Long Term Evolution LTE), 4G+ (Enhanced Long Term Evolution LTE+), Global Microwave Access Interoperability (WiMax), 5G, 6G, etc.
[0025] In this embodiment, the request terminal 100 is used by a user to generate a video generation request that triggers a video generation operation. The video generation request may include a reference video and a target object image. The reference video includes a person and a preset object, with interactive actions between parts of the person's body and the preset object. For example, the reference video could be a video of an interaction between a person and clothing, a person and a home appliance, or a person and a home furnishing product, etc. The target object image includes a target object different from the preset object. For example, the target object could be clothing different from the clothing in the reference video, a home appliance different from the home appliance in the reference video, or a home furnishing product different from the home furnishing product in the reference video. To enable the video generation operation, the video generation request can be sent to the video generation device 200, so that the video generation device 200 can perform the corresponding video generation operation based on the video generation request.
[0026] The video generation device 200 is used to receive a video generation request sent by the requesting end 100. The video generation request includes a reference video and a target object image. After obtaining the reference video and target object image included in the video generation request, the device can determine the context based on the reference video and target object image. This context is used to identify the area where the preset object is replaced with the target object. The device then analyzes and processes the reference video to determine a 3D body partial rendering image corresponding to the interactive action. Subsequently, the device analyzes and processes the target object image, the context, and the 3D body partial rendering image to generate a target video. The generated target video can include a person and a target object, and there is an interactive action between the person's body part and the target object. This effectively enables the generation of a realistic target video based on the reference video and target object image input by the user.
[0027] In this embodiment, a realistic target video can be effectively generated based on the user-input reference video and the target object image. Furthermore, since a 3D body local rendering map is introduced as spatial prior information during the target video generation process, accurate gesture shape and depth information can be obtained. This solves the serious clipping problem that exists when the body local and the target object interact during the video generation process, and can ensure the realism of complex interactive actions to a certain extent, further improving the quality and effect of target video generation.
[0028] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0029] Figure 2A flowchart illustrating a video generation method provided for an exemplary embodiment of this application; see attached diagram. Figure 2 As shown, this embodiment provides a video generation method. The execution subject of this method is a video generation device, which can be implemented as software or a combination of software and hardware. When the video generation device is implemented as hardware, it can specifically be various electronic devices capable of performing video generation operations, including but not limited to servers. When the video generation device is implemented as software, it can be installed in the electronic devices listed above. Specifically, the video generation method may include: Step S201: Acquire a reference video and a target object image. The reference video includes a person and a preset object. There is an interactive action between a part of the person's body and the preset object. The target object image includes a target object that is different from the preset object.
[0030] Step S202: Based on the reference video and the target object image, determine the context, which is used to identify the area where the preset object will be replaced with the target object.
[0031] Step S203: Based on the reference video, determine the 3D body partial rendering image corresponding to the interactive action.
[0032] Step S204: Generate a target video based on the target object image, context, and 3D body partial rendering map. The target video includes a person and the target object, and there are interactive actions between the person's body parts and the target object.
[0033] The specific implementation principles and effects of each of the above steps are explained in detail below: Step S201: Acquire a reference video and a target object image. The reference video includes a person and a preset object. There is an interactive action between a part of the person's body and the preset object. The target object image includes a target object that is different from the preset object.
[0034] When a user has a video generation requirement, the video generation device can acquire a reference video and an image of the target object. The reference video can be a video including a person and a preset object, where there are interactive actions between parts of the person's body and the preset object. For example, if the preset object is clothing and the body part is a hand, the reference video could include actions such as pulling on the hem of the clothing, rolling up the sleeves, or zipping up a zipper. If the preset object is a household appliance and the body part is a hand, the reference video could include actions such as starting, operating, or turning off the appliance. If the preset object is a home furnishing product and the body part is a hand, the reference video could include actions such as contact or adjustment between the person and the product. Similarly, if the preset object is shoes and the body part is a foot, the reference video could include actions such as putting on shoes, tying shoelaces, taking off shoes, or changing how shoes are worn. If the preset object is a soccer ball and the body part is a foot, the reference video could include actions such as juggling, pulling, or controlling the ball.
[0035] In addition, the target object image may include a target object that is different from the preset object. For example, when the preset object is clothing a, the target object is clothing b; when the preset object is home appliance a, the target object can be home appliance b; when the preset object is home furnishing product a, the target object can be home furnishing product b.
[0036] In some instances, the reference video and the target object image can be determined by a requesting end, which includes the reference video and the target object image and is communicatively connected to the video generating device. In this case, acquiring the reference video and the target object image may include: determining the requesting end that is communicatively connected to the video generating device, which may include the reference video and the target object image; and acquiring the reference video and the target object image actively or passively through the requesting end, thereby reliably acquiring the reference video and the target object image used to trigger the video generation operation.
[0037] Step S202: Based on the reference video and the target object image, determine the context, which is used to identify the region in the reference video where the preset object is replaced with the target object.
[0038] In order to generate a target video that includes the target object and the interactive actions in the reference video, after acquiring the reference video and the target object image, the reference video and the target object image can be analyzed and processed to determine the context. This context is used to identify the area in the reference video where the preset object is replaced with the target object. For example, if the reference video is a video of a person interacting with clothing 'a', and the target object image can be an image of clothing 'b', then by analyzing and processing the reference video and the target object image, the context in the reference video used to identify the area where clothing 'a' is replaced with clothing 'b' can be determined.
[0039] In some instances, the context can be determined using a pre-trained context recognition model. In this case, determining the context based on the reference video and the target object image can include: determining the pre-trained context recognition model; inputting the reference video and the target object image into the context recognition model for analysis and processing to obtain the context output by the context recognition model. This effectively ensures the accuracy and reliability of the context determination.
[0040] In this application, the context recognition model can be a Large Language Model (LLM) based on artificial intelligence. This application does not limit the number of model parameters supported by the model, aiming to meet actual needs. If the model has relatively more parameters, the model will be larger and perform better, but it will consume more time and resources during inference or training. If the model has relatively fewer parameters, the model will be smaller and, while meeting performance requirements, more lightweight, consuming less time and resources during inference or training. This context recognition model can be a deep learning model used to process and generate natural language text or multimodal data, implemented based on a neural network architecture, and can be pre-trained on large amounts of data. In an optional implementation, the context recognition model may include an encoder, a decoder, a self-attention layer, and a feed-forward neural network, etc. The encoder is mainly used to convert input data (usually in sequence form) into vector representation. This process can capture the semantic features of the input data. The decoder is responsible for converting the intermediate representation generated by the encoder into output data (usually in sequence form). The self-attention layer is a mechanism that allows the model to pay attention to other positions in the sequence to better encode the current position information. The feedforward neural network can perform nonlinear transformations on the output of the self-attention layer to enhance the model's expressive power. All parts work together, enabling the model built on them to perform well in various complex processing tasks, such as natural language processing, computer vision, speech recognition, machine translation, text summarization, and intelligent question answering.
[0041] In some instances, the context can be determined not only through a pre-trained context recognition model, but also by combining the skeletal pose map of a person. In this case, determining the context based on the reference video and the target object image can include: determining the skeletal pose map of the person based on the reference video; determining the main body region of the object and the mask corresponding to the skeletal pose map based on the target object image, the mask being used to preserve the skeletal region in the skeletal pose map that fits the target object; and determining the context based on the skeletal pose map, the main body region of the object, and the mask.
[0042] To ensure that the generated target video accurately reflects the same interactive actions as the reference video, the reference video is analyzed after acquisition to determine the skeletal pose map of the character. This skeletal pose map includes the main joints and feature points of the human body, such as the head, shoulders, elbows, and wrists, reflecting the character's motion information and the interaction between the character and the preset object in the reference video. Furthermore, since the reference video contains multiple frames, the obtained skeletal pose map can be a sequence of skeletal pose maps, comprising multiple skeletal pose maps. In some instances, the skeletal pose map of the character can be determined by analyzing the reference video using a pre-trained pose extraction model.
[0043] After determining the skeletal pose of the person, the target image and the skeletal pose of the person can be analyzed to determine the main body region of the object and the corresponding mask in the skeletal pose. This mask is used to preserve the skeletal region in the skeletal pose that matches the target object. For example, when the target image is an image of pants b, after analyzing the target image and the skeletal pose, the main body region of the object can be determined to be the lower body region. A mask corresponding to the skeletal pose can also be determined. This mask includes mask values corresponding to each pixel. A mask value of "0" or "1" indicates that the area to be preserved should be a "1", while a mask value of "0" indicates that the area not to be preserved should be a "0". This mask is used to preserve the lower body skeletal region in the skeletal pose that matches pants b. When the target object image is the image of the outer garment c, after analyzing and processing the target object image and the skeletal pose map, it can be determined that the main body area of the object is the upper body area. At the same time, the mask information corresponding to the skeletal pose map is also determined. This mask is used to preserve the upper body skeletal area in the skeletal pose map that is compatible with the outer garment c.
[0044] After determining the main object region and its mask, the skeletal pose map, main object region, and mask can be analyzed to determine the context. In some instances, the context is determined by analyzing the skeletal pose map, main object region, and mask using a pre-trained context recognition model. Alternatively, the context can be determined by performing a weighted summation operation on the skeletal pose map and the main object region. In this case, determining the context based on the skeletal pose map, main object region, and mask can include: determining the subject mask corresponding to the main object region based on the mask; and performing a weighted summation operation on the skeletal pose map and the main object region based on the mask and the subject mask to determine the context.
[0045] Specifically, since the mask corresponds to the entire skeletal pose map, after determining the mask, it can be analyzed to determine the subject mask corresponding to the main body region of the object. In some instances, the subject mask can be determined by the difference between 1 and the mask, i.e., subject mask = 1 - mask. Then, a weighted summation operation can be performed on the skeletal pose map and the main body region of the object based on the mask and the subject mask, i.e., context = mask * skeletal pose map + (1 - mask) * main body region of the object, thus stably determining the context.
[0046] Step S203: Based on the reference video, determine the 3D body partial rendering image corresponding to the interactive action.
[0047] The process of determining the 3D local body rendering map corresponding to the interactive action based on the reference video may include: calling a pre-trained local body estimation model; inputting the reference video into the local body estimation model to obtain the 3D local body rendering map output by the model. It should be noted that the reference video includes multiple video frames. When the local body estimation model performs local body estimation on the reference video, it can analyze and process each video frame and determine the 3D local body rendering map corresponding to the interactive action in each video frame; that is, the number of determined 3D local body rendering maps is multiple.
[0048] Step S204: Generate a target video based on the target object image, context, and 3D body partial rendering map. The target video includes a person and the target object, and there are interactive actions between the person's body parts and the target object.
[0049] After obtaining the target object image, context, and 3D body partial rendering, these elements can be analyzed and processed to generate the target video. In some instances, generating the target video based on the target object image, context, and 3D body partial rendering may include: calling a diffusion model to implement the video generation operation; inputting the target object image, context, and 3D body partial rendering into the diffusion model to stably generate the target video, wherein the generated target video includes a person and the target object, and the person's body parts and the target object have the same interactive actions included in the reference video.
[0050] The video generation method provided in this embodiment acquires a reference video and a target object image, determines the context based on the reference video and the target object image, and determines a 3D body partial rendering map corresponding to the interactive action based on the reference video. Then, it generates a target video based on the target object image, the context, and the 3D body partial rendering map. This effectively enables the generation of realistic target videos based on user-input reference videos and target object images. Furthermore, since the 3D body partial rendering map is introduced as spatial prior information during the target video generation process, accurate gesture shape and depth information can be obtained. This solves the serious clipping problem that exists when interacting between body parts and the target object during video generation. This ensures the realism of complex interactive actions in the target video to a certain extent, further improving the quality and effect of target video generation and guaranteeing the practicality of the method.
[0051] Figure 3 This is a schematic diagram illustrating a process for generating a target video based on a target object image, context, and a 3D body partial rendering map, provided as an exemplary embodiment of this application. Based on the above embodiment, refer to the appendix... Figure 3 As shown, for a target video, it can be determined not only by analyzing and processing the target object image, context, and 3D body local rendering map using a diffusion model, but also by combining action category labels corresponding to interactive actions to generate the target video. In this case, generating a target video based on the target object image, context, and 3D body local rendering map can include: Step S301: Based on the reference video, generate action category labels corresponding to the interactive actions. The action category labels have video timestamps.
[0052] When generating a target video based on a reference video and an image of the target object, to ensure that the interactive actions in the reference video are precisely aligned with those in the target video in terms of time tactile analysis, and to ensure that interactive actions are generated only within the correct time periods, avoiding interference with non-interactive action time periods, the reference video can be analyzed and processed after acquisition to generate action category labels corresponding to the interactive actions. These action category labels can correspond to video timestamps. For example, if the reference video contains 80 frames, the action category label generated through analysis and processing could be "rolling up sleeves," and this action category label could correspond to video timestamps "0 to 32 frames"; or, the action category label could be "pulling the hem of a garment," and this action category label could correspond to video timestamps "10-35 frames," and so on.
[0053] In some instances, action category labels can be determined by analyzing and processing reference videos using a pre-trained label recognition model. In this case, generating action category labels corresponding to interactive actions based on the reference video can include: inputting the reference video into the pre-trained label recognition model for label recognition, and determining the action category labels corresponding to the interactive actions output by the label recognition model. This effectively ensures the accuracy and reliability of determining action category labels.
[0054] In other instances, action category labels can be determined by filtering within a preset label set. In this case, generating action category labels corresponding to interactive actions based on a reference video can include: determining a preset label set, which includes multiple standard action categories; determining the matching degree between the interactive actions in the reference video and each standard action category based on the preset label set and the reference video; and determining the standard action category corresponding to the matching degree as the action category label corresponding to the interactive actions in the reference video if the matching degree is greater than or equal to a preset threshold. This also ensures the flexibility and reliability of determining action category labels.
[0055] Step S302: Generate the target video based on the target object image, context, 3D body partial rendering map, and action category label.
[0056] In some instances, generating a target video based on the target object image, context, 3D body partial rendering map, and action category labels may include: calling a diffusion model to implement the video generation operation; inputting the target object image, context, 3D body partial rendering map, and action category labels into the diffusion model, thereby stably generating the target video, thus ensuring the accuracy and reliability of the target video generation.
[0057] In other instances, the target video can be determined not only by analyzing the target object image, context, 3D body partial rendering, and action category labels, but also by combining global descriptive text corresponding to the reference video. In this case, generating the target video based on the target object image, context, 3D body partial rendering, and action category labels can include: generating global descriptive text based on the reference video; and generating the target video based on the target object image, context, 3D body partial rendering, global descriptive text, and action category labels.
[0058] In order to further ensure the accuracy of interactive actions in the target video after obtaining the reference video, the reference video can be analyzed and processed to generate global descriptive text. This global descriptive text is used to describe the interactive actions in the reference video from a global perspective. For example, the global descriptive text can be "The model adjusted her coat, buttoned it up, and then turned around," etc. In addition, the method for determining this global descriptive text is similar to the method for determining the "action category label" mentioned above. For details, please refer to the above description, which will not be repeated here.
[0059] After generating the global descriptive text, the target object image, context, 3D body partial rendering image, global descriptive text, and action category labels can be analyzed and processed to generate the target video. In some instances, generating the target video based on the target object image, context, 3D body partial rendering image, global descriptive text, and action category labels may include: calling a diffusion model to implement the video generation operation; inputting the target object image, context, 3D body partial rendering image, global descriptive text, and action category labels into the diffusion model, thereby stably generating the target video, thus ensuring the accuracy and reliability of the target video generation.
[0060] In this embodiment, action category labels corresponding to interactive actions are generated based on reference videos. Then, a target video is generated based on the target object image, context, 3D body partial rendering map, and action category labels. Since the target video generation process introduces the 3D body partial rendering map as a spatial prior and fuses it through the action category labels of the specific interactive actions, the video generation process can obtain accurate gesture shape and depth information, completely solving the problem of severe clipping between the body parts and the target object that is easy to occur in the video generation process. This successfully ensures the realistic generation of complex interactive actions and further guarantees the quality and effect of the target video generation.
[0061] Figure 4 This is a schematic diagram illustrating a process for generating a target video based on a target object image, context, 3D body partial rendering image, global descriptive text, and action category tags, provided as an exemplary embodiment of this application; based on the above embodiment, refer to the appendix... Figure 4 As shown, for the target video, it can be determined not only by analyzing and processing the input target object image, context, 3D body local rendering map, global descriptive text, and action category label using a diffusion model, but also by using a diffusion model that includes a context module and a signal injection module. Therefore, in this embodiment, generating the target video based on the target object image, context, 3D body local rendering map, global descriptive text, and action category label can include: Step S401: Call the pre-trained diffusion model for implementing video generation operations. The diffusion model includes at least: a context module and a signal injection module.
[0062] The diffusion model can be a pre-trained model for video generation. To generate the target video based on the model, the diffusion model is trained before calling the pre-trained model for video generation. In this embodiment, the method may further include: acquiring video generation training data, which includes initialized noise latent variables, a preset step size for noise filtering, input data for video generation, and preset standard noise corresponding to a preset standard video. The input data includes a training video and a training object image; inputting the noise latent variables, preset step size, and input data into the diffusion model to be trained to generate predicted noise; generating a temporal mask corresponding to the interactive actions in the training video; generating an interaction loss function corresponding to the interactive actions based on the temporal mask, predicted noise, and preset standard noise; and optimizing the diffusion model to be trained using the interaction loss function to generate the diffusion model.
[0063] Specifically, in order to train a diffusion model for video generation, video generation training data can be obtained first. This training data can be acquired by accessing a preset database, collecting data from a preset server, or through manual configuration. The noise latent variables, preset step size, and input data included in the video generation training data are then input into the diffusion model to be trained to generate predicted noise. ,in, Used to identify the diffusion model to be trained. For noise, latent variables This refers to the preset step size used for filtering noise, or the interval step size used to filter noise information. For input data, To predict noise.
[0064] After obtaining the training video, it can be analyzed and processed to generate a temporal mask corresponding to the interactive actions in the training video. The timing mask Interactive features can be identified by training on interactive and non-interactive video frames in the video, where the mask value corresponding to an interactive video frame differs from that of a non-interactive video frame. In some instances, the mask value for interactive video frames can be set to "1", while the mask value for non-interactive video frames can be set to "0".
[0065] After obtaining temporal masks for interactive actions in both the generated and training videos, the temporal masks, prediction noise, and preset standard noise can be analyzed and processed to generate an interaction loss function corresponding to the interactive actions. In some instances, the interaction loss function can be implemented as follows: ,in, Used to identify desired information, To predict noise, To preset standard noise, This is a temporal mask corresponding to the interactive actions in the training video. This represents element-wise multiplication (Hadamard product). After obtaining the interaction loss function, it can be used to optimize the diffusion model to be trained, thereby generating a diffusion model for implementing video generation operations.
[0066] In other instances, not only can the diffusion model be optimized based on the direct interaction loss function to generate a diffusion model, but it can also be optimized by combining the video generation loss function corresponding to the video generation operation. In this case, optimizing the diffusion model using the interaction loss function to generate the diffusion model can include: generating a generation loss function corresponding to the video generation operation based on predicted noise and preset standard noise; generating a total loss function based on the interaction loss function and the generation loss function; and optimizing the diffusion model based on the total loss function to generate the diffusion model.
[0067] Specifically, after obtaining the predicted noise and the preset standard noise, the predicted noise and the preset standard noise can be analyzed and processed to determine the generation loss function corresponding to the video generation operation. In some instances, this generation loss function can be implemented as follows: ,in, Used to identify desired information, To predict noise, The standard noise is preset. After determining the generation loss function, the interaction loss function and the generation loss function can be analyzed and processed to generate the total loss function. In some instances, the total loss function can be implemented as the sum of the interaction loss function and the generation loss function.
[0068] Alternatively, the total loss function can be determined based on the preset weights corresponding to the interaction loss function. In this case, generating the total loss function based on the interaction loss function and the generation loss function may include: determining the preset weights corresponding to the interaction loss function; determining the product of the preset weights and the interaction loss function; and determining the sum of the product and the generation loss function as the total loss function. Specifically, this total loss function can be determined using the following formula: = ,in, For the total loss function, To generate the loss function, To preset weights, The interaction loss function effectively ensures the accuracy and reliability of determining the total loss function.
[0069] After determining the total loss function, it can be used to optimize the diffusion model to be trained, thus accurately completing the training operation of the diffusion model and determining the diffusion model used to implement the video generation operation. Furthermore, because an additional interaction loss function is used to assign higher penalty weights to interactive action frames during the training process, the model is forced to prioritize and learn the generation operations of interactive actions and complex interactive actions during gradient backpropagation, greatly improving the quality and efficiency of the model's learning of interactive action generation.
[0070] After training and obtaining the diffusion model, the diffusion model can be stored in a preset area so that users can call the diffusion model as needed to perform video generation operations based on the called diffusion model.
[0071] Step S402: Use the signal injection module to process the 3D body local rendering image, global description text and action category label to obtain 3D body local interaction signals.
[0072] Since the diffusion model includes a context module and a signal injection module, the context module is used to fuse and encode multimodal information to construct a latent space representation for action perception; the signal injection module is used for fine-grained semantic parsing and mapping of actions. Therefore, after obtaining the 3D body local rendering image, global descriptive text, and action category labels, these can be input into the signal injection module for analysis and processing to obtain 3D body local interaction signals. These 3D body local interaction signals are used to characterize the changes in local body movements in the reference video.
[0073] In some instances, the signal injection module in the diffusion model may include multiple normalization layers, multiple full attention layers, and a feedforward neural network layer. In this case, the multiple normalization layers and full attention layers included in the signal injection module can be used to process the 3D body local rendering map, global descriptive text, and action category labels in sequence, and the 3D body local interaction signal output by the feedforward neural network layer can be obtained.
[0074] In other instances, the signal injection module may include a global cross-attention layer, a temporal cross-attention layer, and a feedforward neural network. The number of global cross-attention layers and temporal cross-attention layers can be one or more, and the temporal logical relationship between them is not limited; that is, the global cross-attention layer can be located before or after the temporal cross-attention layer. The following example illustrates this with the global cross-attention layer preceding the temporal cross-attention layer. In this case, processing the 3D body local rendering image, global descriptive text, and action category labels using the signal injection module to obtain 3D body local interaction signals may include: processing the 3D body local rendering image and global descriptive text using the global cross-attention layer to obtain a first embedded representation; processing the first embedded feature and action category labels using the temporal cross-attention layer to obtain a second embedded representation; and processing the second embedded representation using the feedforward neural network to obtain the 3D body local interaction signals.
[0075] Specifically, after obtaining the 3D body local rendering image, global descriptive text, and dynamic category labels, the body local rendering image can be input into the signal injection module, and the global descriptive text can be input into the global cross-attention layer. The global cross-attention layer processes the 3D body local rendering image and global descriptive text to obtain the first embedding representation. Then, the first embedding feature and action category label can be processed using the temporal cross-attention layer to obtain the second embedding representation. The second embedding representation can then be processed using a feedforward neural network to stably generate 3D body local interaction signals. These 3D body local interaction signals are used to represent the interactive action information of the body locality in the reference video, thereby improving the accuracy and reliability of interactive action generation in the target video.
[0076] Furthermore, to ensure the stability and reliability of the target video generated by the diffusion model, the diffusion model may also include an image encoding module. Before processing the 3D body local rendering image, global descriptive text, and action category labels using the signal injection module, the method in this embodiment may further include: using the image encoding module to encode the target object image, context, and 3D body local rendering image to obtain object encoding features, context encoding features, and 3D body local encoding features corresponding to the target object image, so that the context module and signal injection module can use the object encoding features, 3D body local encoding features, and context encoding features to achieve stable processing operations on the target object image, context, and 3D body local rendering image.
[0077] Step S403: Use the context module to process the 3D body local interaction signals, context and target object image to obtain context signals. The context signals are used to reflect the background, interactive actions and target objects in the reference video.
[0078] After obtaining the 3D body local interaction signal, the context module can be used to process the 3D body local interaction signal, the context, and the target object image to stably acquire the context signal. This context signal reflects the background, interactive actions, and target actions in the reference video. In some examples, processing the 3D body local interaction signal, the context, and the target object image using the context module to obtain the context signal may include: encoding the target object image and the context separately using the image encoding module to obtain the clothing embedding representation and the context embedding representation; then stitching the clothing embedding representation, the context embedding representation, and the 3D body local interaction signal to obtain the stitched representation; and finally, analyzing and processing the stitched representation using the context module to stably acquire the context signal used to generate the target video.
[0079] Step S404: Generate the target video based on the context signal.
[0080] After obtaining the context signal, it can be analyzed and processed to stably generate the target video. In some instances, the target video can be determined by decoding the context signal using an image decoder. In this case, generating the target video based on the context signal can include: inputting the context signal into the image decoder in the diffusion model for decoding to obtain the target video.
[0081] In some other instances, the diffusion model may also include a denoising module and an image decoder; in this case, generating the target video based on the context signal may also include: using the denoising module to process the initialized noisy latent variables and the context signal to obtain the target noise for generating the target video; and using the image decoder to decode the target noise to obtain the target video.
[0082] Specifically, when the diffusion model includes a denoising module and an image decoder, the denoising module can be used to process the initialized noisy latent variables and context signals to obtain the target noise for generating the target video. The initialized noisy latent variables can be determined based on the reference video. Then, the image decoder can be used to decode the target noise, thereby stably obtaining the target video. This effectively ensures the stability and reliability of the target video generation.
[0083] Furthermore, before processing the initialized noisy latent variables and context signals using the denoising module, the method in this embodiment can identify whether the initialized noisy latent variables and context signals are data aligned. If the initialized noisy latent variables and context signals are not data aligned, the noisy latent variables can be padded with placeholders based on the context signals to obtain adjusted latent variables. Then, the denoising module can be used to process the adjusted latent variables and context signals with relatively aligned data dimensions, thus ensuring the reliability of the target noise generation.
[0084] In this embodiment, a pre-trained diffusion model for video generation is invoked. Then, a signal injection module processes the 3D body local rendering image, global descriptive text, and action category labels to obtain 3D body local interaction signals. Simultaneously, a context module processes the 3D body local interaction signals, context, and target object image to obtain context signals. The target video can be generated based on the obtained context signals. Since the generated target video is generated by the fusion signal between the 3D body local rendering image, global descriptive text, and action category labels, the realism of the interactive actions in the target video is guaranteed to a certain extent.
[0085] Figure 5 This is a flowchart illustrating an exemplary embodiment of the present application, showing how a temporal cross-attention layer processes a first embedded feature and an action category label to obtain a second embedded representation; based on the above embodiment, refer to the appendix... Figure 5 As shown, the second embedding representation can be determined not only by directly analyzing and processing the first embedding features and action category labels based on the temporal cross-attention layer, but also by encoding interactive and non-interactive segments using a rotational position encoding mechanism. In this case, processing the first embedding features and action category labels using the temporal cross-attention layer to obtain the second embedding representation can include: Step S501: Based on the action category labels, determine the interactive and non-interactive segments in the reference video.
[0086] To accurately obtain the second embedded tag, after obtaining the action category tag corresponding to the interactive action, the action category tag can be analyzed to determine the interactive and non-interactive segments in the reference video. Specifically, determining the interactive and non-interactive segments in the reference video based on the action category tag can include: determining the interactive segments in the reference video based on the timestamp corresponding to the action category tag; and determining the non-interactive segments based on the reference video and the interactive segments, i.e., the non-interactive segments are the other video frames in the reference video excluding the interactive segments.
[0087] Step S502: Determine the text embedding representation corresponding to the action category label.
[0088] In order to accurately determine the second embedded representation corresponding to the interactive action, after obtaining the action category label, the action category label can be analyzed to determine the text embedded representation corresponding to the action category label. In some instances, the text embedded features can be determined by analyzing the action category label using a text extractor.
[0089] Step S503: Encode interactive segments and non-interactive segments using a rotation position encoding mechanism to obtain interactive segment representations and non-interactive segment representations.
[0090] Since interactive segments correspond to interactive actions in the reference video, and non-interactive segments correspond to non-interactive actions in the reference video, in order to understand and represent the interactive actions in the reference video, a rotational position encoding mechanism can be used to encode the interactive segments and non-interactive segments to obtain the representations of the interactive segments and non-interactive segments.
[0091] In some instances, using a rotational position encoding mechanism to encode interactive and non-interactive segments to obtain representations of these segments can include: determining a first rotation angle corresponding to the interactive segment and a second rotation angle corresponding to the non-interactive segment based on the rotational position encoding mechanism, wherein the first and second rotation angles are different and can be pre-configured angles for implementing mapping operations; processing based on the first rotation angle and the interactive segment, specifically controlling the encoding of the interactive segment based on the first rotation angle to determine the interactive segment representation; similarly, processing based on the second rotation angle and the non-interactive segment, specifically controlling the encoding of the non-interactive segment based on the second rotation angle, thereby stably determining the non-interactive segment representation.
[0092] In other instances, using a rotational position encoding mechanism to encode interactive and non-interactive segments to obtain interactive and non-interactive segment representations may include: determining a first rotational position code corresponding to the interactive segment and a second rotational position code corresponding to the non-interactive segment using the rotational position encoding mechanism. The first rotational position code may be determined based on a first rotation angle specified by the rotational position encoding mechanism, and the second rotational position code may be determined based on a second rotation angle specified by the rotational position encoding mechanism. The first and second rotation angles are different. After determining the first rotational position code, it can be applied to the interactive segment to obtain its representation; similarly, after determining the second rotational position code, it can be applied to the non-interactive segment to obtain its representation. This effectively ensures the accuracy and reliability of determining the interactive and non-interactive segment representations.
[0093] Step S504: Process the text embedding representation, interactive segment representation and non-interactive segment representation using a temporal cross-attention layer to obtain the second embedding feature.
[0094] After obtaining the text embedding representation, interactive segment representation, and non-interactive segment representation, a temporal cross-attention layer can be used to analyze and process the text embedding representation, interactive segment representation, and non-interactive segment representation in order to stably obtain the second embedding feature.
[0095] In this embodiment, interactive and non-interactive segments in the reference video are determined based on action category labels, and text embedding representations corresponding to the action category labels are determined. Then, a rotation position encoding mechanism is used to encode the interactive and non-interactive segments to obtain interactive segment representations and non-interactive segment representations. A temporal cross-attention layer is used to process the text embedding representations, interactive segment representations, and non-interactive segment representations to obtain a second embedding feature. This effectively achieves precise alignment of the action category labels corresponding to interactive actions with video frames in the time dimension, thereby improving the accuracy of determining 3D body local interaction signals based on the second embedding feature.
[0096] In practical applications, taking clothing in a virtual try-on scenario as the preset object and the hand as a part of the body as an example, this application embodiment provides a method for generating virtual try-on videos based on spatial and semantic guidance. This method can break through the limitation of existing video virtual technology that can only handle simple swaying (non-interactive) of the person. The generated virtual try-on video can include complex interactive actions, such as supporting the AI model's natural display of complex actions such as pulling the corner of the clothes, rolling up the sleeves, and zipping up the zipper. This can not only greatly improve the user's shopping experience and purchase conversion rate, but also reduce the merchant's real-life shooting production cost and product promotion cost.
[0097] This generation method employs the advanced Video DiffusionTransformer (DiT) as its base model, modeling the interactive video virtual try-on task as a conditional generation task guided by both clothing images and action semantics. The method designs a multi-layered interactive signal injection mechanism, introducing 3D hand priors in the spatial dimension to provide fine spatial guidance for the model's generation of hand interaction actions, effectively avoiding the problem of missing depth in 2D information. Semantically, it utilizes timestamped interactive action category labels to accurately define the category and duration of interactive actions, and achieves temporal alignment between interactive action category labels and video features through innovative motion-aware rotational position encoding. Furthermore, addressing the issue of sparse interactive signal distribution, a motion-aware constraint loss is proposed to specifically enhance the supervision of interactive action video frames during training, enabling the model to efficiently learn the generation of complex interactive actions.
[0098] Through the above technical design, in interactive video virtual try-on application scenarios, given an input video and a clothing image, the interactive video virtual try-on task can be modeled as a conditional generation task guided by clothing image features and action semantics, generating virtual try-on videos with interactive actions that conform to real physical laws and high spatiotemporal consistency. For details, please refer to the appendix. Figure 6 As shown, the method for generating virtual fitting videos based on spatial and semantic guidance in this embodiment may include the following steps: Step 1: Obtain the source video and clothing images used as input to the video generation model, and extract and preprocess the multimodal guidance information of the images.
[0099] The process involves acquiring source video and clothing images as input to the video generation model. The source video includes interactive actions between a person and reference clothing, ensuring that the person in the generated virtual fitting video has the same interactive actions as the virtual clothing in the clothing images. After acquiring the source video, information extraction operations can be performed on the source video and clothing images. Typically, the source video length can be 5 seconds, 3 seconds, or 10 seconds, and it can include multiple video frames. When the source video is long, it can be divided into multiple short videos of preset lengths, improving the quality and efficiency of source video analysis and processing. In some scenarios, there are no interactive actions between the person and reference clothing in the source video; in this case, there will also be no interactive actions between the person and clothing images in the generated virtual fitting video.
[0100] Specifically, the basic visual input of the video generation model may include: a clothing image, a video region (determined by a protection region and a protection region mask) used to remove reference clothing from the source video and retain the background, and a 2D skeletal pose map of the human body. The mask can be determined by the clothing image; the 2D skeletal pose map of the human body can be determined by parsing the source video; after determining the 2D skeletal pose map of the human body and the protection region, the context for generating the virtual try-on video can be determined by the mask, the 2D skeletal pose map of the human body, and the protection region. This context can be determined by mask * human skeletal pose map + (1 - mask) * protection region. This context is used to retain the human body area in the source video where the clothing needs to be changed.
[0101] For video generation models, in addition to the clothing images and context mentioned above, spatial guidance features (3D hand priors) can also be included. Specifically, for the key hand movements in the interactive actions of the source video, a pre-trained 3D hand estimation model can be used to extract the 3D hand rendering map of each frame in the input source video. By using the provided 3D hand rendering map as input data for the video generation model, accurate spatial shape and depth information of the hand can be provided to the video generation model, thereby effectively avoiding the depth loss and "hand-clothing clipping" problem caused by using only 2D skeletons.
[0102] In addition to clothing images, context, and 3D hand renderings, the input data for the video generation model can also include semantic guidance features. These semantic guidance features can include two types of textual guidance information: one is “global descriptive text” that describes the overall action of the video (such as “the model straightened her coat, buttoned it up, and then turned around”); the other is “interactive action category labels with timestamps” (such as explicitly indicating that the “roll up sleeves” action occurred in frames 0 to 32).
[0103] Step 2: Use a video generation model to process clothing images, context, 3D hand renderings, and semantic guidance features to generate virtual try-on videos. These videos offer high-fidelity details and natural interactive virtual try-on results.
[0104] The video generation model's backbone employs a parallel multi-branch module design, specifically including: (1) After the visual input is mapped to the latent space by the image encoder, it is divided into multiple token sequences.
[0105] Specifically, the image encoder can encode clothing images, context, and 3D hand renderings to obtain clothing tokens corresponding to the clothing images, context tokens corresponding to the context, and 3D hand tokens corresponding to the 3D hand renderings.
[0106] (2) Context module: responsible for processing clothing tokens and context tokens (composed of background protection area, mask and human 2D pose) to inject background and basic motion information into the generation process.
[0107] Specifically, after obtaining the clothing tokens and context tokens, the clothing tokens and context tokens can be concatenated to generate the first concatenated tokens. Then, based on the first concatenated tokens, a placeholder operation is performed on the 3D hand tokens and pre-initialized denoising tokens (achieved by filling placeholder tokens) to obtain the 3D hand concatenated tokens and denoising concatenated tokens corresponding to the first concatenated tokens. Then, the first concatenated tokens can be analyzed and processed using the context model to inject background and basic motion information into the video generation process.
[0108] (3) Denoising module (DiT module): responsible for receiving noisy tokens and gradually predicting and removing noise in multi-step iterations.
[0109] The interactive motion signal injection module can analyze and process 3D hand stitching tokens to obtain hand injection signals. Then, the context module analyzes and processes the result of adding the hand injection signal and the first stitching tokens to obtain context signals. The context signals are then transmitted to the denoising module so that the denoising module can perform progressive prediction and noise removal processing on the context signals and denoised stitching tokens to generate target noise for generating the target video.
[0110] In some instances, when the interactive action signal injection module analyzes and processes 3D hand-stitched tokens, preliminary spatial and semantic fusion can be achieved. For details, please refer to the appendix. Figure 7As shown, the interactive action signal injection module includes a normalization layer, a full attention layer, a global cross-attention layer, and a temporal cross-attention layer. After the 3D hand rendering image is encoded to form "3D hand tokens," these tokens are sent to the interactive action signal injection module and sequentially processed through the normalization layer, the full attention layer, the global cross-attention layer (integrating global descriptive text semantics), and the temporal cross-attention layer. During the global cross-attention layer analysis, semantic guidance information corresponding to the global descriptive text can be input. This semantic guidance information can be determined by encoding the global descriptive text using a text encoder. During the analysis and processing by the temporal cross-attention layer, semantic guidance information corresponding to the interactive action category label can be input. This semantic guidance information can be determined by encoding the interactive action category label through a text encoder, thereby stably generating a hand injection signal. The logical order between the full attention layer, the global cross-attention layer, and the temporal cross-attention layer can be flexibly adjusted, as long as it can be ensured that the generated hand injection signal is determined by deep fusion analysis of fine gestures in 3D space and text semantic features. Then, this hand injection signal can be added as a control signal bypass to the denoising process of the denoising module.
[0111] Furthermore, when analyzing and processing the semantic guidance information corresponding to "3D hand tokens" and interactive action category labels in the temporal cross-attention layer, a motion-aware rotational position encoding mechanism is introduced to ensure accurate alignment of temporal semantics. In interactive source videos, interactive actions often only occur within a specific time period. To prevent interactive semantics from interfering with non-interactive frames, a motion-aware rotational position encoding mechanism can be configured in the cross-attention layer and made into a temporal cross-attention module. The "motion-aware rotational position encoding mechanism" is used to divide the video frames in the source video into interactive video frames and non-interactive video frames, and then apply different rotational position encodings to interactive video frames and non-interactive video frames. For example, a deflection angle 'a' is applied to interactive video frames, and a deflection angle 'b' is applied to non-interactive segments. The aforementioned deflection angles 'a' and 'b' are different, that is, there is a deviation between deflection angles 'a' and 'b', so that interactive video frames and non-interactive video frames can generate significantly different responses during the cross-attention process.
[0112] For details, please refer to the appendix. Figure 8As shown, after obtaining the interactive action category label, the interactive action category label can be encoded to obtain text embedding. At the same time, the interactive action timestamp corresponding to the source video can be determined based on the interactive action category label. For example, the interactive action timestamp is [0,32], that is, among the multiple video frames included in the source video, frames 0 to 32 are interactive segments, and frames 33 and other video frames are non-interactive segments.
[0113] When calculating the Query(Q) matrix and Key(K) matrix of temporal cross-attention, the relationship between the current video frame and the timestamps of the aforementioned interactive actions can be identified, so as to determine whether the current video frame belongs to an interactive segment based on its time index. Then, for video frames located within the interactive timestamp interval, a rotation position code with a specific rotation angle bound to the interactive action is applied. For non-interactive segments, a position code with a different specific rotation angle is applied, thereby achieving precise alignment of the "interactive text instruction" and the "corresponding video frame" in the time dimension. This helps to ensure the accuracy of recognizing and processing interactive actions using the temporal cross-attention layer.
[0114] Furthermore, in video data containing interactive actions, the proportion of interactive video frames is usually small, and conventional diffusion denoising loss can cause the model to easily overlook these interactive frames. To address this issue while ensuring the quality and effectiveness of the target video generated by the video generation model, motion-aware constraint loss can be introduced as part of the training strategy during the training process. Specifically, the training process for the video generation model can include: (1) During the training phase, an interaction action timing mask is generated based on the interaction action timestamp. The interaction action timing mask may include the values “1” and “0”. The value “1” is used to indicate that there is an interaction, and the value “0” is used to indicate that there is no interaction.
[0115] (2) When calculating the reconstruction loss, the video frames in the training data are divided into "video frames with interactive actions" and "video frames without interactive actions". The generation loss function corresponding to the video generation operation and the interaction loss function corresponding to the interactive action are obtained. Then, the total loss function is generated based on the generation loss function and the interaction loss function, and the diffusion model to be trained is optimized based on the total loss function, so that a video generation model based on the diffusion model can be stably obtained.
[0116] Specifically, the total loss function can be determined using the following formula: = in, For the total loss function, To generate the loss function, To preset weights, Let be the interaction loss function. Used to identify desired information, To predict noise, For noise, latent variables This refers to the preset step size used for filtering noise, or the interval step size used to filter noise information. For input data, To preset standard noise, This is a temporal mask corresponding to the interactive actions in the training video. This indicates element-wise multiplication (Hadamard product), which effectively ensures the accuracy and reliability of determining the total loss function.
[0117] Because the interaction loss function is given a higher penalty weight in the total loss function mentioned above. This forces the video generation model to prioritize and learn the generation of human-clothing interaction actions and complex clothing deformations during gradient backpropagation, greatly improving the efficiency of the video generation model in learning interactive action generation.
[0118] Furthermore, the training process for generating target videos can be divided into two stages, aiming to learn both non-interactive and interactive video virtual try-on capabilities, and following a learning process from easy to difficult. Specifically: In the first stage, the model is trained on a non-interactive video try-on dataset, with the interactive action category label set to an empty string. The goal of this stage is to teach the model how to generate high-fidelity clothing textures. In the second stage, the model continues to be trained on an interactive video try-on dataset. The goal of this stage is to teach the model to follow spatial and semantic guidance and to learn how to generate the physical deformation of clothing caused by interactive actions.
[0119] The virtual try-on video generation method based on spatial and semantic guidance provided in this application embodiment breaks through the limitation of existing video try-on that can only handle simple human movements (non-interactive). For the first time, it systematically solves the problem of generating human-clothing interaction actions that include complex physical deformations (such as pulling clothes, rolling up sleeves, and zipping up zippers), and significantly improves the success rate of video generation including complex human-clothing interaction actions. Specifically, this solution can achieve the following effects: 1) A spatial signal injection technology based on 3D hand prior was designed: For the first time, a 3D hand rendering image was introduced as a spatial prior in a video fitting task. A lightweight interactive action signal injection module was constructed. The fusion operation between the 3D hand rendering image and the text was realized through the constructed interactive action signal injection module. This enabled the video generation model to obtain accurate gesture shape, depth information or 3D spatial depth information, effectively solving the "hand-clothing clipping" and spatial ambiguity problems caused by 2D skeleton guidance. It successfully realized the realistic generation of complex human-clothing interactive actions such as pulling the corner of the clothes and rolling up the sleeves, overcoming the problem of "relying only on 2D key points to generate videos, thus failing to handle the occlusion and pulling relationship between the hands and clothes" in the related video generation process.
[0120] 2) A motion-aware rotational position encoding technique was designed: specifically based on timestamped interactive action tags, motion-aware rotational position encoding can apply different rotational position codes to interactive and non-interactive segments, enabling interactive and non-interactive segments to achieve different responses in the cross-attention mechanism, thus achieving "frame-level" precise control of text action commands. Through the aforementioned timestamped interactive action category tags and the designed motion-aware rotational position encoding, precise alignment of interactive semantic commands with specific video frames on the timeline was achieved. This ensures that the model generates interactive actions only within the correct time period, avoiding semantic interference with non-interactive frames. It also overcomes the problem of "missing action temporal semantics" in the related video generation process, achieving precise action temporal control.
[0121] 3) An action-aware constraint loss training mechanism was proposed: To address the pain point of extremely sparse distribution of interactive action data in real videos, a loss weighting strategy based on interactive action temporal mask was designed. The temporal mask is used to apply higher penalty weights to video frames containing interactive actions, forcing the model to focus on learning the generation of complex interactive actions. This guides the network to prioritize learning interactive action generation during training, thereby overcoming the lack of supervision of complex interactive features, greatly improving the success rate of generating interactive actions in target videos, while maintaining excellent generalization ability and spatiotemporal consistency.
[0122] Furthermore, in some of the processes described in the above embodiments and accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 11, 12, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0123] Figure 9 A schematic diagram of a video generation apparatus provided for an exemplary embodiment of this application; see attached diagram. Figure 9 As shown, this embodiment provides a video generation apparatus for performing the above-described... Figure 2 The video generation method shown may include: The first acquisition module 11 is used to acquire a reference video and a target object image. The reference video includes a person and a preset object. There is an interactive action between a part of the person's body and the preset object. The target object image includes a target object that is different from the preset object. The first determining module 12 is used to determine the context based on the reference video and the target object image. The context is used to identify the area where the preset object will be replaced with the target object. The first determining module 12 is also used to determine a 3D body partial rendering image corresponding to the interactive action based on the reference video. The first processing module 13 is used to generate a target video based on the target object image, context and 3D body partial rendering map. The target video includes a person and the target object, and there are interactive actions between the person's body parts and the target object.
[0124] In some instances, when the first determining module 12 determines the context based on the reference video and the target object image, the first determining module 12 is used to perform the following: determining the skeletal pose diagram of the person based on the reference video; determining the main body region of the object and a mask corresponding to the skeletal pose diagram based on the target object image, the mask being used to preserve the skeletal region in the skeletal pose diagram that is adapted to the target object; and determining the context based on the skeletal pose diagram, the main body region of the object, and the mask.
[0125] In some instances, when the first determining module 12 determines the context based on the skeletal pose map, the main body region of the object, and the mask, the first determining module 12 is used to perform the following: determining the main body mask corresponding to the main body region of the object based on the mask; and performing a weighted summation of the skeletal pose map and the main body region of the object based on the mask and the main body mask to determine the context.
[0126] In some instances, when the first processing module 13 generates a target video based on the target object image, context, and 3D body partial rendering, the first processing module 13 is used to perform the following: generating action category labels corresponding to interactive actions based on the reference video, with each action category label having a corresponding video timestamp; and generating the target video based on the target object image, context, 3D body partial rendering, and action category labels.
[0127] In some instances, when the first processing module 13 generates a target video based on the target object image, context, 3D body partial rendering map, and action category label, the first processing module 13 is used to perform: generating global descriptive text based on the reference video; and generating the target video based on the target object image, context, 3D body partial rendering map, global descriptive text, and action category label.
[0128] In some instances, when the first processing module 13 generates a target video based on the target object image, context, 3D body local rendering, global descriptive text, and action category labels, the first processing module 13 performs the following: calling a pre-trained diffusion model for implementing video generation operations, the diffusion model including at least a context module and a signal injection module; using the signal injection module to process the 3D body local rendering, global descriptive text, and action category labels to obtain 3D body local interaction signals; using the context module to process the 3D body local interaction signals, context, and target object image to obtain a context signal, the context signal being used to reflect the background, interactive actions, and target object in the reference video; and generating the target video based on the context signal.
[0129] In some instances, the diffusion model also includes a denoising module and an image decoder; when the first processing module 13 generates the target video based on the context signal, the first processing module 13 is used to perform: using the denoising module to process the initialized noisy latent variables and the context signal to obtain the target noise for generating the target video; and using the image decoder to decode the target noise to obtain the target video.
[0130] In some instances, the signal injection module includes a global cross-attention layer, a temporal cross-attention layer, and a feedforward neural network. When the first processing module 13 uses the signal injection module to process the 3D body local rendering image, global descriptive text, and action category labels to obtain 3D body local interaction signals, the first processing module 13 performs the following: processing the 3D body local rendering image and global descriptive text using the global cross-attention layer to obtain a first embedding representation; processing the first embedding feature and action category labels using the temporal cross-attention layer to obtain a second embedding representation; and processing the second embedding representation using the feedforward neural network to obtain 3D body local interaction signals.
[0131] In some instances, when the first processing module 13 processes the first embedding feature and action category label using a temporal cross-attention layer to obtain the second embedding representation, the first processing module 13 performs the following: based on the action category label, determining interactive and non-interactive segments in the reference video; determining the text embedding representation corresponding to the action category label; encoding the interactive and non-interactive segments using a rotational position encoding mechanism to obtain interactive segment representation and non-interactive segment representation; and processing the text embedding representation, interactive segment representation, and non-interactive segment representation using a temporal cross-attention layer to obtain the second embedding feature.
[0132] In some instances, when the first processing module 13 uses a rotational position encoding mechanism to encode interactive segments and non-interactive segments to obtain interactive segment representations and non-interactive segment representations, the first processing module 13 performs the following: based on the rotational position encoding mechanism, determining a first rotation angle corresponding to the interactive segment and a second rotation angle corresponding to the non-interactive segment, wherein the first rotation angle and the second rotation angle are different; based on the first rotation angle and the interactive segment, determining the interactive segment representation; based on the second rotation angle and the non-interactive segment, determining the non-interactive segment representation.
[0133] In some instances, before invoking the pre-trained diffusion model for video generation, the first acquisition module 11 and the first processing module 13 in this embodiment perform the following steps: The first acquisition module 11 is used to acquire video generation training data. The video generation training data includes initialized noise latent variables, a preset step size for filtering noise, input data for implementing video generation operations, and preset standard noise corresponding to a preset standard video. The input data includes: training video and training object image. The first processing module 13 is used to input the noise latent variable, the preset step size and the input data into the diffusion model to be trained to generate prediction noise; based on the training video, generate a temporal mask corresponding to the interactive actions in the training video; based on the temporal mask, the prediction noise and the preset standard noise, generate an interaction loss function corresponding to the interactive actions; and use the interaction loss function to optimize the diffusion model to be trained to generate a diffusion model.
[0134] In some instances, when the first processing module 13 optimizes the diffusion model to be trained using the interaction loss function to generate the diffusion model, the first processing module 13 performs the following: generating a generation loss function corresponding to the video generation operation based on the predicted noise and the preset standard noise; generating a total loss function based on the interaction loss function and the generation loss function; and optimizing the diffusion model to be trained based on the total loss function to generate the diffusion model.
[0135] The video generation device in this embodiment can also perform the aforementioned embodiments. Figures 1-8 The description of the embodiments shown is for reference only, and will not be elaborated upon here.
[0136] like Figure 10 As shown, this embodiment provides an electronic device, which can refer to a complete set of hardware and software infrastructure supporting software operation, data processing, and task execution. It can provide computing resources, storage, communication, and a runtime environment for applications (such as large model inference, scientific simulations, and web services). This electronic device is used to perform the above-described tasks. Figure 2 In practice, the electronic device shown in the video generation method may include a memory 24 and a processor 25.
[0137] Memory 24 is used to store computer programs and can be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device, data structures, contact data, phone book data, messages, pictures, videos, etc.
[0138] The processor 25, coupled to the memory 24, is configured to execute a computer program in the memory 24 for: acquiring a reference video and a target object image, the reference video including a person and a preset object, wherein there is an interactive action between a part of the person's body and the preset object, and the target object image including a target object different from the preset object; determining a context based on the reference video and the target object image, the context being used to identify the region in the reference video where the preset object is replaced with the target object; determining a 3D body part rendering image corresponding to the interactive action based on the reference video; and generating a target video based on the target object image, the context, and the 3D body part rendering image, the target video including a person and the target object, wherein there is an interactive action between a part of the person's body and the target object.
[0139] Furthermore, such as Figure 10 As shown, the electronic device also includes other components such as a communication component 26, a display 27, a power supply component 28, and an audio component 29. Figure 10 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 10 The components shown. Additionally... Figure 10 The components within the center frame are optional, not mandatory, and their specific requirements depend on the product form of the work node. In this embodiment, the work node can be a terminal device such as a desktop computer, laptop computer, smartphone, or Internet of Things (IoT) device, or a server-side device such as a conventional server, cloud server, or server array. If the work node in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 10 The components within the center frame; if the working node in this embodiment is implemented as a server-side device such as a conventional server, cloud server, or server array, it may not include... Figure 10 The component within the center frame.
[0140] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0141] The aforementioned communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.
[0142] The aforementioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0143] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0144] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0145] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile components, or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium.
[0146] Accordingly, this application also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor is able to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. In addition, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, so that the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device can be implemented as a means to implement the corresponding functions in the above method embodiments.
[0147] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0148] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A video generation method, characterized in that, include: Acquire a reference video and a target object image. The reference video includes a person and a preset object. There is an interactive action between a part of the person's body and the preset object. The target object image includes a target object that is different from the preset object. Based on the reference video and the target object image, a context is determined, wherein the context is used to identify the region in the reference video where the preset object is replaced with the target object; Based on the reference video, a 3D partial body rendering image corresponding to the interactive action is determined; Based on the target object image, context, and the 3D body partial rendering map, a target video is generated. The target video includes the person and the target object, and there is an interactive action between the person's body part and the target object.
2. The method according to claim 1, characterized in that, Based on the reference video and the target object image, the context is determined, including: Based on the reference video, the skeletal posture diagram of the person is determined; Based on the target object image, the main body region of the object and a mask corresponding to the skeletal pose map are determined. The mask is used to preserve the skeletal region in the skeletal pose map that is adapted to the target object. The context is determined based on the skeletal pose map, the main body region of the object, and the mask.
3. The method according to claim 2, characterized in that, Based on the skeletal pose map, the main body region of the object, and the mask, the context is determined, including: Based on the mask, determine the main body mask corresponding to the main body region of the object; The context is determined by performing a weighted summation of the skeletal pose map and the object's main body region based on the mask and the main body mask.
4. The method according to claim 1, characterized in that, Based on the target object image, context, and the 3D body partial rendering image, a target video is generated, including: Based on the reference video, generate action category tags corresponding to the interactive actions, and the action category tags correspond to video timestamps; The target video is generated based on the target object image, context, the 3D body partial rendering image, and the action category label.
5. The method according to claim 4, characterized in that, Based on the target object image, context, the 3D body partial rendering map, and the action category label, the target video is generated, including: Based on the reference video, generate global description text; The target video is generated based on the target object image, context, the 3D body partial rendering image, global descriptive text, and the action category label.
6. The method according to claim 5, characterized in that, Based on the target object image, context, the 3D body partial rendering map, global descriptive text, and the action category label, the target video is generated, including: The pre-trained diffusion model for video generation is invoked, and the diffusion model includes at least: a context module and a signal injection module; The signal injection module is used to process the 3D body local rendering image, global description text, and action category label to obtain 3D body local interaction signals. The context module is used to process the 3D body local interaction signal, the context, and the target object image to obtain a context signal, which is used to reflect the background, interactive actions, and target object in the reference video; The target video is generated based on the context signals.
7. The method according to claim 6, characterized in that, The diffusion model further includes: a denoising module and an image decoder; generating the target video based on the context signal includes: The denoising module is used to process the initialized noisy latent variables and the context signal to obtain target noise for generating the target video; The target noise is decoded using the image decoder to obtain the target video.
8. The method according to claim 6, characterized in that, The signal injection module includes: a global cross-attention layer, a temporal cross-attention layer, and a feedforward neural network; the signal injection module processes the 3D body local rendering image, global descriptive text, and action category labels to obtain 3D body local interaction signals, including: The 3D body local rendering map and the global descriptive text are processed using the global cross-attention layer to obtain the first embedded representation; The first embedded feature and the action category label are processed using the temporal cross-attention layer to obtain the second embedded representation; The second embedded representation is processed using the feedforward neural network to obtain the 3D body local interaction signal.
9. The method according to claim 8, characterized in that, The first embedded feature and the action category label are processed using the temporal cross-attention layer to obtain a second embedded representation, including: Based on the action category labels, determine the interactive and non-interactive segments in the reference video; Determine the text embedding representation corresponding to the action category label; The interactive segment and the non-interactive segment are encoded using a rotation position encoding mechanism to obtain the interactive segment representation and the non-interactive segment representation; The second embedding feature is obtained by processing the text embedding representation, the interactive segment representation, and the non-interactive segment representation using the temporal cross-attention layer.
10. The method according to claim 9, characterized in that, The interactive segment and the non-interactive segment are encoded using a rotational position encoding mechanism to obtain representations of the interactive segment and the non-interactive segment, including: Based on the rotation position encoding mechanism, a first rotation angle corresponding to the interactive segment and a second rotation angle corresponding to the non-interactive segment are determined, wherein the first rotation angle and the second rotation angle are different. Based on the first rotation angle and the interaction segment, the representation of the interaction segment is determined; Based on the second rotation angle and the non-interactive segment, the characterization of the non-interactive segment is determined.
11. The method according to any one of claims 6-10, characterized in that, Before invoking the pre-trained diffusion model used to perform the video generation operation, the method further includes: Acquire video generation training data, which includes initialized noise latent variables, a preset step size for filtering noise, input data for implementing video generation operations, and preset standard noise corresponding to a preset standard video. The input data includes: training video and training object image. The noise latent variable, the preset step size, and the input data are input into the diffusion model to be trained to generate predicted noise; Based on the training video, generate a temporal mask corresponding to the interactive actions in the training video; Based on the temporal mask, the predicted noise, and the preset standard noise, an interaction loss function corresponding to the interaction action is generated. The diffusion model to be trained is optimized using the interaction loss function to generate the diffusion model.
12. The method according to claim 11, characterized in that, Optimizing the diffusion model to be trained using the interaction loss function to generate the diffusion model includes: Based on the predicted noise and the preset standard noise, a generation loss function corresponding to the video generation operation is generated; Based on the interaction loss function and the generation loss function, a total loss function is generated; The diffusion model to be trained is optimized based on the total loss function to generate the diffusion model.
13. A video generation apparatus, characterized in that, include: The first acquisition module is used to acquire a reference video and a target object image. The reference video includes a person and a preset object. There is an interactive action between a part of the person's body and the preset object. The target object image includes a target object that is different from the preset object. The first determining module is used to determine a context based on the reference video and the target object image, wherein the context is used to identify the region where the preset object is replaced with the target object; The first determining module is further configured to determine a 3D body partial rendering image corresponding to the interactive action based on the reference video; The first processing module is used to generate a target video based on the target object image, the context, and the 3D body partial rendering image. The target video includes the person and the target object, and there is an interactive action between the person's body partial and the target object.
14. An electronic device, characterized in that, include: A memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the method of any one of claims 1-12.
15. A computer storage medium, characterized in that, Used to store a computer program that, when executed by a computer, implements the method of any one of claims 1-12.
16. A computer program product, characterized in that, include: A computer-readable storage medium storing computer instructions that, when executed by one or more processors, cause the processors to perform the steps of the method of any one of claims 1-12.