Motion control method, motion control model training method, motion control model training device and motion control model training equipment
By extracting video and image features and using frequency-aware weights for noise reduction, the problem of unsatisfactory appearance consistency and flexibility of action transfer in the prior art is solved, and a higher quality action control effect is achieved.
Patent Information
- Application Number
- CN202510199735.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art shows that the appearance consistency and flexibility are not ideal when transferring the action of the reference subject in the reference video to the target subject in the target image.
By extracting video features from the first video and image features from the first image, the initial fusion features are obtained, and the initial fusion features are performed on the initial fusion features in N stages through the diffusion neural network based on the frequency perception weight to obtain the final fusion features. The final fusion feature is used to generate a second video containing the target subject, and the action of the target subject is controlled based on the action of the reference subject.
The appearance consistency and accuracy of the action of the target subject in the generated video are improved, and the action is more similar to the reference subject and more flexibility.
Smart Images

Figure CN120050486A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a motion control method, a training method of a motion control model, a device, and a device. Background Art
[0002] With the continuous development of image processing technology, motion customization has gradually come into people's view at present. Motion customization involves generating a video, and the motion of the subject in the video is the motion indicated by the input control signal.
[0003] In related technologies, it is supported to transfer the motion of the reference subject in the reference video to the target subject in the target image to generate an action video of the target subject. For example, a specific low-rank adaptation method is used to transfer the motion of the reference subject in the reference video to the target subject in any target image to generate an action video of the target subject.
[0004] However, the action video of the target subject generated by the above method is still not ideal in terms of appearance consistency and flexibility. Summary of the Invention
[0005] The embodiments of the present application provide a motion control method, a training method of a motion control model, a device, and a device. The technical solutions are as follows:
[0006] According to one aspect of the embodiments of the present application, a motion control method is provided, and the method includes:
[0007] Extract a first video feature from a first video and a first image feature from a first image, where the first video feature is used to characterize the semantic feature of the reference subject in the first video, and the first image feature is used to characterize the semantic feature of the target subject in the first image;
[0008] Obtain an initial fusion feature according to the first video feature and the first image feature, where the initial fusion feature is used to fuse the semantic feature of the reference subject and the semantic feature of the target subject;
[0009] Based on the frequency-aware weight, perform noise reduction processing on the initial fusion feature through a diffusion neural network for N stages to obtain a final fusion feature, where the frequency-aware weight is used to control the attention weight of the first image at different frequencies, the attention weight is used to control the attention to the motion feature of the reference subject in the noise reduction processing, and the final fusion feature is used to fuse the motion feature of the reference subject and the spatial structure feature of the target subject, and N is a positive integer;
[0010] Generate a second video including the target subject based on the final fusion feature, and the action of the target subject in the second video is controlled based on the action of the reference subject in the first video.
[0011] According to one aspect of the embodiments of the present application, a method for training an action control model is provided, and the method includes:
[0012] Extract sample video features from a sample video and extract sample image features from a sample image. The sample video features are used to characterize the semantic features of the sample reference subject in the sample video, and the sample image features are used to characterize the semantic features of the sample target subject in the sample image;
[0013] Based on the sample video features, train a diffusion neural network to obtain a trained diffusion neural network. The diffusion neural network is used to perform noise reduction processing on the sample video features in N stages to obtain an output video, and the output video is a denoised video determined based on the sample video;
[0014] Based on the sample video features and the sample image features, train a frequency perception module to obtain a trained frequency perception module. The frequency perception module is used to determine a frequency perception weight, and the frequency perception weight is used to control the attention weight of the sample image at different frequencies. The attention weight is used to control the attention to the action features of the sample reference subject during the noise reduction processing;
[0015] Based on the trained diffusion neural network and the trained frequency perception module, construct an action control model. The action control model is used to generate a second video including the target subject based on a first image including the target subject and a first video including the reference subject, and the action of the target subject in the second video is controlled based on the action of the reference subject in the first video.
[0016] According to one aspect of the embodiments of the present application, an action control device is provided, and the device includes:
[0017] A feature extraction module, configured to extract first video features from a first video and extract first image features from a first image. The first video features are used to characterize the semantic features of the reference subject in the first video, and the first image features are used to characterize the semantic features of the target subject in the first image;
[0018] A feature fusion module, configured to obtain an initial fusion feature according to the first video features and the first image features. The initial fusion feature is used to fuse the semantic features of the reference subject and the semantic features of the target subject;
[0019] A noise reduction processing module, which is configured to perform N-stage noise reduction processing on the initial fusion features through a diffusion neural network based on frequency perception weights to obtain final fusion features. The frequency perception weights are used to control the attention weights of the first image at different frequencies, and the attention weights are used to control the attention to the action features of the reference subject during the noise reduction processing. The final fusion features are used to fuse the action features of the reference subject and the spatial structure features of the target subject, where N is a positive integer;
[0020] A video generation module, which is configured to generate a second video containing the target subject based on the final fusion features, and the action of the target subject in the second video is controlled based on the action of the reference subject in the first video.
[0021] According to one aspect of the embodiments of the present application, there is provided a training device for an action control model, and the device includes:
[0022] An extraction module, which is configured to extract sample video features from a sample video and extract sample image features from a sample image. The sample video features are used to characterize the semantic features of the sample reference subject in the sample video, and the sample image features are used to characterize the semantic features of the sample target subject in the sample image;
[0023] A first training module, which is configured to train a diffusion neural network based on the sample video features to obtain a trained diffusion neural network. The diffusion neural network is used to perform N-stage noise reduction processing on the sample video features to obtain an output video, and the output video is a denoised video determined based on the sample video;
[0024] A second training module, which is configured to train a frequency perception module based on the sample video features and the sample image features to obtain a trained frequency perception module. The frequency perception module is used to determine frequency perception weights, and the frequency perception weights are used to control the attention weights of the sample image at different frequencies, and the attention weights are used to control the attention to the action features of the sample reference subject during the noise reduction processing;
[0025] A construction module, which is configured to construct an action control model based on the trained diffusion neural network and the trained frequency perception module. The action control model is used to generate a second video containing the target subject based on a first image containing the target subject and a first video containing the reference subject, and the action of the target subject in the second video is controlled based on the action of the reference subject in the first video.
[0026] According to one aspect of the embodiments of the present application, a computer device is provided. The computer device includes a processor and a memory. A computer program is stored in the memory and is loaded and executed by the processor to implement the above-mentioned action control method or to implement the training method of the above-mentioned action control model.
[0027] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided. A computer program is stored in the computer-readable storage medium and is loaded and executed by a processor to implement the above-mentioned action control method or to implement the training method of the above-mentioned action control model.
[0028] According to one aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes a computer program, and the computer program is loaded and executed by a processor to implement the above-mentioned action control method or to implement the training method of the above-mentioned action control model.
[0029] The technical solution provided by the embodiments of the present application may include the following beneficial effects:
[0030] An initial fusion feature is obtained based on the first video feature and the first image feature, and then the attention weight of the first image at different frequencies is adjusted based on the frequency-aware weight, so that during the process of the diffusion neural network performing N-stage noise reduction processing on the initial fusion feature, it can shift from focusing on the spatial structure features of the target subject to focusing on the action features of the reference subject at different frequencies, so that the final fusion feature obtained can include more detailed action features, and the appearance consistency of the target subject in the finally generated second video is better, the action is more accurate, the action similarity with the reference subject in the first video is higher, and the flexibility is stronger. Description of the Drawings
[0031] Figure 1 is a schematic diagram of a computer system provided by an embodiment of the present application;
[0032] Figure 2 is a schematic diagram of a first video and a first image provided by an embodiment of the present application;
[0033] Figure 3 is a flowchart of an action control method provided by an embodiment of the present application;
[0034] Figure 4 is a schematic diagram of an action control model provided by an embodiment of the present application;
[0035] Figure 5 is an attention map between a frequency-aware module and a first image provided by an embodiment of the present application;
[0036] Figure 6 It is a flowchart of a motion control method in an animation production scenario provided by an embodiment of the present application;
[0037] Figure 7 It is a flowchart of a training method for a motion control model provided by an embodiment of the present application;
[0038] Figure 8 It is a schematic diagram of the training process of a diffusion neural network provided by an embodiment of the present application;
[0039] Figure 9 It is a schematic diagram of the training process of a frequency perception module provided by an embodiment of the present application;
[0040] Figure 10 It is a block diagram of a motion control device provided by an embodiment of the present application;
[0041] Figure 11 It is a block diagram of a training device for a motion control model provided by an embodiment of the present application;
[0042] Figure 12 It is a block diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0043] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0044] Please refer to Figure 1 , which shows a schematic diagram of a computer system provided by an embodiment of the present application. The computer system may include: a terminal device 100 and a server 200.
[0045] The terminal device 100 can be an electronic device such as a PC (Personal Computer), a mobile phone, a wearable device, a vehicle-mounted terminal, a VR (Virtual Reality) device, an AR (Augmented Reality) device, an MR (Mixed Reality) device, etc. An action control model can be arranged in the terminal device 100. The action control model can be applied to transfer the action of the reference subject in the reference video to the target subject in the target image, and generate an action video containing the target subject. Exemplarily, the first video contains the reference subject, and the first image contains the target subject. Optionally, the action control model is used to generate a second video containing the target subject based on the first video features extracted from the first video and the first image features extracted from the first image. The action of the target subject in the second video is controlled based on the action of the reference subject in the first video. Exemplarily, first, an initial fusion feature is obtained according to the first video features and the first image features. The initial fusion feature is used to fuse the semantic features of the reference subject and the semantic features of the target subject. Then, based on the frequency-aware weight, the initial fusion feature is subjected to N stages of noise reduction processing through a diffusion neural network to obtain a final fusion feature. The frequency-aware weight is used to control the attention weight of the first image at different frequencies. The attention weight is used to control the attention to the action features of the reference subject during the noise reduction processing. The final fusion feature is used to fuse the action features of the reference subject and the spatial structure features of the target subject. N is a positive integer. Finally, based on the final fusion feature, a second video containing the target subject is generated.
[0046] In some embodiments, the action control model can also be arranged in the server 200. The terminal device 100 obtains the input of the action control model and sends it to the server 200. The action control model in the server 200 generates a second video containing the target subject based on the input of the action control model sent by the terminal device 100 (the input of the action control model can include one or more of the first video, the first image, the first video features, the first image features, and the initial fusion feature). Then, the server 200 sends the second video to the terminal device 100.
[0047] The server 200 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0048] The terminal device 100 and the server 200 can communicate with each other through a network, such as a wired or wireless network.
[0049] In the action control method provided by the embodiments of this application, the execution entity of each step can be a computer device. The computer device refers to an electronic device with data calculation, processing, and storage capabilities. Taking Figure 1 the action control system shown as an example, the action control method can be executed by the terminal device 100, or by the server 200, or by the interaction and cooperation of the terminal device 100 and the server 200. This application does not make any limitations in this regard. For the convenience of description, in the following method embodiments, only the execution entity of each step of the action control method being a computer device will be introduced and described.
[0050] Action transfer involves applying a specific action pattern to a target subject and is widely used in film production, game development, and animation. However, this usually requires a large amount of capital and human resources. For example, a professional motion capture system may cost thousands to tens of thousands of dollars and requires skilled technicians. Similarly, creating a 30-second animation at 12 frames per second requires approximately six professional animators working a total of 20 days. These high costs pose a significant challenge and limit the access of many potential creators.
[0051] To facilitate controllable actions in video generation, methods can be roughly divided into two categories: (1) methods based on predefined signals such as poses and depth maps; (2) methods that focus on global actions. Despite the progress, these methods still have significant limitations. Methods based on predefined signals require strict alignment of the spatial structure (such as shape, skeleton, perspective) between the target image and the reference video, which is often not feasible in real-world scenarios. They also have difficulty obtaining pose information of non-human subjects. On the contrary, global action methods usually generate actions with a fixed layout and lack the ability to transfer actions across diverse subjects. Some methods use identity-specific Low-Rank Adaptation (LoRAs) to create animations for the target subject but still encounter difficulties in terms of appearance consistency and flexibility.
[0052] Action customization involves generating a video in which a subject performs an action indicated by an input control signal. Current methods use pose guidance or global motion customization, but are severely restricted by spatial structures such as layout, skeleton, and perspective consistency, reducing adaptability across different subjects and scenarios. To overcome these limitations, FlexiAct is proposed in this application, which can transfer the action in a reference video to an arbitrary target image. Different from the methods in the related art, FlexiAct allows differences in layout, perspective, and skeleton structure between the subject in the reference video and the target image while maintaining identity consistency. Achieving this requires precise action control, spatial structure adaptation, and consistency preservation. For this purpose, RefAdapter is introduced, which is a lightweight image-conditioned adapter that excels in spatial adaptation and consistency preservation and outperforms existing methods in balancing appearance consistency and structural flexibility. In addition, based on observations, the denoising process shows different degrees of attention to motion (low frequency) and appearance details (high frequency) at different time steps. Therefore, this application proposes FAE (Frequency-aware Action Extraction), which, different from related techniques that rely on separate spatio-temporal architectures, directly implements action extraction during the denoising process. Experiments show that the method of this application effectively transfers actions to subjects with different layouts, skeletons, and perspectives.
[0053] This application introduces FlexiAct, a versatile framework for action customization in heterogeneous scenarios. The method in this application can transfer actions from a reference video to a target image without alignment in layout, shape, or perspective, while preserving action dynamics and appearance details. It addresses three key challenges: (1) Spatial structure adaptation: Transferring actions to images with different poses, layouts, or perspectives. (2) Appearance domain bias: Solving the appearance inconsistency problem caused by the backbone model, especially when the reference video and the target image belong to different domains (e.g., real-world videos and anime images). (3) Precise action extraction and control: Accurately replicating the actions in the reference video rather than just capturing general actions.
[0054] To address the first two challenges, this application introduces RefAdapter (Rank Adapter), an image-conditioned architecture that decouples training from action extraction to maintain consistency. RefAdapter uses any frame as an image condition to adapt reference actions with different spatial structures, combining the consistency of Image-to-Video Generation (Image Prompt Adapter, Image Prompt Adapter) and the flexibility of architectures such as IPAdapter (Image Prompt Adapter). Different from ReferenceNet (Reference Network) and ControlNet (Control Network), RefAdapter fine-tunes a small group of LoRA parameters for appearance consistency, avoiding large-scale parameter replication.
[0055] For precise action control, this application proposes Frequency-Aware Action Extraction (FAE). This method uses trainable frequency-aware embeddings to learn to decouple actions from videos. The FAE embeddings adjust their sensitivity to different frequencies at denoising time steps, focusing on motion information (low frequencies) early and appearance details (high frequencies) later. By varying the attention weights at different time steps, FAE can directly perform action extraction during the denoising process without relying on a separate spatio-temporal architecture.
[0056] As Figure 2 shown, given a target image, FlexiAct can transfer the actions in the reference video to the target subject, enabling precise action adaptation and appearance consistency even in heterogeneous scenarios with different spatial structures or cross-domain subjects. For example Figure 2 The three groups a, b, and c included in it. In group a, the reference subject in the reference video is a human, and the target subject in the target image is also a human. The reference subject and the target subject have the same spatial structure; in group b, the reference subject in the reference video is a dog, and the target subject in the target image is a tiger. The reference subject and the target subject have similar spatial structures; in group c, the reference subject in the reference video is a human, and the target subject in the target image is a dog. The reference subject and the target subject are cross-domain subjects. Several possible implementations of the target image and the reference video are given in groups a, b, and c, which can achieve in heterogeneous scenarios where the reference subject and the target subject have different spatial structures and cross-domain subjects. The spatial structure can refer to the spatial structural components such as the shape, skeleton, and perspective of the subject (such as the target subject). Cross-domain subjects refer to the situation where the reference subject and the target subject are in different domains. The domain can include the time and space, species, object, etc. to which the reference subject and the target subject belong.
[0057] With this method, FlexiAct demonstrates its ability to handle complex action transfer tasks. It can not only transfer actions across different postures, layouts, or perspectives but also effectively address the appearance inconsistency issue caused by the reference video and the target image belonging to different domains (such as real-world videos and anime images). This enables creators to achieve high-quality action customization and animation production using FlexiAct even in resource-constrained situations, greatly expanding the possibilities of creative expression.
[0058] Please refer to Figure 3 , which shows a flowchart of an action control method provided by an embodiment of the present application, and this method is executed by a computer device. The method includes at least one of the following steps 310 to 340.
[0059] Step 310: Extract a first video feature from a first video and a first image feature from a first image. The first video feature is used to characterize the semantic feature of a reference subject in the first video, and the first image feature is used to characterize the semantic feature of a target subject in the first image.
[0060] In some embodiments, the computer device executes the above action control method through an action control model. In some embodiments, the input of the action control model can be the first video and the first image, or the first video feature and the first image feature. In one example, if the input of the action control model is the first video and the first image, then the action control model includes a preprocessing module for extracting the first video feature from the first video and the first image feature from the first image. In another example, if the input of the action control model is the first video feature and the first image feature, the first video feature is extracted from the first video by a preposed preprocessing model, and the first image feature is extracted from the first image by a preposed preprocessing model. The preprocessing module or the preprocessing model includes at least one feature extraction layer, and the feature extraction layer is used to extract the first video feature from the first video or the first image feature from the first image. The feature extraction layer for extracting the first video feature and the feature extraction layer for extracting the first image feature can be the same feature extraction layer or different feature extraction layers.
[0061] In some embodiments, the first video can be any video containing the action of a reference subject, and the first image can be any image containing a target subject. The reference subject refers to the moving object included in the first video, and the target subject refers to the static object included in the first image.
[0062] This application does not limit the space-time in which the reference entity and the target entity are located. Exemplarily, the reference entity and the target entity can be objects in different space-times. For example, the reference entity is an anime character and the target entity is a person in the real world. Exemplarily, the reference entity and the target entity can be different objects in the same space-time. For example, both the reference entity and the target entity are objects in the real world, such as the reference entity is a person and the target entity is a puppy.
[0063] This application also does not limit the species of the reference entity and the target entity (here, reproductive isolation is used as the basis for species classification). Exemplarily, the reference entity and the target entity can be of the same species. For example, both the reference entity and the target entity are animals. Exemplarily, the reference entity and the target entity can be of different species. For example, the reference entity is a person and the target entity is another animal (such as a dog, a cat, a tiger, a leopard). In some embodiments, when the reference entity and the target entity are of the same species, they can be the same object or different objects. Exemplarily, both the reference entity and the target entity are animals. In one example, the reference entity and the target entity are the same object, such as both the reference entity and the target entity are cats. In another example, the reference entity and the target entity are different objects, such as the reference entity is a puppy and the target entity is a kitten.
[0064] In some embodiments, the above several possibilities can be combined. Exemplarily, the reference entity and the target entity can be different species in different space-times, such as the reference entity is an animal character in an anime and the target entity is a puppy in the real world. Exemplarily, the reference entity and the target entity can be different objects of the same species in different space-times, such as the reference entity is a puppy in an anime and the target entity is a kitten in the real world. Exemplarily, the reference entity and the target entity can be different species in the same space-time, such as the reference entity is a puppy in the real world and the target entity is a person in the real world. Exemplarily, the reference entity and the target entity can be different objects of the same species in the same space-time, such as the reference entity is a puppy in the real world and the target entity is a wolf in the real world.
[0065] This application does not limit the method for extracting the first video feature from the first video. For example, any one of HOG (Histogram of Oriented Gradients), SIFT (Scale-Invariant Feature Transform), SUR (Speeded Up Robust Features), and ORB (Oriented FAST and Rotated BRIEF) can be used to extract the first video feature from the first video.
[0066] The present application does not limit the method for extracting the first image features from the first image either. For example, the first image features can be extracted through edge detection, corner detection, texture features, color features, and shape features.
[0067] In some embodiments, when selecting the method for extracting the first video features and / or the first image features, a method that is biased towards extracting the contour features in the video and / or image can be selected, so that the obtained first video features can better reflect the actions of the reference subject in the first video, and the first image features can better reflect the spatial structure of the target subject.
[0068] Step 320: Obtain an initial fusion feature according to the first video features and the first image features, where the initial fusion feature is used to fuse the semantic features of the reference subject and the target subject.
[0069] In some embodiments, the first video features and the first image features are concatenated to obtain the initial fusion feature. In some embodiments, the initial fusion feature includes the semantic features of the reference subject and the target subject. In some embodiments, the first video features are concatenated after the first image features to obtain an initial fusion feature starting with the first image features. In some embodiments, no changes need to be made to the first video features and the first image features during the concatenation operation, and only need to concatenate them in order.
[0070] In some embodiments, in order to assist the diffusion neural network to better understand the action features of the reference subject, when obtaining the initial fusion feature, in addition to the first video features and the first image features, the first text features are also concatenated.
[0071] In some embodiments, the first text features are extracted from the first text information, where the first text information is used to describe the actions of the reference subject in the first video, and the first text features are used to characterize the action features of the reference subject in the first video. Exemplarily, as shown in group b of Figure 2 the first text information can be that the reference subject is performing an action of leaping up and then falling down, during which the front limbs go from touching the ground to leaving the ground and gradually rise, and then fall back to the ground, and the hind limbs never leave the ground. The first text features are obtained by extracting the semantic features of the first text information. In some embodiments, the input of the action control model further includes the first text information or the first text features. In some embodiments, when the input of the action control model further includes the first text information, the preprocessing module also needs to have the ability to extract the semantic features of the text information. When the input of the action control model further includes the first text features, the preprocessing module also needs to have the ability to extract the semantic features of the text information. Exemplarily, the preprocessing module or the preprocessing model includes a text feature extraction layer for extracting the first text features from the first text information.
[0072] In some embodiments, the first video feature, the first image feature, and the first text feature are spliced to obtain an initial fusion feature. In this way, during the process of the diffusion neural network performing noise reduction processing on the initial fusion feature, it can better extract the action features of the reference subject from the initial fusion feature and fuse them with the spatial structure features of the target subject to generate a final fusion feature. In some embodiments, the first image feature, the first video feature, and the first text feature are connected together in sequence to obtain an initial fusion feature, and no special processing is required for the first video feature, the first image feature, and the first text feature during the splicing process.
[0073] Step 330, based on the frequency-aware weight, perform N stages of noise reduction processing on the initial fusion feature through a diffusion neural network to obtain a final fusion feature. The frequency-aware weight is used to control the attention weight of the first image at different frequencies, and the attention weight is used to control the attention to the action features of the reference subject during the noise reduction processing. The final fusion feature is used to fuse the action features of the reference subject and the spatial structure features of the target subject, where N is a positive integer.
[0074] In some embodiments, the attention weight applied during the noise reduction process is determined based on the frequency-aware weight and / or the initial attention weight. In some embodiments, the initial attention weight is a fixed parameter after the training of the action control model and can be regarded as a constant. In some embodiments, the frequency-aware weight can be a weight bias applied to the initial attention weight, or it can be the attention weight applied to the noise reduction process after applying a weight bias to the initial attention weight. In some embodiments, the frequency-aware weight is generated based on a frequency-aware module. In some embodiments, the frequency-aware module is used to control the frequency-aware weight of the target subject or the image containing the target subject at different frequencies. For example, the frequency-aware module is used to control the frequency-aware weight of the first image at different frequencies. The processing process of the frequency-aware module is not affected by the first image, but is related to the process of the N-stage noise reduction processing. The frequency-aware module will be described in detail in the subsequent training process of the action control model.
[0075] In some embodiments, the core idea of the diffusion neural network is to gradually add noise to the data (in this application, the data is an image or a video) by simulating the diffusion process, and learn to reverse this process to reconstruct the original data sample from the noise. Its core advantages are the ability to generate high-quality and high-resolution image and video content, and at the same time have strong scalability and flexibility. In some embodiments, the diffusion neural network includes N diffusion sub-networks, and the N diffusion sub-networks are connected in series in sequence. One diffusion sub-network is used to perform one-stage noise reduction processing on the intermediate fusion feature output by the previous diffusion sub-network. The intermediate fusion feature output by the Nth diffusion sub-network is the above-mentioned final fusion feature, and the input of the first diffusion sub-network is the above-mentioned initial fusion feature.
[0076] In some embodiments, the initial fusion feature is subjected to N-stage noise reduction processing through the diffusion neural network, from high frequency to low frequency, gradually changing from focusing on the spatial structure features of the target subject to focusing on the action features of the reference subject, so that the output final fusion feature can better represent the spatial structure features of the target subject and the fine action features of the reference subject.
[0077] Step 340, based on the final fusion feature, generate a second video including the target subject, and the action of the target subject in the second video is controlled based on the action of the reference subject in the first video.
[0078] In some embodiments, the final fusion feature is restored to obtain a second video including the target subject. In some embodiments, the action of the target subject in the second video is controlled based on the action of the reference subject in the first video. For example, the action of the target subject in the second video is the same as or similar to the action of the reference subject in the first video. It should be noted that since there may be differences in the space-time, species, and object between the reference subject and the target subject, there may be differences in the spatial structure between the reference subject and the target subject. Therefore, in the process of generating the second video including the target subject based on the first video including the reference subject and the first image including the target subject, the action of the target subject in the generated second video may not be exactly the same as the action of the reference subject in the first video, but an action adapted to the spatial structure characteristics of the target subject. For example, in the first video, it is the action of a puppy raising its front paw, and in the generated second video, it is the action of a tiger raising its front paw.
[0079] The technical solution provided by the embodiments of the present application obtains an initial fusion feature based on the first video feature and the first image feature, and then adjusts the attention weight of the first image at different frequencies based on the frequency-aware weight, so that during the process of the diffusion neural network performing N-stage noise reduction processing on the initial fusion feature, it can shift from focusing on the spatial structure features of the target subject to focusing on the action features of the reference subject at different frequencies, so that the final fusion feature obtained can include more detailed action features, and the appearance consistency of the target subject in the finally generated second video is better, the action is more accurate, the action similarity with the reference subject in the first video is higher, and the flexibility is stronger.
[0080] Please refer to Figure 4 , and next, taking the structure of the action control model shown in Figure 4 as an example, the action control method provided by the embodiments of the present application will be described in detail.
[0081] Step 310, extract the first video feature from the first video, and extract the first image feature from the first image. The first video feature is used to characterize the semantic feature of the reference subject in the first video, and the first image feature is used to characterize the semantic feature of the target subject in the first image.
[0082] 1. Extract the first video feature from the first video
[0083] In some embodiments, before extracting the first video feature from the first video, the first video is segmented into multiple video frames, and the first video feature is extracted based on the multiple video frames.
[0084] In some embodiments, the action extraction process may damage the consistency maintenance of the I2V model, and I2V is a strongly constrained image-conditioned framework. I2V uses the first video frame of the video as the conditional image to ensure the consistency of the video, but if the spatial structure of the first image is different from that of the first video, it may hinder the smooth action transfer. Therefore, during the training process, the conditional image needs to be replaced with other video frames to enhance the adaptability of the action control model to this situation. Of course, in the actual application process, the conditional image can also be replaced with other video frames. To reduce the complexity of the video frame replacing the conditional image, a video frame can be selected from the multiple video frames of the first video as the conditional image. Exemplarily, a video frame can be randomly selected from the multiple video frames of the first video as the conditional image. Exemplarily, the middle video frame or the last video frame of the multiple video frames of the first video can be used as the conditional image.
[0085] Exemplarily, the process of extracting the first video feature may include at least one of the following steps a to d.
[0086] Step a, obtain a first video, where the first video includes M video frames, and M is an integer greater than 1.
[0087] In some embodiments, the first video is a reference video provided by a video creator. In some embodiments, the first video can be a video intercepted from the reference video, and this interception can be performed in terms of time or on the video screen. In some embodiments, in the case of intercepting on the video screen, the first video can be randomly intercepted. Exemplarily, as Figure 4 shown, the reference video is randomly intercepted to obtain 6 intercepted videos as shown in Figure 4 shown, and one of the 6 intercepted videos is randomly selected as the first video.
[0088] In some embodiments, the M video frames are divided according to the frame rate of the first video. In some embodiments, after dividing according to the frame rate of the first video, the number of video frames of the first video obtained may be relatively large, and the processing pressure on the action control model will be relatively large. Therefore, the M video frames can be sampled from the video frames of the first video. In some embodiments, the value of M can be predefined, for example, it can be predefined by the video creator, or it can be predefined by the model developer during the training process of the action control model. In some embodiments, the first video frame set is obtained by dividing according to the frame rate of the first video, and the first video frame set includes P video frames, where P is an integer greater than or equal to M. M video frames are sampled from the P video frames. In some embodiments, the M video frames are evenly sampled from the P video frames. In some embodiments, the M video frames are randomly sampled from the P video frames. In some embodiments, the M video frames are arranged in the playing time sequence in the first video.
[0089] Step b, randomly select a video frame from the M video frames as the conditional image.
[0090] Step c, replace the first video frame of the first video with the conditional image to obtain the adjusted M video frames.
[0091] In some embodiments, the first video frame among the M video frames is replaced with the conditional image. In some embodiments, here, the conditional image is taken as an example of a randomly selected video frame for illustration, and the conditional image can also be a specified video frame other than the first video frame among the M video frames. Exemplarily, the conditional image can be the Mth video frame among the M video frames, or the Qth video frame among the M video frames, where Q is an integer greater than 1 and less than M.
[0092] Step d, extract the features of the conditional image and the adjusted M video frames to obtain the first video features.
[0093] In some embodiments, the features of the conditional image and the adjusted M video frames are respectively extracted to obtain the first video feature. In some embodiments, it may also be to extract the features of the conditional image and the features of the M video frames, and use the features of the conditional image to replace the features of the first video frame among the features of the M video frames to obtain the adjusted M video frames.
[0094] In some embodiments, the features of the conditional image and the adjusted M video frames are concatenated to obtain the first video feature. In some embodiments, the features of the conditional image and the adjusted M video frames are concatenated in order without special processing.
[0095] In some embodiments, in order to ensure the smoothness of the transition of the actions of the reference subject included in the conditional image and the remaining video frames, a parameter also needs to be concatenated to the features of the conditional image. Exemplarily, the features of the conditional image are extracted to obtain the first conditional feature, and the first conditional feature is used to characterize the semantic features of the reference subject in the conditional image; the first conditional feature and the first parameter are concatenated to obtain the second conditional feature, and the first parameter is used to smoothly transition the actions of the reference subject in the conditional image and the first video. In some embodiments, the first parameter is concatenated after the first conditional feature to obtain the second conditional feature, and the second conditional feature can be regarded as a set in which the first conditional feature and the first parameter are arranged in order.
[0096] In some embodiments, the features of the adjusted M video frames are extracted to obtain the second video feature, and the second video feature is used to characterize the semantic features of the reference subject in the adjusted M video frames; a first noise is added to the second video feature to obtain the third video feature, and the first noise is used to blur the spatial structure features of the reference subject in the second video feature; according to the second conditional feature and the third video feature, the first video feature is obtained.
[0097] In some embodiments, the above first noise includes but is not limited to Gaussian noise, Poisson noise, multiplicative noise, and salt-and-pepper noise. In some embodiments, the first noise is randomly generated. Exemplarily, the first noise is randomly generated Gaussian noise. In some embodiments, adding noise means adding noise data that satisfies the distribution of the first noise to the second video feature. Taking the first noise as Gaussian noise as an example, adding the first noise to the second video feature means adding noise data that satisfies the Gaussian noise distribution to the second video feature.
[0098] In some embodiments, the second conditional feature and the third video feature are concatenated to obtain the first video feature. In some embodiments, the second conditional feature and the third video feature are connected in order to obtain the first video feature.
[0099] Add noise to the second video feature to blur the spatial structure feature of the reference subject in the second video feature, so as to avoid being affected by the spatial structure feature of the reference subject during the subsequent noise reduction process of the diffusion neural network and ensure the effect of the finally output final fusion feature.
[0100] In some embodiments, the duration corresponding to the first video feature needs to be the same as the duration of the expected second video. Therefore, attention also needs to be paid to the duration of the video feature obtained based on the second conditional feature and the third video feature. The duration corresponding to the first video feature refers to the duration of the video restored based on the first video feature. In some embodiments, the second conditional feature and the third video feature are spliced to obtain a fourth video feature. If the duration corresponding to the fourth video feature is less than the first duration, then the fourth video feature needs to be filled so that the duration corresponding to the fourth video feature is not less than the first duration to ensure the effect of the finally generated second video. Herein, the first duration is the duration of the expected second video.
[0101] In some embodiments, when the duration corresponding to the fourth video feature is not less than the first duration, the fourth video feature is determined as the first video feature, where the first duration is the duration of the expected second video; when the duration corresponding to the fourth video feature is less than the first duration, preset parameters are filled for the fourth video feature to obtain the first video feature.
[0102] In some embodiments, when filling the preset parameters for the fourth video feature, the filling is performed between the second conditional feature and the third video feature, so that the transition between the actions of the first video frame and the subsequent video frames in the finally obtained final fusion feature based on the first video feature will be smoother. In some embodiments, even if the duration corresponding to the fourth video feature is not less than the first duration, preset parameters can also be filled between the second conditional feature and the third video feature to make the transition between the actions of the first video frame and the subsequent video frames in the final fusion feature smoother.
[0103] In some embodiments, the preset parameter refers to a preset feature value, and filling the preset parameter means filling the parameter used to represent the features of at least one video frame. Exemplarily, the preset parameter is 0, that is, 0 is used to represent the features of the at least one filled video frame.
[0104] 2. Extract the first image feature from the first image
[0105] In some embodiments, the process of extracting the first image feature is relatively simple. Exemplarily, the process of extracting the first image feature may include at least one of the following steps a to c.
[0106] Step a, obtain the first image.
[0107] Step b: Extract the features of the first image to obtain second image features, which are used to characterize the semantic features of the first image.
[0108] Step c: Add second noise to the second image features to obtain first image features. The second noise is used to blur the action features of the target subject in the second image features.
[0109] In some embodiments, the above-mentioned second noise includes but is not limited to Gaussian noise, Poisson noise, multiplicative noise, and salt-and-pepper noise. In some embodiments, the second noise is randomly generated. Exemplarily, the second noise is randomly generated Gaussian noise.
[0110] Regarding the relevant content of extracting the first image features, reference can be made to the introduction of extracting the first video features in the above embodiments. The difference is that only one first image needs to be processed in the process of extracting the first image features, and there is no need to process multiple video frames. This application will not elaborate here.
[0111] Adding noise to the second image features to blur the action features of the target subject in the second image features can avoid being affected by the action features of the target subject during the subsequent noise reduction process of the diffusion neural network, and ensure the effect of the finally output final fusion features.
[0112] Step 320: Obtain initial fusion features based on the first video features and the first image features. The initial fusion features are used to fuse the semantic features of the reference subject and the target subject.
[0113] In some embodiments, in order to assist the diffusion neural network to better understand the action features of the reference subject, when obtaining the initial fusion features, in addition to the first video features and the first image features, the first text features are also concatenated.
[0114] In some embodiments, the first text features are extracted from the first text information, where the first text information is used to describe the actions of the reference subject in the first video, and the first text features are used to characterize the action features of the reference subject in the first video; the first video features, the first image features, and the first text features are concatenated to obtain the initial fusion features. Exemplarily, as Figure 4 shown, after the first text features, the first video features, and the first image features are concatenated, they are used as the initial fusion features and input into the diffusion neural network.
[0115] Step 330: Based on the frequency-aware weights, perform N-stage noise reduction processing on the initial fusion features through the diffusion neural network to obtain the final fusion features.
[0116] In some embodiments, the diffusion neural network includes N diffusion sub-networks, where the i-th diffusion sub-network is used to perform noise reduction processing in the i-th stage. In some embodiments, the noise reduction processing includes at least one of the following steps a to c.
[0117] Step a, for the i-th stage among the N stages, concatenate the intermediate fusion feature of the (i - 1)-th stage and the frequency-aware weight of the i-th stage to obtain the concatenated feature of the i-th stage, and the concatenated feature is used as the input of the i-th diffusion sub-network among the N sub-networks; where i is a positive integer less than or equal to N.
[0118] In some embodiments, when i equals 1, the intermediate fusion feature of the (i - 1)-th stage is the initial fusion feature. In some embodiments, the intermediate fusion feature of the (i - 1)-th stage is the output of the (i - 1)-th diffusion sub-network.
[0119] In some embodiments, the frequency-aware weight of the i-th stage is determined based on the initial attention weight. In some embodiments, the initial attention weight is a parameter of the diffusion sub-network, and this parameter is determined during the training process.
[0120] In some embodiments, the frequency-aware weight is determined based on the initial attention weight and the weight bias, so it is necessary to obtain the initial attention weight first. In some embodiments, the frequency-aware weight is a weight configuration based on the initial attention weight, and in this case, it is not necessary to obtain the initial attention weight.
[0121] Exemplarily, regarding the method of obtaining the initial attention weight, the steps of determining the frequency-aware weight may include at least one of the following steps 1 to 2.
[0122] Step 1, obtain the initial attention weight.
[0123] Step 2, determine the relationship between i and at least one threshold.
[0124] In some embodiments, for the i-th stage among the N stages, when i is less than the first threshold, the initial attention weight is determined as the frequency-aware weight of the i-th stage; when i is greater than or equal to the first threshold and less than the second threshold, the sum of the initial attention weight and the first weight bias is determined as the frequency-aware weight of the i-th stage; when i is greater than or equal to the second threshold and less than or equal to N, the sum of the initial attention weight and the second weight bias is determined as the frequency-aware weight of the i-th stage.
[0125] In some embodiments, the second weight bias is greater than the first weight bias.
[0126] Exemplarily, the frequency-aware weight can be determined based on the following formula:
[0127]
[0128] W attn = W ori + W bias
[0129] wherein, W bias represents the weight bias, where is the first weight bias, α is the second weight bias, W ori is the initial attention weight, W attn is the frequency perception weight, t h is the first threshold, t l is the second threshold.
[0130] In some embodiments, t represents the time step, which can be understood as i, that is, the W at the t-th time step attn is the attention weight applied by the i-th diffusion sub-network, and t = i.
[0131] Exemplarily, for a method that does not require obtaining the initial attention weight, the steps of determining the frequency perception weight may include the following step 1.
[0132] Step 1, determine the relationship between i and at least one threshold.
[0133] In some embodiments, for the i-th stage among N stages, when i is less than the first threshold, the frequency perception weight of the i-th stage is determined to be 0; when i is greater than or equal to the first threshold and less than the second threshold, the frequency perception weight of the i-th stage is determined to be the first weight bias; when i is greater than or equal to the second threshold and less than or equal to N, the frequency perception weight of the i-th stage is determined to be the second weight bias.
[0134] It should be noted that in this case, the weight applied by the attention layer of the i-th diffusion sub-network is the sum of the initial attention weight and the frequency perception weight.
[0135] Exemplarily, the frequency perception weight can be determined based on the following formula:
[0136]
[0137] W attn = W ori + W bias
[0138] wherein, W bias represents the frequency perception weight, where is the first weight bias, α is the second weight bias, W ori is the initial attention weight, W attnis the weight applied by the attention layer of the i-th diffusion sub-network, t h is the first threshold, t l is the second threshold.
[0139] In some embodiments, t represents the time step, which can be understood as i, that is, W at the t-th time step attn is the attention weight applied to the i-th diffusion sub-network, and t = i.
[0140] The method for calculating the frequency-aware weight is given in the above embodiments, so that the attention weight of the first image to the action features of the reference subject changes with frequency. In the case of low frequency, it focuses on the action features, and in the case of high frequency, the attention is distributed over the entire image, paying attention to global details such as the background, so that the final fused features obtained can take into account both the action features and the spatial structure features.
[0141] Step b, perform noise reduction processing on the concatenated features of the i-th stage through the i-th diffusion sub-network to obtain the intermediate fused features of the i-th stage.
[0142] In some embodiments, the diffusion sub-network includes an attention layer, a first rank adapter, a feed-forward network, and a second rank adapter connected in sequence. Among them, the attention layer is used to capture the spatial layout features included in the concatenated features, the first rank adapter is used to reduce the dimension of the parameters output by the attention layer, the feed-forward network is used to perform a non-linear transformation on the features input to the feed-forward network from the first rank adapter, and the second rank adapter is used to reduce the dimension of the parameters output by the feed-forward network. The output of the second rank adapter is the intermediate fused feature.
[0143] The attention layer is one of the core components in a diffusion neural network (for example, the diffusion neural network is a DIT module), mainly used to capture complex relationships and patterns in the input data. The attention layer mainly includes the following implementation methods: 3D full attention, multi-head self-attention, cross-attention, and linear attention. Figure 4 The diffusion sub-network shown in includes a 3D full attention layer.
[0144] The rank adapter is an efficient parameter fine-tuning method, usually used for task adaptation based on a pre-trained model. By inserting an additional adapter module (such as after the attention layer or the feed-forward network) in the Transformer layer, the model is adjusted in a lightweight manner. The core idea of the rank adapter is to project the high-dimensional output of the original model into a low-dimensional space for processing through a feed-forward network for dimension reduction and then restore it to the original dimension. The advantage of this method is that only a small number of parameters (such as less than 5% of the model parameters) need to be updated to significantly improve the performance of the model on specific tasks.
[0145] The feed-forward network is an important part of the diffusion neural network (e.g., the diffusion neural network is a DIT module), and is usually located in each Transformer block. Its role is to perform non-linear transformation and feature extraction on the input features. In the DIT module, the feed-forward network usually consists of two linear layers, with a non-linear activation function (such as GELU or ReLU) in between. Through this design, the FFN can further enhance the model's expressive ability and feature learning ability. In addition, the FFN can also be combined with the adapter module for fine-tuning specific tasks.
[0146] In some embodiments, taking the noise reduction process of the i-th stage as an example, the i-th diffusion sub-network extracts the global information of the first video feature based on the attention layer to obtain the attention feature. The first rank adapter reduces the dimension of the attention feature to obtain the dimension-reduced attention feature. Then, the feed-forward network extracts the action feature and the spatial structure feature of the target subject in the dimension-reduced attention feature, and outputs the feed-forward feature. Then, the second rank adapter reduces the dimension of the feed-forward feature, and determines the dimension-reduced feed-forward feature as the intermediate fusion feature of the i-th stage.
[0147] Step c, determine the size of i. When i is less than N, let i = i + 1, and start executing again from the step of splicing the intermediate fusion feature of the (i - 1)-th stage and the frequency perception weight of the i-th stage; when i is equal to N, determine the intermediate fusion feature of the i-th stage as the final fusion feature.
[0148] In some embodiments, N is the number of diffusion sub-networks.
[0149] Through the above method, after the noise reduction process of N stages, the diffusion neural network can fuse the action feature of the reference subject and the spatial structure feature of the target subject in the initial fusion feature to obtain the final fusion feature. In the second video restored based on the final fusion feature, the target subject can present the action controlled by the action of the reference subject in the first video.
[0150] Step 340, generate a second video containing the target subject based on the final fusion feature, and the action of the target subject in the second video is controlled by the action of the reference subject in the first video.
[0151] In some embodiments, a second video containing the target subject is restored based on the final fusion feature. In some embodiments, the process of generating the second video based on the final fusion feature is the inverse process of extracting the first video feature from the first video.
[0152] In some embodiments, before generating a second video including a target subject based on the final fusion feature, it is also necessary to determine whether the final fusion feature meets the first condition. In the case where the final fusion feature meets the first condition, the step of generating a second video including the target subject based on the final fusion feature is executed; in the case where the final fusion feature does not meet the first condition, the final fusion feature is used as the updated initial fusion feature, and the step of obtaining the final fusion feature by performing N-stage noise reduction processing on the initial fusion feature through a diffusion neural network based on frequency-aware weights is executed again. After the final fusion feature meets the first condition, a second video including the target subject is generated to ensure the quality of the generated second video.
[0153] Regarding the setting of the first condition, it is also necessary to determine it in combination with the actual situation of the motion control model. Exemplarily, the first condition includes: the final fusion feature is obtained after T times of N-stage noise reduction processing, and T is a positive integer.
[0154] In this case, the frequency-aware weight can be determined based on T. Taking the following formula as an example:
[0155]
[0156] W attn =W ori +W bias
[0157] Wherein, W bias represents the frequency-aware weight, where is the first weight bias, α is the second weight bias, W ori is the initial attention weight, W attn is the weight applied by the attention layer of the i-th diffusion sub-network, t h is the first threshold, t l is the second threshold.
[0158] Wherein, t represents the time step, that is, after t times of N-stage noise reduction processing. Among them, in 1 time of N-stage noise reduction processing, the weights applied by the diffusion sub-network in each stage of noise reduction processing are the same.
[0159] Exemplarily, such as Figure 5The figure shows an attention map between the frequency-aware weight and the first image. At a relatively large time step (e.g., step = 800), the frequency-aware embedding focuses on action information (low frequency) and pays high attention to the moving parts of the reference subject. At an intermediate time step (e.g., step = 500), the attention turns to the fine-grained details of the reference subject. At a relatively late time step (e.g., step = 200), the attention is distributed across the entire image, indicating that global details such as the background are being attended to. The attention weight of the first image to the frequency-aware embedding is increased at relatively large time steps, while the original weight is maintained at other time steps. This enhances the ability of the generated second video to perceive and replicate the actions of the reference subject in the first video.
[0160] In some embodiments, the technical solutions provided by the embodiments of the present application can be applied to different scenarios. Exemplarily, it can be applied to an animation production scenario. In this case, the first video is a reference animation video, and the first image is a target animation image. Different animation characters can adopt the same actions, and animators do not need to re-configure a new action video for the target animation characters in the target animation image, reducing the workload of animators.
[0161] Exemplarily, as Figure 6 shown, the action control method in the animation production scenario includes at least one of the following steps 610 to 640.
[0162] Step 610: Extract reference animation video features from the reference animation video and target animation image features from the target animation image. The reference animation video features are used to represent the semantic features of the reference animation character in the reference animation video, and the target animation image features are used to represent the semantic features of the target animation character in the target animation image.
[0163] Step 620: Obtain an initial fusion feature based on the reference animation video features and the target animation image features. The initial fusion feature is used to fuse the semantic features of the reference animation character and the semantic features of the target animation character.
[0164] Step 630: Based on the frequency-aware weight, perform N-stage noise reduction processing on the initial fusion feature through a diffusion neural network to obtain a final fusion feature. The frequency-aware weight is used to control the attention weight of the target animation image at different frequencies, and the attention weight is used to control the attention to the action features of the reference animation character during the noise reduction process. The final fusion feature is used to fuse the action features of the reference animation character and the spatial structure features of the target animation character. N is a positive integer.
[0165] Step 640: Based on the final fused features, generate a second video including the target anime character, where the actions of the target anime character in the second video are controlled based on the actions of the reference anime character in the reference anime video.
[0166] For the specific details of the embodiments of the present application, reference may be made to the content in the above embodiments, and the present application will not elaborate herein.
[0167] Next, the training process of the action control model will be described. Please refer to Figure 7 , which shows a flowchart of a method for training an action control model provided by an embodiment of the present application. The method includes at least one of the following steps 710 to 740.
[0168] Step 710: Extract sample video features from the sample video and sample image features from the sample image. The sample video features are used to represent the semantic features of the sample reference subject in the sample video, and the sample image features are used to represent the semantic features of the sample target subject in the sample image.
[0169] In some embodiments, the sample video can be any video containing the sample reference subject. The sample reference subjects included in different sample videos can be the same or different. In some embodiments, the sample video can be obtained from an existing sample set. In some embodiments, the sample image can be any image containing the sample target subject.
[0170] In some embodiments, the method for extracting sample video features from the sample video can refer to the method for extracting first video features from the first video in the above embodiments, and the method for extracting sample image features from the sample image can refer to the method for extracting first image features from the first image in the above embodiments. The present application will not elaborate herein.
[0171] Step 720: Based on the sample video features, train a diffusion neural network to obtain a trained diffusion neural network. The diffusion neural network is used to perform noise reduction processing on the sample video features in N stages to obtain an output video, and the output video is a denoised video determined based on the sample video.
[0172] Step 730: Based on the sample video features and sample image features, train a frequency perception module to obtain a trained frequency perception module. The frequency perception module is used to determine frequency perception weights, and the frequency perception weights are used to control the attention weights of the sample image at different frequencies. The attention weights are used to control the attention to the action features of the sample reference subject during the noise reduction processing.
[0173] In some embodiments, during the training process, the diffusion neural network and the frequency perception module are trained separately to ensure the consistency between the trained diffusion neural network and the trained frequency perception module.
[0174] The training processes of the diffusion neural network and the frequency perception module will be described in detail in the following embodiments.
[0175] Step 740: Based on the trained diffusion neural network and the trained frequency perception module, construct an action control model. The action control model is used to generate a second video containing the target subject based on a first image containing the target subject and a first video containing a reference subject, and the action of the target subject in the second video is controlled based on the action of the reference subject in the first video.
[0176] In some embodiments, based on the preset structure of the action control model, splice the trained diffusion neural network and the trained frequency perception module to obtain the action control model. In some embodiments, connect the trained diffusion neural network and the trained frequency perception module according to the preset structure to obtain the action control model.
[0177] In one example, the action control model includes the trained diffusion neural network and the trained frequency perception module. The input of the action control model can be the first video feature, the first image feature, and the first text feature, or the first video feature, the first image feature, and the first text information. The output of the action control model is a second video containing the target subject. In this case, there is also a preprocessing model before the action control model.
[0178] In another example, the action control model includes a preprocessing module, the trained diffusion neural network, and the trained frequency perception module. The input of the action control model can be the first video, the first image, and the first text feature, or the first video, the first image, and the first text information. The output of the action control model is a second video containing the target subject. Among them, the preprocessing module is used to perform preprocessing on one or more of the first video, the first image, and the first text information respectively. In this case, during the process of constructing the action control model, it is necessary to connect the preprocessing module, the diffusion neural network, and the frequency perception module according to the preset structure.
[0179] The technical solution provided by the embodiments of the present application gives a training method for the action control model. Separately training the diffusion neural network and the frequency perception module can ensure the consistency between the trained diffusion neural network and the trained frequency perception module.
[0180] Next, the training processes of the diffusion neural network and the frequency perception module will be described in detail respectively.
[0181] 1. Diffusion Neural Network
[0182] In some embodiments, the diffusion neural network includes N diffusion sub-networks, and during the training process, the parameters of the N diffusion sub-networks also need to be adjusted.
[0183] In some embodiments, the training process of the diffusion neural network includes at least one of the following steps a to b.
[0184] Step a, for the i-th stage among the N stages, perform noise reduction processing on the intermediate output features of the (i - 1)-th stage through the i-th diffusion sub-network among the N diffusion sub-networks to obtain the intermediate output features of the i-th stage; where i is a positive integer less than or equal to N. When i equals 1, the intermediate output features of the (i - 1)-th stage are sample video features, and when i equals N, the intermediate output features of the i-th stage are used to generate the output video.
[0185] In some embodiments, during the process of training the diffusion neural network, there is no need to input features related to the sample image, which is mainly used to train the adaptability of the diffusion neural network to the spatial structure.
[0186] In some embodiments, when i equals 1, the intermediate output features of the (i - 1)-th stage are the initial fusion features generated based on the sample video features. In some embodiments, the initial fusion features are obtained by concatenating the sample video features and the sample text features. In some embodiments, steps c and d are further included before step a.
[0187] Step c, extract sample text features from the sample text information, where the sample text information is used to describe the actions of the sample reference subject in the sample video, and the sample text features are used to characterize the action features of the sample reference subject in the sample video.
[0188] Step d, concatenate the sample video features and the sample text features to obtain the initial fusion features.
[0189] Regarding the working process of the diffusion neural network, please refer to the introduction in the above embodiments regarding the action control method, and the present application will not elaborate here.
[0190] Step b, adjust the parameters of the diffusion neural network based on the output video and the sample video to obtain the trained diffusion neural network.
[0191] In some embodiments, the diffusion sub-network includes: an attention layer, a first rank adapter, a feed-forward network, and a second rank adapter; where the attention layer is used to capture the spatial layout features included in the sample video features, the first rank adapter is used to reduce the dimension of the parameters output by the attention layer, the feed-forward network is used to perform a non-linear transformation on the features input to the feed-forward network from the first rank adapter, and the second rank adapter is used to reduce the dimension of the parameters output by the feed-forward network.
[0192] In some embodiments, during the training process of the diffusion neural network, only the parameters of the first rank adapter and / or the second rank adapter need to be adjusted. Exemplarily, based on the output video and the sample video, the parameters of at least one of the first rank adapter and the second rank adapter are adjusted to obtain the trained diffusion neural network.
[0193] In some embodiments, a first loss function is calculated based on the output video and the sample video, and the parameters of at least one of the first rank adapter and the second rank adapter are adjusted based on the first loss function to obtain the trained diffusion neural network. In some embodiments, the first loss function can be any function that can be applied to calculate the image loss. Exemplarily, the first loss function includes but is not limited to at least one of the following: MAE (Mean Absolute Error), GCE (Generalized Cross Entropy), FL (Focal Loss), NLNL (Negative Learning of Noisy Labels), RCE (Reverse Cross Entropy), SCE (Symmetric Cross Entropy).
[0194] Exemplarily, Figure 8 Shown is a schematic diagram of the training process of the diffusion neural network. Directly using I2V for spatial structure adaptation is challenging because the action extraction process may damage the consistency maintenance of the I2V model, and I2V is a strongly constrained image-conditioned framework. During the training process, I2V uses the first image frame of the video as the conditional image to ensure the consistency of the video, but if the initial spatial structure is different from the reference video, it may hinder smooth action transfer.
[0195] To solve this problem, a gap is introduced between the conditional image and the initial spatial structure, and a sample video frame randomly extracted from the sample video is used as the conditional image during the training process. Specifically, during the I2V training process, the first sample video frame is replaced with a sample video frame randomly extracted from the sample video as the conditional image to maximize the spatial structure difference. This method utilizes the context learning ability of MMDiT, performs comprehensive conditional perception through a 3D full attention layer, and does not require replicating the network architecture to obtain strong control for each layer. By doing so, the network can be trained with the fewest additional parameters. Instead of adopting a heavy architecture like ControlNet or ReferenceNet, appearance learning is achieved by training a low-rank adapter with a small number of parameters. During the RefAdapter training, the input video latent variable used is different from the latent variable in the I2V training, and the first sample video frame embedding is replaced with the conditional image embedding to ensure stable training.
[0196] In the above embodiments, the training process of the diffusion neural network is given. In this process, a sample video frame randomly selected from the sample video is used as the conditional image to maximize the spatial structure difference, and then the attention layer included in the diffusion neural network is used for comprehensive conditional perception, reducing the volume of the diffusion neural network while ensuring stable training, so that the trained diffusion neural network can better identify the spatial structure features.
[0197] 2. Frequency Perception Module
[0198] In some embodiments, the training process of the frequency perception module may include at least one of the following steps a to f.
[0199] Step a, extracting sample text features from the sample text information, where the sample text information is used to describe the actions of the sample reference subject in the sample video, and the sample text features are used to characterize the action features of the sample reference subject in the sample video.
[0200] Step b, concatenating the sample video features, sample image features, and sample text features to obtain the initial fusion features.
[0201] Step c, for the i-th stage among the N stages, concatenating the intermediate output features of the (i - 1)-th stage and the frequency perception weights of the i-th stage to obtain the concatenated features of the i-th stage. The concatenated features are used as the input of the i-th diffusion sub-network among the N diffusion sub-networks, and the frequency perception weights of the i-th stage are determined based on the frequency perception module; where i is a positive integer less than or equal to N. When i equals 1, the intermediate fusion features of the (i - 1)-th stage are the initial fusion features.
[0202] Step d, performing noise reduction processing on the concatenated features of the i-th stage through the i-th diffusion sub-network to obtain the intermediate fusion features of the i-th stage.
[0203] Step e, determining the size relationship between i and N. When i is less than N, let i = i + 1, and start executing the step of concatenating the intermediate output features of the (i - 1)-th stage and the frequency perception weights of the i-th stage again to obtain the concatenated features of the i-th stage; when i equals N, determine the intermediate fusion features of the i-th stage as the final fusion features, and the final fusion features are used to fuse the action features of the sample reference subject and the spatial structure features of the sample target subject.
[0204] Step f, adjusting the parameters of the frequency perception module based on the final fusion features and the initial fusion features to obtain the trained frequency perception module.
[0205] In some embodiments, the diffusion neural network applied during the process of training the frequency perception module is a trained diffusion neural network. Regarding the relevant content of the diffusion neural network during the training process, reference can be made to the introduction in the above embodiments, and details are not described herein again in this application.
[0206] In some embodiments, based on the final fusion feature and the initial fusion feature, a second loss function is calculated, and the parameters of the frequency perception module are adjusted based on the second loss function. In some embodiments, the second loss function can be any function that can be applied to calculate the image loss. Exemplarily, the second loss function includes but is not limited to at least one of the following: MAE, GCE, FL, NLNL, RCE, SCE. In some embodiments, the first loss function and the second loss function can be the same or different.
[0207] To extract action information from a sample video, a direct method is to train action embeddings to adapt to the action information, similar to Genie or text inversion. However, the results obtained by this method are not ideal. By examining the attention maps between the action embeddings and the first image at different time steps, it can be found that the embeddings mainly focus on low-frequency action information in the early stage of denoising and gradually shift their attention to high-frequency details in the later stage. Based on this insight, this application proposes Frequency-aware Action Extraction (FAE). Specifically, Figure 9 FIG. shows the training process of FAE. Learnable frequency-aware embeddings are used at each DiT layer to extract information from the sample video. These embeddings are concatenated with the input latent variables along the sequence length. During training, the frequency-aware embeddings are trained separately from the RefAdapter to maintain their consistency. During the training process, the sample video is preprocessed by random cropping to help the action control model adapt to different layouts. This ensures that the target subject can appear in various positions and prevents the embeddings from paying too much attention to the layout information. During this process, following the original I2V training protocol, the first video frame is used as the conditional image.
[0208] The above embodiments give the training process of the frequency perception module. The frequency perception module and the diffusion neural network are trained separately so that the frequency perception module can adapt to different layouts to prevent the first image from paying too much attention to the layout information.
[0209] The following is the device embodiment of this application, which can be used to execute the method embodiment of this application. For details not disclosed in the device embodiment of this application, please refer to the method embodiment of this application.
[0210] Please refer to Figure 10, which shows a block diagram of an action control device provided by an embodiment of the present application. The device has the functions of implementing the above method examples, and the functions can be implemented by hardware or by hardware executing corresponding software. The device can be the computer device introduced above or can be set in the computer device. As Figure 10 shown, the device 1000 includes: a feature extraction module 1010, a feature fusion module 1020, a noise reduction processing module 1030, and a video generation module 1040.
[0211] The feature extraction module 1010 is configured to extract first video features from a first video and first image features from a first image. The first video features are used to characterize the semantic features of a reference subject in the first video, and the first image features are used to characterize the semantic features of a target subject in the first image.
[0212] The feature fusion module 1020 is configured to obtain an initial fusion feature according to the first video features and the first image features. The initial fusion feature is used to fuse the semantic features of the reference subject and the semantic features of the target subject.
[0213] The noise reduction processing module 1030 is configured to perform N-stage noise reduction processing on the initial fusion feature through a diffusion neural network based on frequency-aware weights. The frequency-aware weights are used to control the attention weights of the first image at different frequencies. The attention weights are used to control the attention to the action features of the reference subject during the noise reduction processing. The final fusion feature is used to fuse the action features of the reference subject and the spatial structure features of the target subject, where N is a positive integer.
[0214] The video generation module 1040 is configured to generate a second video including the target subject based on the final fusion feature. The action of the target subject in the second video is controlled based on the action of the reference subject in the first video.
[0215] In some embodiments, the diffusion neural network includes N diffusion sub-networks; the noise reduction processing module 1030 is configured to, for the i-th stage among the N stages, splice the intermediate fusion feature of the (i - 1)-th stage and the frequency-aware weight of the i-th stage to obtain the spliced feature of the i-th stage, and use the spliced feature as the input of the i-th diffusion sub-network among the N sub-networks; where i is a positive integer less than or equal to N, and when i is equal to 1, the intermediate fusion feature of the (i - 1)-th stage is the initial fusion feature; perform noise reduction processing on the spliced feature of the i-th stage through the i-th diffusion sub-network to obtain the intermediate fusion feature of the i-th stage; in the case where i is less than N, let i = i + 1, and start executing again from the step of splicing the intermediate fusion feature of the (i - 1)-th stage and the frequency-aware weight of the i-th stage; in the case where i is equal to N, determine the intermediate fusion feature of the i-th stage as the final fusion feature.
[0216] In some embodiments, the noise reduction processing module 1030 is further configured to obtain an initial attention weight; for the i-th stage among the N stages, in the case where i is less than a first threshold, determine the initial attention weight as the frequency-aware weight of the i-th stage; in the case where i is greater than or equal to the first threshold and less than a second threshold, determine the sum of the initial attention weight and a first weight bias as the frequency-aware weight of the i-th stage; in the case where i is greater than or equal to the second threshold and less than or equal to N, determine the sum of the initial attention weight and a second weight bias as the frequency-aware weight of the i-th stage; where the second weight bias is greater than the first weight bias.
[0217] In some embodiments, the feature extraction module 1010 is further configured to extract a first text feature from first text information, where the first text information is used to describe the action of the reference subject in the first video, and the first text feature is used to characterize the action feature of the reference subject in the first video; the feature fusion module 1020 is configured to splice the first video feature, the first image feature, and the first text feature to obtain the initial fusion feature.
[0218] In some embodiments, the video generation module 1040 is further configured to, in the case where the final fusion feature meets a first condition, perform the step of generating a second video including the target subject based on the final fusion feature; in the case where the final fusion feature does not meet the first condition, use the final fusion feature as the updated initial fusion feature, and start executing again from the step of performing N-stage noise reduction processing on the initial fusion feature through the diffusion neural network based on the frequency-aware weight to obtain the final fusion feature.
[0219] In some embodiments, the first condition includes that the final fused feature is obtained after T times of denoising processing in the N stages, where T is a positive integer.
[0220] In some embodiments, the feature extraction module 1010 is configured to obtain the first video, where the first video includes M video frames and M is an integer greater than 1; randomly select one video frame from the M video frames as the conditional image; replace the first video frame of the first video with the conditional image to obtain M adjusted video frames; extract the features of the conditional image and the M adjusted video frames to obtain the first video feature.
[0221] In some embodiments, the feature extraction module 1010 is configured to extract the features of the conditional image to obtain a first conditional feature, where the first conditional feature is used to characterize the semantic features of the reference subject in the conditional image; splice the first conditional feature and a first parameter to obtain a second conditional feature, where the first parameter is used to smoothly transition the action of the reference subject between the conditional image and the first video; extract the features of the M adjusted video frames to obtain a second video feature, where the second video feature is used to characterize the semantic features of the reference subject in the M adjusted video frames; add a first noise to the second video feature to obtain a third video feature, where the first noise is used to blur the spatial structure features of the reference subject in the second video feature; obtain the first video feature according to the second conditional feature and the third video feature.
[0222] In some embodiments, the feature extraction module 1010 is configured to splice the second conditional feature and the third video feature to obtain a fourth video feature; when the duration corresponding to the fourth video feature is not less than a first duration, determine the fourth video feature as the first video feature, where the first duration is the expected duration of the second video; when the duration corresponding to the fourth video feature is less than the first duration, fill the fourth video feature with a preset parameter to obtain the first video feature.
[0223] In some embodiments, the feature extraction module 1010 is configured to obtain the first image; extract the features of the first image to obtain a second image feature, where the second image feature is used to characterize the semantic features of the first image; add a second noise to the second image feature to obtain the first image feature, where the second noise is used to blur the action features of the target subject in the second image feature.
[0224] The technical solution provided by the embodiment of the present application obtains an initial fusion feature based on the first video feature and the first image feature, and then adjusts the attention weight of the first image at different frequencies based on the frequency-aware weight, so that during the process of the diffusion neural network performing noise reduction processing on the initial fusion feature for N stages, it can shift from focusing on the spatial structure features of the target subject to focusing on the action features of the reference subject at different frequencies, so that the final fusion feature obtained can include more detailed action features, and the appearance consistency of the target subject in the finally generated second video is better, the action is more accurate, the action similarity with the reference subject in the first video is higher, and the flexibility is stronger.
[0225] Please refer to Figure 11 , which shows a block diagram of a training device for an action control model provided by an embodiment of the present application. This device has the functions of implementing the above method examples, and the functions can be implemented by hardware or by hardware executing corresponding software. This device can be the computer device introduced above, or can be set in the computer device. As Figure 11 shown, the device 1100 includes: an extraction module 1110, a first training module 1120, a second training module 1130, and a construction module 1140.
[0226] The extraction module 1110 is configured to extract sample video features from a sample video and extract sample image features from a sample image. The sample video features are used to characterize the semantic features of the sample reference subject in the sample video, and the sample image features are used to characterize the semantic features of the sample target subject in the sample image.
[0227] The first training module 1120 is configured to train a diffusion neural network based on the sample video features to obtain a trained diffusion neural network. The diffusion neural network is configured to perform noise reduction processing on the sample video features for N stages to obtain an output video, and the output video is a denoised video determined based on the sample video.
[0228] The second training module 1130 is configured to train a frequency-aware module based on the sample video features and the sample image features to obtain a trained frequency-aware module. The frequency-aware module is configured to determine a frequency-aware weight, and the frequency-aware weight is used to control the attention weight of the sample image at different frequencies, and the attention weight is used to control the attention to the action features of the sample reference subject during the noise reduction processing.
[0229] A building module 1140 is configured to build an action control model based on the trained diffusion neural network and the trained frequency perception module. The action control model is used to generate a second video including the target subject based on a first image including the target subject and a first video including a reference subject, and the action of the target subject in the second video is controlled based on the action of the reference subject in the first video.
[0230] In some embodiments, the diffusion neural network includes N diffusion sub-networks; the first training module 1120 is configured to, for the i-th stage of the N stages, perform noise reduction processing on the intermediate output features of the (i - 1)-th stage through the i-th diffusion sub-network among the N diffusion sub-networks to obtain the intermediate output features of the i-th stage; where i is a positive integer less than or equal to N. When i equals 1, the intermediate output features of the (i - 1)-th stage are the sample video features. When i equals N, the intermediate output features of the i-th stage are used to generate the output video; adjust the parameters of the diffusion neural network based on the output video and the sample video to obtain the trained diffusion neural network.
[0231] In some embodiments, the diffusion sub-network includes: an attention layer, a first rank adapter, a feed-forward network, and a second rank adapter; the first training module 1120 is configured to adjust the parameters of at least one of the first rank adapter and the second rank adapter based on the output video and the sample video to obtain the trained diffusion neural network; where the attention layer is used to capture the spatial layout features included in the sample video features, the first rank adapter is used to reduce the dimension of the parameters output by the attention layer, the feed-forward network is used to perform a non-linear transformation on the features input to the feed-forward network from the first rank adapter, and the second rank adapter is used to reduce the dimension of the parameters output by the feed-forward network.
[0232] In some embodiments, the diffusion neural network includes N diffusion sub-networks; the second training module 1130 is further configured to extract sample text features from the sample text information, where the sample text information is used to describe the actions of the sample reference subject in the sample video, and the sample text features are used to characterize the action features of the sample reference subject in the sample video; splice the sample video features, the sample image features, and the sample text features to obtain an initial fusion feature; for the i-th stage of the N stages, splice the intermediate output feature of the (i - 1)-th stage and the frequency perception weight of the i-th stage to obtain the splicing feature of the i-th stage, and the splicing feature is used as the input of the i-th diffusion sub-network among the N diffusion sub-networks, and the frequency perception weight of the i-th stage is determined based on the frequency perception module; where i is a positive integer less than or equal to N, and when i is equal to 1, the intermediate fusion feature of the (i - 1)-th stage is the initial fusion feature; perform noise reduction processing on the splicing feature of the i-th stage through the i-th diffusion sub-network to obtain the intermediate fusion feature of the i-th stage; when i is less than N, let i = i + 1, and start to execute the step of splicing the intermediate output feature of the (i - 1)-th stage and the frequency perception weight of the i-th stage again to obtain the splicing feature of the i-th stage; when i is equal to N, determine the intermediate fusion feature of the i-th stage as the final fusion feature, and the final fusion feature is used to fuse the action features of the sample reference subject and the spatial structure features of the sample target subject; based on the final fusion feature and the initial fusion feature, adjust the parameters of the frequency perception module to obtain the trained frequency perception module.
[0233] The technical solution provided in the embodiments of the present application gives a training method for an action control model. Separately training the diffusion neural network and the frequency perception module can ensure the consistency of the trained diffusion neural network and the trained frequency perception module.
[0234] Please refer to Figure 12 , which shows a schematic structural diagram of a computer device provided in an embodiment of the present application. This computer device can be used to implement the action control method or the training method of the action control model provided in the above embodiments. Specifically:
[0235] The computer device 1200 includes a central processing unit (such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array), etc.) 1201, a system memory 1204 including a RAM (Random-Access Memory) 1202 and a ROM (Read-Only Memory) 1203, and a system bus 1205 connecting the system memory 1204 and the central processing unit 1201. The computer device 1200 also includes a basic input / output system (Input Output System, I / O system) 1206 for facilitating the transfer of information between various components within the server, and a mass storage device 1207 for storing an operating system 1210, application programs 1214, and other program modules 1215.
[0236] In some embodiments, the basic input / output system 1206 includes a display 1208 for displaying information and input devices 1209 such as a mouse and a keyboard for user input of information. Among them, both the display 1208 and the input devices 1209 are connected to the central processing unit 1201 through an input / output controller 1210 connected to the system bus 1205. The basic input / output system 1206 may also include an input / output controller 1210 for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1210 also provides output to a display screen, a printer, or other types of output devices.
[0237] The mass storage device 1207 is connected to the central processing unit 1201 through a mass storage controller (not shown) connected to the system bus 1205. The mass storage device 1207 and its associated computer-readable medium provide non-volatile storage for the computer device 1200. That is to say, the mass storage device 1207 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0238] Without loss of generality, the computer-readable medium may include a computer storage medium and a communication medium. The computer storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer storage medium includes RAM, ROM, EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cartridges, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will understand that the computer storage medium is not limited to the above several types. The above-mentioned system memory 1204 and mass storage device 1207 may be collectively referred to as memory.
[0239] According to an embodiment of the present application, the computer device 1200 may also run on a remote computer on the network through a network such as the Internet. That is, the computer device 1200 may be connected to the network 1212 through the network interface unit 1211 connected to the system bus 1205, or in other words, the network interface unit 1211 may also be used to connect to other types of networks or remote computer systems (not shown).
[0240] The memory further includes a computer program, which is stored in the memory and configured to be executed by one or more processors to implement the above-mentioned action control method or the training method of the action control model.
[0241] In an exemplary embodiment, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor of a computer device, the above-mentioned action control method or the training method of the action control model is implemented.
[0242] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or optical discs, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0243] In an exemplary embodiment, a computer program product is further provided. The computer program includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned action control method or the training method of the action control model.
[0244] It should be noted that, before and during the process of collecting the relevant data of the user, this application can display a prompt interface, a pop-up window or output voice prompt information. The prompt interface, pop-up window or voice prompt information is used to prompt the user that their relevant data is being collected currently, so that this application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the confirmation operation issued by the user for the prompt interface or the pop-up window. Otherwise (that is, when the confirmation operation issued by the user for the prompt interface or the pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are ended, that is, the relevant data of the user is not obtained. In other words, all user data collected by this application (including the first video, the first image, etc.) is processed strictly in accordance with the requirements of relevant national laws and regulations. Obtaining the informed consent or separate consent of the personal information subject is carried out under the condition that the user agrees and authorizes, and within the scope authorized by laws and regulations and the personal information subject, subsequent data use and processing behaviors are carried out, and the collection, use and processing of relevant user data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0245] It should be understood that "a plurality of" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0246] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A motion control method, characterized in that: The method comprises: Extracting a first video feature from a first video and extracting a first image feature from a first image, wherein the first video feature is used to characterize a semantic feature of a reference subject in the first video, and the first image feature is used to characterize a semantic feature of a target subject in the first image; Obtaining an initial fusion feature according to the first video feature and the first image feature, wherein the initial fusion feature is used to fuse the semantic feature of the reference subject with the semantic feature of the target subject; Based on the frequency perception weight, performing N stages of denoising processing on the initial fused features through a diffusion neural network to obtain a final fused feature, wherein the frequency perception weight is used to control the attention weight of the first image at different frequencies, and the attention weight is used to control the attention to the action features of the reference subject in the denoising process, and the final fused feature is used to fuse the action features of the reference subject and the spatial structure features of the target subject, and N is a positive integer; Based on the final fusion features, a second video including the target subject is generated, and the action of the target subject in the second video is controlled based on the action of the reference subject in the first video.
2. The method according to claim 1, characterized in that The diffusion neural network includes N diffusion sub-networks; The method of performing N stages of noise reduction processing on the initial fusion features through a diffusion neural network based on the frequency perception weight to obtain the final fusion features includes: For the i-th stage in the N stages, the intermediate fusion features of the i-1th stage and the frequency-aware weights of the i-th stage are concatenated to obtain the concatenated features of the i-th stage, and the concatenated features are used as inputs of the i-th diffusion subnetwork in the N subnetworks; wherein i is a positive integer less than or equal to N, and when i is equal to 1, the intermediate fusion features of the i-1th stage are the initial fusion features; Performing denoising on the splicing features of the i-th stage through the i-th diffusion sub-network to obtain the intermediate fusion features of the i-th stage; When i is less than N, let i=i+1, and start again from the step of concatenating the intermediate fusion features of the i-1th stage and the frequency perception weights of the i-th stage; When i is equal to N, the intermediate fusion feature of the i-th stage is determined as the final fusion feature.
3. The method according to claim 2, characterized in that The method further comprises: Get the initial attention weights; For an i-th stage in the N stages, when i is less than a first threshold, determining the initial attention weight as the frequency perception weight of the i-th stage; When i is greater than or equal to the first threshold and less than the second threshold, the sum of the initial attention weight and the first weight bias is determined as the frequency perception weight of the i-th stage; When i is greater than or equal to the second threshold value and less than or equal to N, the sum of the initial attention weight and the second weight bias is determined as the frequency perception weight of the i-th stage; The second weight bias is greater than the first weight bias.
4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: extracting a first text feature from first text information, where the first text information is used to describe an action of the reference subject in the first video, and the first text feature is used to characterize an action feature of the reference subject in the first video; The obtaining of the initial fusion feature according to the first video feature and the first image feature includes: The first video feature, the first image feature, and the first text feature are concatenated to obtain the initial fusion feature.
5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: In the case where the final fusion feature satisfies the first condition, performing the step of generating a second video containing the target subject based on the final fusion feature; When the final fusion feature does not satisfy the first condition, the final fusion feature is used as the updated initial fusion feature, and the step of performing N stages of noise reduction processing on the initial fusion feature through a diffusion neural network based on frequency perception weights is started again to obtain the final fusion feature.
6. The method according to claim 5, characterized in that The first condition includes: the final fusion feature is obtained after T times of the N stages of denoising processing, where T is a positive integer.
7. The method according to any one of claims 1 to 6, characterized in that: The step of extracting a first video feature from the first video includes: Acquire the first video, where the first video includes M video frames, where M is an integer greater than 1; Randomly select a video frame from the M video frames as a conditional image; Replacing the first video frame of the first video with the conditional image to obtain M adjusted video frames; Features of the conditional image and the adjusted M video frames are extracted to obtain the first video features.
8. The method according to claim 7, characterized in that The extracting the features of the conditional image and the adjusted M video frames to obtain the first video features includes: Extracting a feature of the conditional image to obtain a first conditional feature, where the first conditional feature is used to characterize a semantic feature of the reference subject in the conditional image; splicing the first condition feature and the first parameter to obtain a second condition feature, wherein the first parameter is used to smoothly transition the condition image and the action of the reference subject in the first video; Extracting features of the adjusted M video frames to obtain second video features, where the second video features are used to characterize semantic features of the reference subject in the adjusted M video frames; adding a first noise to the second video feature to obtain a third video feature, wherein the first noise is used to blur a spatial structural feature of the reference subject in the second video feature; The first video feature is obtained according to the second condition feature and the third video feature.
9. The method according to claim 8, characterized in that The obtaining the first video feature according to the second condition feature and the third video feature includes: splicing the second condition feature and the third video feature to obtain a fourth video feature; In a case where the duration corresponding to the fourth video feature is not less than the first duration, determining the fourth video feature as the first video feature, and the first duration is the expected duration of the second video; When the duration corresponding to the fourth video feature is shorter than the first duration, the fourth video feature is filled with preset parameters to obtain the first video feature.
10. The method according to any one of claims 1 to 9, characterized in that: The step of extracting a first image feature from the first image comprises: acquiring the first image; Extracting features of the first image to obtain second image features, where the second image features are used to represent semantic features of the first image; A second noise is added to the second image feature to obtain the first image feature, wherein the second noise is used to blur the action feature of the target subject in the second image feature.
11. A method for training an action control model, characterized in that: The method comprises: Extracting sample video features from the sample video and extracting sample image features from the sample image, wherein the sample video features are used to characterize semantic features of the sample reference subject in the sample video, and the sample image features are used to characterize semantic features of the sample target subject in the sample image; Based on the sample video features, a diffusion neural network is trained to obtain a trained diffusion neural network, wherein the diffusion neural network is used to perform N stages of noise reduction processing on the sample video features to obtain an output video, wherein the output video is a denoised video determined based on the sample video, and N is a positive integer; Based on the sample video features and the sample image features, a frequency perception module is trained to obtain a trained frequency perception module, wherein the frequency perception module is used to determine a frequency perception weight, wherein the frequency perception weight is used to control an attention weight of the sample image at different frequencies, wherein the attention weight is used to control attention to the action features of the sample reference subject in the noise reduction process; Based on the trained diffusion neural network and the trained frequency perception module, a motion control model is constructed, and the motion control model is used to generate a second video containing the target subject based on a first image containing the target subject and a first video containing a reference subject, and the action of the target subject in the second video is controlled based on the action of the reference subject in the first video.
12. The method according to claim 11, characterized in that The diffusion neural network includes N diffusion sub-networks; The step of training a diffusion neural network based on the sample video features to obtain a trained diffusion neural network includes: For the i-th stage in the N stages, the i-th diffusion subnetwork in the N diffusion subnetworks performs noise reduction processing on the intermediate output features of the i-1-th stage to obtain the intermediate output features of the i-th stage; wherein i is a positive integer less than or equal to N, when i is equal to 1, the intermediate output features of the i-1-th stage are the sample video features, and when i is equal to N, the intermediate output features of the i-th stage are used to generate the output video; The parameters of the diffusion neural network are adjusted based on the output video and the sample video to obtain the trained diffusion neural network.
13. The method according to claim 12, characterized in that The diffusion subnetwork includes: an attention layer, a first-rank adapter, a feedforward network, and a second-rank adapter; The step of adjusting the parameters of the diffusion neural network based on the output video and the sample video to obtain the trained diffusion neural network includes: Adjusting parameters of at least one of the first rank adapter and the second rank adapter based on the output video and the sample video to obtain the trained diffusion neural network; Among them, the attention layer is used to capture the spatial layout features included in the sample video features, the first rank adapter is used to reduce the dimension of the parameters output by the attention layer, the feedforward network is used to perform nonlinear transformation on the features input from the first rank adapter to the feedforward network, and the second rank adapter is used to reduce the dimension of the parameters output by the feedforward network.
14. The method according to any one of claims 11 to 13, characterized in that The diffusion neural network includes N diffusion sub-networks; The step of training a frequency perception module based on the sample video features and the sample image features to obtain a trained frequency perception module includes: Extracting sample text features from sample text information, wherein the sample text information is used to describe the action of the sample reference subject in the sample video, and the sample text features are used to characterize the action features of the sample reference subject in the sample video; splicing the sample video features, the sample image features and the sample text features to obtain an initial fusion feature; For the i-th stage in the N stages, the intermediate output features of the i-1th stage and the frequency perception weight of the i-th stage are concatenated to obtain the concatenated features of the i-th stage, and the concatenated features are used as the input of the i-th diffusion subnetwork in the N diffusion subnetworks, and the frequency perception weight of the i-th stage is determined based on the frequency perception module; wherein i is a positive integer less than or equal to N, and when i is equal to 1, the intermediate fusion feature of the i-1th stage is the initial fusion feature; Performing denoising on the splicing features of the i-th stage through the i-th diffusion sub-network to obtain the intermediate fusion features of the i-th stage; When i is less than N, let i=i+1, and start again from the step of concatenating the intermediate output features of the i-1th stage and the frequency perception weights of the i-th stage to obtain the concatenated features of the i-th stage; When i is equal to N, the intermediate fusion feature of the i-th stage is determined as the final fusion feature, and the final fusion feature is used to fuse the action feature of the sample reference subject and the spatial structure feature of the sample target subject; Based on the final fusion feature and the initial fusion feature, the parameters of the frequency perception module are adjusted to obtain the trained frequency perception module.
15. A motion control device, characterized in that: The device comprises: a feature extraction module, configured to extract a first video feature from a first video and a first image feature from a first image, wherein the first video feature is used to characterize a semantic feature of a reference subject in the first video, and the first image feature is used to characterize a semantic feature of a target subject in the first image; A feature fusion module, used for obtaining an initial fusion feature according to the first video feature and the first image feature, wherein the initial fusion feature is used for fusing the semantic feature of the reference subject with the semantic feature of the target subject; a denoising processing module, configured to perform N stages of denoising processing on the initial fused features through a diffusion neural network based on frequency perception weights to obtain final fused features, wherein the frequency perception weights are used to control the attention weights of the first image at different frequencies, and the attention weights are used to control the attention to the action features of the reference subject in the denoising processing, and the final fused features are used to fuse the action features of the reference subject and the spatial structure features of the target subject, and N is a positive integer; A video generation module is used to generate a second video containing the target subject based on the final fusion feature, wherein the action of the target subject in the second video is controlled based on the action of the reference subject in the first video.
16. A training device for a motion control model, characterized in that: The device comprises: An extraction module, used to extract sample video features from the sample video and sample image features from the sample image, wherein the sample video features are used to characterize the semantic features of the sample reference subject in the sample video, and the sample image features are used to characterize the semantic features of the sample target subject in the sample image; A first training module is used to train a diffusion neural network based on the sample video features to obtain a trained diffusion neural network, wherein the diffusion neural network is used to perform N stages of noise reduction processing on the sample video features to obtain an output video, wherein the output video is a denoised video determined based on the sample video; a second training module, configured to train a frequency perception module based on the sample video features and the sample image features to obtain a trained frequency perception module, wherein the frequency perception module is configured to determine a frequency perception weight, wherein the frequency perception weight is configured to control an attention weight of the sample image at different frequencies, wherein the attention weight is configured to control attention to an action feature of the sample reference subject in the noise reduction process; A construction module is used to construct a motion control model based on the trained diffusion neural network and the trained frequency perception module, wherein the motion control model is used to generate a second video containing the target subject based on a first image containing the target subject and a first video containing a reference subject, wherein the action of the target subject in the second video is controlled based on the action of the reference subject in the first video.
17. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the motion control method as described in any one of claims 1 to 10, or to implement the training method of the motion control model as described in any one of claims 11 to 14.
18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the motion control method as described in any one of claims 1 to 10, or to implement the training method of the motion control model as described in any one of claims 11 to 14.
19. A computer program product, characterized in that The computer program product comprises a computer program, which is loaded and executed by a processor to implement the motion control method as described in any one of claims 1 to 10, or the training method of the motion control model as described in any one of claims 11 to 14.
Citation Information
Cited By
Action migration method and device, electronic equipment, computer readable storage medium and program product
CN121354212A
Dance video generation method based on motion focusing attention and decoupling control
CN121531199A
Dance video generation method based on motion focus attention and decoupled control
CN121531199B