Mobile terminal movie feeling lens image to video generation method based on diffusion model

CN122845889APending Publication Date: 2026-09-29SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610984498.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0009]综上,现有技术至少存在以下问题:一是大规模图像到视频扩散模型参数量大、采样步数多、推理内存占用高,难以在移动端部署;二是现有轻量化图像动画或视频生成方法难以生成稳定的子弹时间、滑动变焦和慢动作等电影感镜头效果;三是现有模型剪枝方法容易损害视频时序特征和主体细节;四是现有少步数蒸馏方法难以同时保证时间连续性、主体一致性和镜头运动稳定性;五是现有量化方法未充分考虑视频扩散模型不同模块的量化敏感性

Benefits of technology

1、本发明通过构建教师图像到视频扩散模型,并针对子弹时间效果、滑动变焦效果和慢动作效果等不同电影感镜头类型分别配置镜头效果适配模块,能够生成具有对应镜头运动特征的电影感镜头参考视频数据;由此,学生模型在训练过程中不仅能够学习普通图像到视频生成能力,还能够建立目标电影感镜头类型与主体保持、背景透视变化、时间连续性等视频运动特征之间的映射关系,从而提高移动端图像到视频生成结果的镜头表现力和电影感效果稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845889A_ABST
    Figure CN122845889A_ABST
Patent Text Reader

Abstract

This invention relates to the field of image processing and discloses a mobile-device image-to-video generation method based on a diffusion model. The method involves acquiring an input image and the target cinematic shot type, constructing a teacher image-to-video diffusion model, configuring corresponding adaptation modules for different shot types, and generating ordinary image-to-video data and cinematic shot reference video data. A student model is then constructed, and the diffusion Transformer backbone network undergoes structured deep pruning. Supervised fine-tuning, few-step distillation training, and hybrid post-training quantization are performed to obtain a quantized image-to-video diffusion model suitable for mobile deployment. After loading the model on the mobile device, the corresponding adaptation module is activated according to the target shot type, generating and decoding a video sequence using a few-step diffusion denoising method to obtain a target video with bullet time, zoom-in, or slow-motion effects. This invention reduces model size, sampling steps, and memory usage, while improving video temporal continuity, subject consistency, and shot motion stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically, to a method for generating video from cinematic-style images on mobile devices based on a diffusion model. Background Technology

[0002] In recent years, diffusion models have made significant progress in image generation, video generation, and image-to-video generation tasks. Image-to-video generation techniques typically use one or more input images as conditions, generating consecutive video frames through temporal modeling, motion prediction, and progressive denoising, thus giving static images dynamic representation capabilities. With the increasing demand for short videos, mobile image creation, social media content generation, and smart terminal video editing, users want to generate videos with specific cinematic language and cinematic motion effects from just one photo, such as bullet time, zoom-in, and slow motion effects.

[0003] Existing image-to-video diffusion models typically rely on large-scale diffusion Transformer networks, high-resolution latent variable modeling, and multi-step diffusion sampling processes. While these models offer strong video generation quality, they suffer from large parameter counts, high memory consumption, and long inference latency, usually requiring server-side GPU environments and making direct deployment on mobile devices such as smartphones, tablets, or edge computing devices difficult. Especially in video generation tasks, the model not only needs to preserve the appearance of the subject in the input image but also needs to generate continuous, stable video frames that conform to physical motion laws, thus further increasing computational overhead compared to static image generation.

[0004] On the other hand, existing mobile image-to-video generation solutions mostly employ lightweight networks, templated animations, or simple motion transformations, such as local image deformation, parallax simulation, scaling and translation, or frame interpolation. While these methods have low computational overhead, the generated videos typically feature monotonous camera movements, making it difficult to express complex cinematic effects. For bullet-time effects, existing methods often struggle to maintain both subject stability and background perspective changes simultaneously; for zoom-in effects, existing methods fail to accurately simulate the coupling between changes in subject scale and background perspective; and for slow-motion effects, existing methods are prone to issues such as motion discontinuity, subject edge artifacts, or inter-frame jitter.

[0005] Furthermore, existing image-to-video diffusion models typically require overall fine-tuning of the large model when adapting to different cinematic shot types, resulting in high training costs and inflexible deployment. While low-rank adaptive modules can reduce fine-tuning costs to some extent, existing solutions mostly focus on style transfer or general motion control, with fewer establishing unified adaptation mechanisms for specific cinematic shot motion effects such as bullet time, zoom sliding, and slow motion. Different shot types correspond to significantly different motion patterns; without targeted training data generation and conditional mapping mechanisms, the model is prone to problems such as unstable shot motion, subject drift, background distortion, or unclear target effects.

[0006] To reduce the deployment cost of diffusion models, compression methods such as model pruning, knowledge distillation, and quantization have been proposed in existing technologies. However, conventional pruning methods are mostly designed for convolutional networks or static image models. When directly applied to video diffusion Transformers, they can easily disrupt temporal modeling capabilities, leading to decreased inter-frame continuity, degraded subject details, or unstable motion trajectories. In particular, simply deleting a single layer or proportionally cropping the network width may impair the hidden feature dimensions and temporal attention representation capabilities, making it unsuitable for stable modeling of cinematic camera movements.

[0007] Existing few-step distillation methods can reduce the number of diffusion sampling steps, but in image-to-video generation tasks, relying solely on teacher output supervision or ordinary reconstruction loss is often insufficient to guarantee the quality of the generated video in terms of temporal continuity, subject consistency, and camera motion stability. Student models, under low sampling steps, are prone to issues such as screen flickering, changes in character identity, subject distortion, background stretching, or camera motion that does not conform to the target cinematic effect. Therefore, it is necessary to introduce joint optimization constraints more suitable for video generation tasks into the few-step distillation process, enabling the student model to still approximate the video generation distribution of the teacher model even with low sampling steps.

[0008] Meanwhile, while post-training quantization can reduce model storage and inference memory usage, different modules in the video diffusion model have varying sensitivities to quantization errors. For example, activation values ​​in the diffusion denoising process require high numerical stability, attention projection layers and conditional embedding layers affect temporal relationships and conditional control effects, and linear layers in the feedforward network account for a large proportion of parameters. Applying the same quantization precision to all modules may lead to a significant decrease in generation quality or fail to adequately reduce mobile storage and memory overhead. Therefore, how to perform hybrid post-training quantization based on the importance and quantization sensitivity of different modules is also a problem that needs to be addressed in mobile deployment.

[0009] In summary, existing technologies have at least the following problems: First, large-scale image-to-video diffusion models have a large number of parameters, many sampling steps, and high inference memory consumption, making them difficult to deploy on mobile devices; second, existing lightweight image animation or video generation methods struggle to generate stable cinematic effects such as bullet time, zoom-in, and slow motion; third, existing model pruning methods easily damage video temporal features and subject details; fourth, existing low-step distillation methods struggle to simultaneously guarantee temporal continuity, subject consistency, and camera motion stability; and fifth, existing quantization methods do not fully consider the quantization sensitivity of different modules in the video diffusion model.

[0010] Therefore, this application proposes a mobile cinematic shot image to video generation method based on a diffusion model to solve the above problems. Summary of the Invention

[0011] The purpose of this invention is to solve the technical problems mentioned in the background art above, and to provide a method for generating video from cinematic shot images on mobile devices based on a diffusion model.

[0012] The above-mentioned objective of the present invention is achieved as follows: A mobile-based cinematic-style shot image to video generation method based on a diffusion model includes the following steps: S1. Acquire the input image and receive the target cinematic shot type input by the user. The target cinematic shot type includes at least bullet time effect, zoom-in effect and slow motion effect. S2. Construct a teacher image-to-video diffusion model, and configure corresponding low-rank adaptive modules for different target cinematic shot types. Use the teacher image-to-video diffusion model to generate reference video data for the corresponding shot motion effects. S3. Construct a student image-to-video diffusion model, perform structured deep pruning on the diffusion Transformer backbone network in the student model, and compress and replace continuous redundant layer intervals by analyzing the redundancy relationship between Transformer layers to obtain a lightweight student model. S4. Use the image-to-video data generated by the original base model and the cinematic shot reference video data to perform supervised fine-tuning and warm-up on the lightweight student model in order to restore the student model's basic image-to-video generation capability and establish the mapping relationship between the target cinematic shot type and video motion features. S5. Perform a few-step distillation training on the student model that has completed the supervised fine-tuning, so that the student model can generate video using a preset few-step diffusion denoising method. In the distillation process, the teacher model's true score estimation model, online pseudo score estimation model and lightweight discriminator are introduced to jointly optimize the temporal continuity, subject consistency and camera motion stability of the video results generated by the student model. S6. Perform hybrid post-training quantization on the student model that has completed a few distillation steps, and use different quantization accuracies for different network modules to reduce the storage of model parameters and the memory overhead of mobile inference. S7. Load the quantized image into the video diffusion model in the mobile terminal, encode the input image, and load the corresponding low-rank adaptive module according to the target cinematic shot type, and use the target cinematic shot type as a conditional input to participate in the diffusion denoising generation process. S8. Decode and output the generated video sequence to obtain the target video with corresponding cinematic camera motion effects.

[0013] Furthermore, in step S2, the teacher image-to-video diffusion model is constructed using an image-to-video generation model based on the diffusion Transformer architecture. The teacher model loads corresponding low-rank adaptive modules for different cinematic shot types, wherein: The bullet time effect and the sliding zoom effect use a low-rank adaptive module for public lens effects to generate training videos; The slow-motion effect is achieved by constructing a slow-motion low-rank adaptive module to generate training videos after frame interpolation processing of motion videos.

[0014] Furthermore, the structured deep pruning of the diffuse Transformer backbone network in step S3 includes the following steps: S31. Using calibration samples as input to the student model, extract the hidden state features of each Transformer layer; S32. The feature redundancy between different Transformer layers is calculated using residual linear probes and CKA similarity analysis. S33. Identify continuous redundant layer intervals based on inter-layer similarity results; S34. Replace the corresponding continuous redundant layer interval with a trainable alternative Transformer module. S35. The replacement Transformer module is trained by interval distillation through hidden state alignment, so that the output features of the replacement Transformer module approximate the output features of the original continuous layer interval.

[0015] Furthermore, during the structured depth pruning process, the hidden dimension of the diffusion Transformer backbone network remains unchanged, and only the network depth direction is subjected to structured compression to reduce the loss of video temporal features and the degradation of details of the main subject.

[0016] Furthermore, the supervised fine-tuning preheating in step S4 includes the following stages: The first stage uses image-to-video data generated by the original base model to train the lightweight student model to restore its basic generative capabilities. In the second stage, the student model is trained using cinematic shot reference video data generated by the teacher model to adapt to the shot motion distribution. This enables the student model to learn the subject retention characteristics, background perspective change characteristics, and temporal continuity characteristics under corresponding shot motion.

[0017] Furthermore, the few-step distillation training in step S5 is performed using an adversarial distillation method, including the following steps: S51. Initialize the student model that has completed supervised fine-tuning as a few-step diffusion generator; S52. Generate student video results through a few-step diffusion sampling; S53. Generate teacher reference video results under the same input conditions using the teacher model; S54. Use the true score estimation model to estimate the true distribution of the teacher reference video results; S55. Use an online pseudo-fraction estimation model to estimate the pseudo-distribution of student video results; S56. Use a lightweight discriminator to perform adversarial discrimination between student video results and teacher reference video results; S57. Combine distribution matching loss, adversarial loss, and reinforcement learning reward to jointly optimize the few-step diffusion generator.

[0018] Furthermore, the few-step diffusion generator uses a 4-step diffusion denoising method to generate video, thereby reducing the inference time on mobile devices and the computational overhead of diffusion sampling.

[0019] Furthermore, the hybrid post-training quantization in step S6 includes: The activation values ​​of the diffusion model are maintained with 16-bit floating-point precision. FP8 precision quantization is used for the attention projection layer, conditional embedding projection layer, and temporal embedding projection layer. The linear layers of the feedforward network are quantized using 4-bit weights.

[0020] Furthermore, the training samples include portrait image samples and motion scene image samples, wherein: The bullet time effect and the sliding zoom effect use human portrait image samples as input; The slow-motion effect uses scene image samples containing moving targets as input.

[0021] Furthermore, the mobile terminal includes at least a smartphone, a tablet terminal, or an edge computing device, and the quantized image-to-video diffusion model generates a video sequence with a preset number of frames and a preset resolution using a few-step diffusion sampling method.

[0022] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention constructs a teacher image-to-video diffusion model and configures lens effect adaptation modules for different cinematic lens types such as bullet time effect, sliding zoom effect, and slow motion effect. It can generate cinematic lens reference video data with corresponding lens motion characteristics. As a result, the student model can not only learn the ability to generate ordinary images to videos during the training process, but also establish the mapping relationship between the target cinematic lens type and video motion characteristics such as subject preservation, background perspective change, and temporal continuity. This improves the lens expressiveness and cinematic effect stability of the mobile image-to-video generation results.

[0023] 2. This invention performs structured deep pruning on the diffusion Transformer backbone network, identifies continuous redundant layer intervals based on hidden state similarity, and uses trainable alternative Transformer modules for compression and replacement. This reduces the network depth while maintaining the hidden dimension, thereby reducing the model parameter size and computational overhead. It also reduces the loss of video temporal features and degradation of main details caused by directly pruning the network width or deleting single layers. Furthermore, by combining supervised fine-tuning warm-up and low-step distillation training, the lightweight student model can complete video generation with fewer diffusion denoising steps, thereby reducing mobile inference latency and diffusion sampling computation costs.

[0024] 3. This invention introduces a teacher model's true score estimation model, an online pseudo score estimation model, and a lightweight discriminator in the low-step distillation process. The lightweight discriminator's scores on intermediate latent variable features or multi-layer spatiotemporal features are used as reinforcement learning reward signals, which can improve the distribution approximation ability and visual quality of the student model's generated videos with low sampling steps. Furthermore, through hybrid post-training quantization, differentiated quantization accuracy is applied to activation values, attention projection layers, conditional embedding projection layers, temporal embedding projection layers, and feedforward network linear layers. This can reduce the storage footprint of model parameters and the memory overhead of mobile devices while maintaining the generation quality, making the quantized image-to-video diffusion model more suitable for deployment on smartphones, tablets, or edge computing devices. Attached Figure Description

[0025] Figure 1 This is a logic diagram of the mobile terminal cinematic shot image to video generation method based on diffusion model of the present invention; Figure 2 This is a flowchart of the mobile terminal cinematic shot image to video generation method based on the diffusion model of the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the following description is provided in conjunction with embodiments and appendices. Figure 1-2 The present invention will be further described in detail below. It should be understood that the following embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.

[0027] Example 1: This example provides a mobile-device cinematic motion image-to-video generation method based on a diffusion model. This method can be applied to smartphones, tablets, or edge computing devices to generate videos with cinematic motion effects from input images. The cinematic motion effects include at least one of bullet time, zoom-in, and slow-motion effects.

[0028] The method in this embodiment specifically includes the following steps: S1. Acquire the input image and receive the target cinematic shot type input or selected by the user.

[0029] Specifically, the mobile terminal or training device acquires the image input by the user and receives the target cinematic shot type input or selected by the user. The target cinematic shot type is selected from at least one of bullet time effect, zoom-in effect, and slow motion effect.

[0030] In this embodiment, the input images include portrait image samples and / or motion scene image samples. Specifically, the bullet-time effect and the zoom-in effect can use portrait image samples as training or inference input, while the slow-motion effect can use scene image samples containing moving targets as training or inference input. For example, portrait image samples can include single-person portraits, half-body images, standing images, etc.; motion scene image samples can include images of people running, vehicles moving, ball sports, dance movements, or other images containing moving targets.

[0031] S2. Construct a teacher image-to-video diffusion model and configure corresponding lens effect adaptation modules for different target cinematic lens types.

[0032] Specifically, a teacher image-to-video diffusion model is constructed. This teacher image-to-video diffusion model is built using an image-to-video generation model based on the diffusion Transformer architecture, which can generate video sequences based on input images and conditional information.

[0033] For different types of cinematic shots, corresponding shot effect adaptation modules are configured. These modules are preferably low-rank adaptive modules, with different low-rank adaptive modules corresponding to different types of cinematic shots.

[0034] In this embodiment, the low-rank adaptive module corresponding to the bullet time effect is trained based on video samples with varying perspectives (around viewpoint, relatively stable subject, or varying background viewpoint). The low-rank adaptive module corresponding to the zoom effect is trained based on video samples with varying subject scale, background perspective compression, or background perspective stretching. The low-rank adaptive module corresponding to the slow-motion effect is trained based on slow-motion training samples, which are obtained by performing frame interpolation on videos containing moving targets.

[0035] After loading the corresponding low-rank adaptive module, the teacher image-to-video diffusion model generates ordinary image-to-video data and cinematic shot reference video data with corresponding camera motion effects. The ordinary image-to-video data provides supervision for basic image-to-video generation, while the cinematic shot reference video data provides supervision for cinematic shot motion.

[0036] S3. Construct a student image-to-video diffusion model and perform structured deep pruning on the diffusion Transformer backbone network in the student image-to-video diffusion model.

[0037] Specifically, a student image-to-video diffusion model is constructed. The student image-to-video diffusion model can adopt the same or similar diffusion Transformer architecture as the teacher image-to-video diffusion model, but its parameter scale, network depth, or computational cost is smaller than that of the teacher image-to-video diffusion model.

[0038] Structured deep pruning is performed on the diffusion Transformer backbone network in the student image-to-video diffusion model. During pruning, continuous redundant layer intervals are identified based on the similarity of hidden states in different Transformer layers, and trainable replacement Transformer modules are used to compress and replace these continuous redundant layer intervals, resulting in a lightweight student model.

[0039] In this embodiment, step S3 further includes the following steps: S31. Input student images into the video diffusion model using calibration samples and extract the hidden state features of each Transformer layer.

[0040] Specifically, human image samples, motion scene image samples, or samples generated by the teacher model are used as calibration samples and input into the student image into the video diffusion model. The hidden state features output by each Transformer layer in the diffusion Transformer backbone network are recorded.

[0041] S32. Use residual linear probe and / or center kernel alignment similarity analysis methods to calculate the degree of feature redundancy between different Transformer layers.

[0042] Specifically, the predictability between hidden states in adjacent or cross-layer layers can be analyzed using residual linear probes, and the similarity between hidden states in different Transformer layers can be analyzed using center kernel alignment similarity, thereby determining the degree of feature redundancy between different Transformer layers.

[0043] S33. Identify continuous redundant layer intervals based on inter-layer similarity results.

[0044] Specifically, when the similarity of the hidden states between multiple consecutive Transformer layers is higher than a preset threshold, or when the residual linear probe shows that the outputs of multiple consecutive Transformer layers can be predicted well by the hidden states of the previous layer, the multiple consecutive Transformer layers are identified as a continuous redundant layer interval.

[0045] S34. Replace the continuous redundant layer intervals with trainable alternative Transformer modules.

[0046] Specifically, a trainable alternative Transformer block module with the same structure as the original model is used, and this trainable alternative Transformer block module is used to replace the continuous redundant layer intervals.

[0047] S35. Train the trainable alternative Transformer module by interval distillation using the hidden state alignment method.

[0048] Specifically, the output hidden state of the replaced continuous redundant layer interval is used as the supervision target, and interval distillation training is performed on the trainable replacement Transformer module to make the output features of the trainable replacement Transformer module approximate the output features of the replaced continuous redundant layer interval.

[0049] In the above structured deep pruning process, the hidden dimensions of the diffuse Transformer backbone network are kept unchanged, and the Transformer layers are only structured compressed in the network depth direction to reduce the loss of video temporal features and the degradation of main details.

[0050] S4. Supervised fine-tuning of the lightweight student model using ordinary image-to-video data and cinematic shot reference video data.

[0051] Specifically, the lightweight student model is fine-tuned under supervision using ordinary image-to-video data generated by the teacher's image-to-video diffusion model and cinematic shot reference video data. This enables the lightweight student model to restore its basic image-to-video generation capabilities and learn the mapping relationship between the target cinematic shot type and video motion features.

[0052] In this embodiment, the supervised fine-tuning in step S4 includes the following stages: In the first stage, ordinary image-to-video data generated by the teacher's image-to-video diffusion model is used to train the lightweight student model to restore its basic generative capabilities. This stage is used to enable the lightweight student model to learn image content preservation, subject appearance continuation, basic motion prediction, and video frame continuity.

[0053] In the second stage, cinematic-style reference video data generated by the teacher's image-to-video diffusion model is used to train the lightweight student model to adapt to the camera motion distribution. This enables the lightweight student model to learn the subject retention characteristics, background perspective change characteristics, and temporal continuity characteristics under corresponding camera motions. For example, the bullet time effect corresponds to a relatively stable subject with the viewpoint changing around the subject, the zoom effect corresponds to changes in subject scale and background perspective, and the slow motion effect corresponds to a slowed motion trajectory and smooth inter-frame changes.

[0054] S5. Perform a few-step distillation training on the lightweight student model that has completed supervised fine-tuning.

[0055] Specifically, the lightweight student model that has completed supervised fine-tuning is subjected to a few-step distillation training, which enables it to generate videos using a preset few-step diffusion denoising method. During the distillation training process, the temporal continuity, subject consistency and camera motion stability of the student-generated videos are jointly optimized based on the distribution difference between the teacher reference video and the student-generated video, the adversarial discrimination results and the video quality reward.

[0056] In this embodiment, step S5 further includes the following steps: S51. Initialize the lightweight student model that has completed supervised fine-tuning as a few-step diffusion generator.

[0057] Specifically, the lightweight student model that has undergone supervised fine-tuning is used as the initial network parameters of the few-step diffusion generator, enabling it to have basic image-to-video generation capabilities and cinematic shot motion generation capabilities.

[0058] S52. Generate student video results under the same input image and target cinematic shot type conditions using a low-step diffusion generator.

[0059] Specifically, the input image and the target cinematic shot type are input into the few-step diffusion generator, and the student video results are generated using the few-step diffusion sampling method.

[0060] S53. Using the teacher image-to-video diffusion model, generate teacher reference video results under the same input image and target cinematic shot type conditions.

[0061] Specifically, under the same input image and target cinematic shot type as the student video results, a teacher reference video result is generated using a teacher image-to-video diffusion model.

[0062] S54. Use the true score estimation model to estimate the true distribution score of the teacher reference video results.

[0063] Specifically, the teacher reference video results are input into the true score estimation model to obtain the true distribution score, which characterizes how close the teacher reference video results are to the distribution of the true video.

[0064] S55. Use an online pseudo-score estimation model to estimate the pseudo-distribution scores of student video results.

[0065] Specifically, the student video results are input into an online pseudo-score estimation model to obtain pseudo-distribution scores that characterize the distribution features of the student video results.

[0066] S56. Use a lightweight discriminator to perform adversarial discrimination between student video results and teacher reference video results.

[0067] Specifically, the student video results and the teacher reference video results are input into a lightweight discriminator, which determines whether the input video comes from the student model or the teacher model, and outputs the adversarial discrimination result.

[0068] S57. Based on the real score estimation results, pseudo score estimation results, adversarial discrimination results, and reinforcement learning rewards, a joint training loss is constructed to optimize the few-step diffusion generator.

[0069] Specifically, the real score estimation model and the online pseudo-score estimation model provide the real distribution direction and the student distribution direction, respectively, and the difference between the two is used to construct the distribution matching distillation loss. A lightweight discriminator performs adversarial discrimination between student-generated videos and teacher reference videos, and constructs a generator adversarial loss. Furthermore, the scores of the lightweight discriminator on intermediate latent variable features or multi-layer spatiotemporal features are used as reinforcement learning reward signals, and a group relative strategy optimization method is adopted to optimize the few-step diffusion generator, thereby improving the visual quality, temporal continuity, subject consistency, and camera motion stability of the generated videos.

[0070] In this embodiment, the low-step diffusion generator generates video using a 2 to 8-step diffusion denoising method, preferably using a 4-step diffusion denoising method, to reduce mobile inference time and diffusion sampling computation overhead.

[0071] S6. Perform hybrid post-training quantization on the lightweight student model that has completed a few steps of distillation training.

[0072] Specifically, a hybrid post-training quantization is performed on the lightweight student model that has completed a few-step distillation training, and different quantization accuracies are used for different network modules to obtain a quantized image-to-video diffusion model suitable for mobile deployment.

[0073] In this embodiment, the hybrid post-training quantization described above includes: maintaining 16-bit floating-point precision for the activation values ​​of the diffusion model; using FP8 precision quantization for the attention projection layer, conditional embedding projection layer, and temporal embedding projection layer; and using 4-bit weight quantization for the linear layers of the feedforward network. This hybrid post-training quantization reduces the storage footprint of model parameters and the memory overhead for mobile inference.

[0074] S7. Load the quantized image into the video diffusion model in the mobile terminal, encode the input image, and load, fuse, or activate the corresponding lens effect adaptation module according to the target cinematic lens type.

[0075] Specifically, quantized images are loaded into a video diffusion model on smartphones, tablets, or edge computing devices. The mobile terminal encodes the input image to obtain latent variables or conditional features; based on the target cinematic shot type, it loads, fuses, or activates the corresponding low-rank adaptive module, and uses the target cinematic shot type as a conditional input in the few-step diffusion denoising generation process.

[0076] In this embodiment, when the target cinematic shot type is a bullet time effect, the low-rank adaptive module corresponding to the bullet time effect is loaded, merged, or activated; when the target cinematic shot type is a zoom-in effect, the low-rank adaptive module corresponding to the zoom-in effect is loaded, merged, or activated; when the target cinematic shot type is a slow motion effect, the low-rank adaptive module corresponding to the slow motion effect is loaded, merged, or activated.

[0077] S8. Decode and output the generated video latent variable sequence to obtain the target video with corresponding cinematic camera motion effects.

[0078] Specifically, the quantized image-to-video diffusion model uses a low-step diffusion sampling method to generate a video latent variable sequence with a preset number of frames and a preset resolution. Then, a video decoder decodes the video latent variable sequence to output the target video with corresponding cinematic camera motion effects. By reducing the model parameter size, the number of diffusion sampling steps, and the amount of RAM used, this embodiment can achieve low-latency image-to-video generation on mobile devices.

[0079] Example 2: This example, based on Example 1, further provides a specific implementation method with formula constraints. This example corresponds to the illustrated process, including the base model, teacher model, student model, structured pruning, supervised fine-tuning warm-up, few-step distillation, hybrid post-training quantization, and mobile deployment output.

[0080] S1. Acquire the input image and receive the target cinematic shot type input or selected by the user.

[0081] In this embodiment, the input image is denoted as The target cinematic shot type is denoted as ,in Selected from at least one of the following: bullet time effect, zoom-in effect, and slow motion effect. Input image This includes human image samples and / or moving scene image samples. The bullet time effect and the zoom-in effect use human image samples as training or inference input, while the slow-motion effect uses scene image samples containing moving targets as training or inference input.

[0082] S2. Construct a teacher image-to-video diffusion model and configure corresponding lens effect adaptation modules for different target cinematic lens types.

[0083] like Figure 1 As shown on the left, in this embodiment, a basic image-to-video diffusion model is first constructed, and a teacher model is formed on top of this basic model. Teacher Model An image-to-video generation model based on a diffusion Transformer architecture is employed, targeting different target cinematic shot types. Configure low-rank adaptive modules respectively The teacher model loads the low-rank adaptive module. Then, based on the input image Target cinematic shot type and diffused noise Reference video for generating cinematic shots: ; in, Indicates the number of video frames. and These represent the height and width of the video frame, respectively. This indicates the number of video channels. Therefore, the data generated by the teacher model includes both regular image and video data. And cinematic shot reference video data In the diagram, "Generatedvideos" corresponds to ordinary image-to-video data, and "Cinematiceffectvideos" corresponds to cinematic effect reference video data.

[0084] Among them, the low-rank adaptive modules corresponding to the bullet time effect and the sliding zoom effect are trained based on video samples with corresponding lens motion effects; the low-rank adaptive module corresponding to the slow motion effect is trained based on slow motion training samples, and the slow motion training samples are obtained by performing frame interpolation processing on videos containing moving targets.

[0085] S3. Construct a student image-to-video diffusion model and perform structured deep pruning on the diffusion Transformer backbone network in the student image-to-video diffusion model.

[0086] like Figure 1 As shown in the “◯Structured Pruning” section, this embodiment constructs a student image-to-video diffusion model from the base model or teacher model structure. Furthermore, it performs structured deep pruning on its diffused Transformer backbone network.

[0087] Specifically, step S3 includes the following steps S31 to S35.

[0088] S31. Using calibration samples as input to the student image, input the image into the video diffusion model and extract the hidden state features of each Transformer layer. Input the calibration samples into the student model. Extract the first Hidden state features output by each Transformer layer To ensure consistent dimensions in subsequent similarity calculations, the hidden states are expanded into matrices along the sample dimension: ; in, This indicates the number of samples / tokens obtained by merging batch samples, time frames, and spatial tokens. This represents the hidden dimension. Because structured depth pruning preserves the hidden dimension, the hidden states of different Transformer layers all have the same dimension. .

[0089] S32. Using residual linear probes and / or center kernel alignment similarity analysis methods, calculate the degree of feature redundancy between different Transformer layers. For the first... Layer and first The hidden state of the layer is first centralized to obtain... and The center kernel alignment similarity is defined as: ; in, This is a dimensionless similarity value, and its value is used to characterize the redundancy of the two hidden states. Let Frobenius norm be denoted as . Since both the numerator and denominator are combinations of the inner product norms of the hidden states, the formula as a whole is dimensionless. Alternatively, the residual linear probe can be used to calculate the . Hidden state of layer 1 Predictive power of hidden states in layers: ; in, It is a trainable linear probe. This is the normalized prediction error, which is a dimensionless error. higher and When it is lower, it is considered that the first Layer to the first There is significant redundancy between the layers.

[0090] S33. Identify consecutive redundant layer intervals based on inter-layer similarity results. Identify consecutive Transformer layers that meet the redundancy criteria as consecutive redundant layer intervals: ; in, Indicates the first Layer to the first A continuous redundant layer interval consisting of layers. For example, when adjacent layers or across layers are within a continuous layer interval... Greater than the preset similarity threshold, and When the error is less than the preset error threshold, the continuous layer interval is determined as a compressible interval.

[0091] S34. Replace consecutive redundant layer intervals with trainable alternative Transformer modules. Replace continuous redundant layer intervals Let the input hidden state of the consecutive redundant layer interval be... The original continuous redundant layer interval output is The output of the alternative Transformer module is: ; in, and Having the same dimension, they all belong to .

[0092] S35. Train the trainable alternative Transformer module by interval distillation using the hidden state alignment method.

[0093] To ensure that the replacement Transformer module retains the expressive power of the original continuous redundant layer intervals, the interval distillation loss is defined as the normalized hidden state alignment loss: ; in, The loss is dimensionless. By minimizing this loss, the output features of the replacement Transformer module approximate the output features of the replaced continuous redundant layer intervals. The hidden dimensions of the diffusion Transformer backbone network are preserved during pruning. The compression remains unchanged, but only in the network depth direction is structured to reduce the loss of video temporal features and the degradation of subject details.

[0094] S4. Supervised fine-tuning of the lightweight student model is performed using image-to-video data generated from the original base model and cinematic shot reference video data.

[0095] like Figure 1 As shown in "②Supervised Fine-tuning Warm-up", the structured pruned student model is denoted as First, we use ordinary image-to-video data. The system restores the student model's base image to video generation capabilities, and then utilizes cinematic-style shot reference video data. Perform camera motion distribution adaptation training.

[0096] For training samples The output of the lightweight student model is: ; in, and They have the same video tensor dimension. The supervised fine-tuning loss is defined as: ; The first term is the element-wise average video reconstruction error, which has a normalized form. A dimensionless measure of lens motion difference, used to constrain the learning target cinematic lens type in student models. The corresponding features include subject retention, background perspective change, and temporal continuity. and These are the weighting coefficients.

[0097] This step enables the student model not only to recover the ability to generate videos from ordinary images, but also to establish a mapping relationship between the target cinematic shot type and video motion features. After completing the supervised fine-tuning warm-up, the student model parameters are used as the initialization parameters for the subsequent few-step generator, i.e., Init. in the diagram.

[0098] S5. Perform a few-step distillation training on the lightweight student model that has completed supervised fine-tuning. For example... Figure 1 As shown in "③ Few-step distillation", the student model that has completed supervised fine-tuning is initialized as a 4-step generator. The generator produces fake videos (student video results) during the sampling process; the teacher model or reference data provides reference videos (teacher reference video results).

[0099] Specifically, step S5 includes the following steps S51 to S57.

[0100] S51. Initialize the lightweight student model that has completed supervised fine-tuning as a few-step diffusion generator.

[0101] The student model, after supervised fine-tuning, is initialized as a few-step diffusion generator: ; in, This represents a few-step diffusion generator. In this embodiment, A 4-step diffusion denoising method is preferred, but 2 to 8-step diffusion denoising methods can also be used.

[0102] S52. Generate student video results using a few-step diffusion generator under the same input image and target cinematic shot type conditions. The few-step diffusion generator uses the input image... Target cinematic shot type Adapter module and noise Under the given conditions, generate student video results: ; in, This corresponds to the fakevideos shown in the diagram.

[0103] S53. Using the teacher image-to-video diffusion model, generate teacher reference video results under the same input image and target cinematic shot type. The teacher model generates teacher reference video results under the same conditions: ; in, The corresponding reference videos are shown in the diagram. and Having the same dimension, they all belong to S54. Use the true score estimation model to estimate the true distribution score of the teacher reference video results.

[0104] like Figure 1 RealModel in As shown, the true score estimation model yields results based on the teacher reference video. By estimating the true distribution fractions, we obtain: ; in, This represents the true score estimation model. It is a dimensionless fractional vector or scalar used to represent the score of the teacher's reference video on the real video distribution.

[0105] S55. Use an online pseudo-score estimation model to estimate the pseudo-distribution scores of student video results.

[0106] like Figure 1 FakeModel in As shown, the online pseudo-score estimation model evaluates student video results. By performing pseudo-distribution fraction estimation, we obtain: ; in, This represents an online pseudo-score estimation model. and They have the same dimension and are all dimensionless fractions. Based on and Construct a distribution-matched distillation loss, as shown in the diagram. ; in, This represents the dimension of the fraction vector; when the fraction is a scalar This loss is a dimensionless loss, used to constrain the distribution score of student video results to approximate the true distribution score of teacher reference video results.

[0107] S56. Use a lightweight discriminator to perform adversarial discrimination between student video results and teacher reference video results.

[0108] like Figure 1 As shown in the Discriminator, the lightweight discriminator Given a video, an input image, and a shot type condition, output the probability that the video belongs to the reference video distribution.

[0109] The discriminator training loss is: ; The adversarial loss corresponding to the few-step diffusion generator, as shown in the diagram. ,for: ; Since the discriminator output is a probability value, the logarithmic loss mentioned above is a dimensionless loss.

[0110] S57. Based on the true distribution score, pseudo distribution score, adversarial discrimination result and video quality reward, a joint training loss is constructed to optimize the few-step diffusion generator.

[0111] To further constrain temporal continuity, subject consistency, and camera motion stability, a video quality bonus is introduced. (See the diagram below.) This embodiment employs a group-based strategy to optimize the reward structure. A group of student videos is obtained by sampling the same input image and the target cinematic shot type. ; A lightweight discriminator scores the intermediate latent variable features or multi-layered spatiotemporal features of each student's video to obtain a reward value: .

[0112] Advantages of in-group normalization: ; in, The average reward within the group. The standard deviation of the group's rewards. To prevent constants with a denominator of zero. Since both the numerator and denominator are bonus values, This is a dimensionless dominance value.

[0113] The GRPO loss of the generator is defined as: ; in, This indicates that the few-step diffusion generator generates the first step given the input image, the target cinematic shot type, and the adaptation module. The conditional probability of each student video is used. This loss is used to encourage the generator to improve the probability of generated results with higher temporal continuity, subject consistency, and camera motion stability. Therefore, the joint optimization objective of the few-step diffusion generator is: ; in, These are dimensionless weighting coefficients. All the above losses are dimensionless or normalized losses. By minimizing... The low-step diffusion generator can generate videos that simultaneously meet the requirements of distribution matching, visual realism, temporal continuity, subject consistency, and stable camera movement under the condition of 4-step diffusion denoising.

[0114] S6. Perform hybrid post-training quantization on the lightweight student model that has completed a few steps of distillation training.

[0115] like Figure 1As shown in "④ Hybrid Post-Training Quantization", hybrid post-training quantization is performed on the 4-step generator that has completed a few steps of distillation training to obtain a quantized image-to-video diffusion model suitable for mobile terminal deployment. Specifically, the activation values ​​of the diffusion model are maintained with 16-bit floating-point precision; the attention projection layer, conditional embedding projection layer, and temporal embedding projection layer are quantized with FP8 precision; and the linear layers of the feedforward network are quantized with 4-bit weights.

[0116] For the weights of the linear layer of the feedforward network Its 4-bit weighted quantization can be expressed as: ; in, This represents the weights after quantization and reverse calibration. and Having the same dimensions It is a dimensionless quantity. and The lower and upper limits of the range are represented by 4-digit integers. Therefore, Compared with the original weight They have the same dimensions.

[0117] S7. Load the quantized image into the video diffusion model on the mobile terminal, encode the input image, and load, fuse, or activate the corresponding shot effect adaptation module according to the target cinematic shot type. The quantized image is loaded into a video diffusion model on a mobile terminal. The mobile terminal includes smartphones, tablets, or edge computing devices. The mobile terminal processes the input image... Encode the image to obtain latent variables or conditional features: ; in, This refers to the image encoder. It is based on the target cinematic shot type. The mobile terminal loads, integrates, or activates the corresponding low-rank adaptive module. and will , and Together they serve as conditional inputs for the few-step diffusion denoising process.

[0118] S8. Decode and output the generated video latent variable sequence to obtain the target video with corresponding cinematic camera motion effects.

[0119] The quantized image-to-video diffusion model generates a sequence of latent variables for the video using a few-step diffusion sampling method. ; in, This represents a quantized, low-step image-to-video diffusion model. This represents the sequence of latent variables in the video. Subsequently, the video decoder Dec decodes the sequence of latent variables: ; in, This indicates the final output target video. The target video has a cinematic feel and style. The corresponding camera movement effects, such as bullet time effect, zoom-in effect, or slow motion effect.

[0120] Through the above methods, this embodiment can reduce the model parameter size, diffusion sampling steps, and memory usage, and achieve the generation of video sequences with preset frame counts and preset resolutions on mobile devices.

[0121] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating video from cinematic-style shot images on mobile devices based on a diffusion model, characterized in that, The steps include the following: S1. Acquire the input image and receive the target cinematic shot type input by the user. The target cinematic shot type includes at least bullet time effect, zoom-in effect and slow motion effect. S2. Construct a teacher image-to-video diffusion model, and configure corresponding low-rank adaptive modules for different target cinematic shot types. Use the teacher image-to-video diffusion model to generate reference video data for the corresponding shot motion effects. S3. Construct a student image-to-video diffusion model, perform structured deep pruning on the diffusion Transformer backbone network in the student model, and compress and replace continuous redundant layer intervals by analyzing the redundancy relationship between Transformer layers to obtain a lightweight student model. S4. Use the image-to-video data generated by the original base model and the cinematic shot reference video data to perform supervised fine-tuning and warm-up on the lightweight student model in order to restore the student model's basic image-to-video generation capability and establish the mapping relationship between the target cinematic shot type and video motion features. S5. Perform a few-step distillation training on the student model that has completed the supervised fine-tuning, so that the student model can generate video using a preset few-step diffusion denoising method. In the distillation process, the teacher model's true score estimation model, online pseudo score estimation model and lightweight discriminator are introduced to jointly optimize the temporal continuity, subject consistency and camera motion stability of the video results generated by the student model. S6. Perform hybrid post-training quantization on the student model that has completed a few distillation steps, and use different quantization accuracies for different network modules to reduce the storage of model parameters and the memory overhead of mobile inference. S7. Load the quantized image into the video diffusion model in the mobile terminal, encode the input image, and load the corresponding low-rank adaptive module according to the target cinematic shot type, and use the target cinematic shot type as a conditional input to participate in the diffusion denoising generation process. S8. Decode and output the generated video sequence to obtain the target video with corresponding cinematic camera motion effects.

2. The method for generating video from cinematic-style shot images on mobile devices based on a diffusion model according to claim 1, characterized in that, In step S2, the teacher image-to-video diffusion model is constructed using an image-to-video generation model based on the diffusion Transformer architecture. The teacher model loads corresponding low-rank adaptive modules for different cinematic shot types, wherein: The bullet time effect and the sliding zoom effect use a low-rank adaptive module for public lens effects to generate training videos; The slow-motion effect is achieved by constructing a slow-motion low-rank adaptive module to generate training videos after frame interpolation processing of motion videos.

3. The method for generating video from cinematic-style shot images on mobile devices based on a diffusion model according to claim 1, characterized in that, When performing structured deep pruning on the diffuse Transformer backbone network in step S3, the following steps are included: S31. Using calibration samples as input to the student model, extract the hidden state features of each Transformer layer; S32. The feature redundancy between different Transformer layers is calculated using residual linear probes and CKA similarity analysis. S33. Identify continuous redundant layer intervals based on inter-layer similarity results; S34. Replace the corresponding continuous redundant layer interval with a trainable alternative Transformer module. S35. The replacement Transformer module is trained by interval distillation through hidden state alignment, so that the output features of the replacement Transformer module approximate the output features of the original continuous layer interval.

4. The mobile terminal cinematic shot image to video generation method based on diffusion model according to claim 3, characterized in that, The structured depth pruning process maintains the hidden dimension of the diffusion Transformer backbone network unchanged, and only performs structured compression in the network depth direction to reduce the loss of video temporal features and the degradation of details of the main subject.

5. The mobile terminal cinematic shot image to video generation method based on diffusion model according to claim 1, characterized in that, The supervised fine-tuning preheating in step S4 includes the following stages: The first stage uses image-to-video data generated by the original base model to train the lightweight student model to restore its basic generative capabilities. In the second stage, the student model is trained using cinematic shot reference video data generated by the teacher model to adapt to the shot motion distribution. This enables the student model to learn the subject retention characteristics, background perspective change characteristics, and temporal continuity characteristics under corresponding shot motion. 6.The method of claim 1, wherein, The few-step distillation training in step S5 is conducted using an adversarial distillation method, including the following steps: S51. Initialize the student model that has completed supervised fine-tuning as a few-step diffusion generator; S52. Generate student video results through a few-step diffusion sampling; S53. Generate teacher reference video results under the same input conditions using the teacher model; S54. Use the true score estimation model to estimate the true distribution of the teacher reference video results; S55. Use an online pseudo-fraction estimation model to estimate the pseudo-distribution of student video results; S56. Use a lightweight discriminator to perform adversarial discrimination between student video results and teacher reference video results; S57. Combine distribution matching loss, adversarial loss, and reinforcement learning reward to jointly optimize the few-step diffusion generator. 7.The method of claim 6, wherein, The low-step diffusion generator uses a 4-step diffusion denoising method to generate video, thereby reducing the inference time on mobile devices and the computational overhead of diffusion sampling. 8.The method of claim 1, wherein, The hybrid post-training quantization in step S6 includes: The activation values ​​of the diffusion model are maintained with 16-bit floating-point precision. FP8 precision quantization is used for the attention projection layer, conditional embedding projection layer, and temporal embedding projection layer. The linear layers of the feedforward network are quantized using 4-bit weights. 9.The method of claim 1, wherein, The training samples include portrait image samples and moving scene image samples, wherein: The bullet time effect and the sliding zoom effect use human portrait image samples as input; The slow-motion effect uses scene image samples containing moving targets as input.

10. The mobile terminal cinematic shot image to video generation method based on diffusion model according to claim 1, characterized in that, The mobile terminal includes at least a smartphone, a tablet terminal, or an edge computing device, and the quantized image-to-video diffusion model generates a video sequence with a preset number of frames and a preset resolution using a few-step diffusion sampling method.