Video processing methods, apparatus, computer-readable storage media, and computer program products

CN121509749BActive Publication Date: 2026-08-14XIAMEN MEITUZHIJIA TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

这类方法主要依赖于摄像设备的姿态感应器、取景框提示、实时人脸识别或目标追踪算法等手段,在用户拍摄过程中提供构图建议或自动校正运镜角度,从而帮助拍摄者维持合理的构图结构,其存在的问题包括:很难对已经拍摄好的素材进行智能重构或重跟拍处理,其应用场景存在较大限制

Benefits of technology

[0020]上述视频处理方法、装置、计算机设备、计算机可读存储介质和计算机程序产品,通过获取原始视频;对所述原始视频中的目标对象进行跟踪检测,得到所述目标对象在所述原始视频中的原始运动轨迹;获取目标画面尺寸与目标拍摄风格;根据所述原始运动轨迹、所述目标画面尺寸以及所述原始视频的原始画面尺寸,确定画面扩展参数;根据所述画面扩展参数,对所述原始视频进行画面扩展,得到扩展后视频;将所述原始运动轨迹映射到所述扩展后视频中,得到映射后运动轨迹;根据所述目标拍摄风格以及所述原始运动轨迹,确定所述目标对象在所述扩展后视频中的目标运动轨迹;根据所述目标运动轨迹与所述映射后运动轨迹,确定裁剪框序列;根据所述裁剪框序列,按所述目标画面尺寸对所述扩展后视频进行画面裁剪,得到所述原始视频在所述目标拍摄风格下的目标视频。本发明实施例中基于主体跟踪得到目标对象的原始运动轨迹,作为已生成视频的重构图的基础,根据目标拍摄风格以及目标画面尺寸对该原始运动轨迹进行调整,得到目标对象在目标拍摄风格下在画面中应具有的目标运动轨迹,从而根据该目标运动轨迹对于原始视频进行画面裁剪以及补全,得到重构图后的视频,无需更改原视频帧内容,即可快速实现多种运镜风格的自动重构,既保留了素材细节,也提升了构图美感,特别适用于普通用户在短视频拍摄中存在的构图失衡问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509749B_ABST
    Figure CN121509749B_ABST
Patent Text Reader

Abstract

This application relates to a video processing method, apparatus, computer device, computer-readable storage medium, and computer program product. The method includes: acquiring an original video; tracking and detecting a target object in the original video to obtain the original motion trajectory of the target object in the original video; acquiring the target screen size and target shooting style; determining screen expansion parameters based on the original motion trajectory, target screen size, and the original screen size of the original video; expanding the original video according to the screen expansion parameters to obtain an expanded video; determining a cropping box sequence based on the target motion trajectory and the mapped motion trajectory; and cropping the expanded video according to the target screen size based on the cropping box sequence to obtain a target video of the original video under the target shooting style. This method can improve the efficiency and effectiveness of video reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer data processing technology, and in particular to a video processing method, apparatus, computer-readable storage medium, and computer program product. Background Technology

[0002] Video composition, as an extremely important component of visual art expression, is an extension and expansion based on image composition.

[0003] In related technologies, research and application of video composition primarily focus on the video shooting stage. This involves designing hardware devices or intelligent algorithms to guide users in real-time, ensuring the center of the frame remains aligned with the subject. These methods mainly rely on camera posture sensors, viewfinder cues, real-time face recognition, or target tracking algorithms to provide composition suggestions or automatically correct camera angles during shooting, helping the photographer maintain a reasonable composition. However, these methods have limitations, including difficulty in intelligently reconstructing or reshooting already captured footage, and significant restrictions on their application scenarios. Therefore, a video processing solution capable of reconstructing images from generated video is needed. Summary of the Invention

[0004] Based on this, this application provides a video processing method, apparatus, computer device, computer-readable storage medium, and computer program product, which can improve the efficiency and effectiveness of video reconstruction.

[0005] On one hand, this application provides a video processing method, the method comprising:

[0006] The process involves: acquiring the original video; tracking and detecting the target object in the original video to obtain its original motion trajectory; acquiring the target screen size and target shooting style; determining screen expansion parameters based on the original motion trajectory, the target screen size, and the original screen size of the original video; expanding the original video according to the screen expansion parameters to obtain an expanded video; mapping the original motion trajectory onto the expanded video to obtain a mapped motion trajectory; determining the target motion trajectory of the target object in the expanded video based on the target shooting style and the original motion trajectory; determining a cropping box sequence based on the target motion trajectory and the mapped motion trajectory; and cropping the expanded video according to the target screen size based on the cropping box sequence to obtain the target video of the original video under the target shooting style.

[0007] On the one hand, this application also provides a video processing apparatus, the apparatus comprising:

[0008] The first acquisition module is used to acquire the original video;

[0009] The detection module is used to track and detect target objects in the original video to obtain the original motion trajectory of the target objects in the original video;

[0010] The second acquisition module is used to acquire the target image size and the target shooting style;

[0011] The first determining module is used to determine the image expansion parameters based on the original motion trajectory, the target image size, and the original image size of the original video.

[0012] An expansion module is used to expand the original video according to the image expansion parameters to obtain an expanded video;

[0013] A mapping module is used to map the original motion trajectory onto the expanded video to obtain the mapped motion trajectory;

[0014] The second determining module is used to determine the target motion trajectory of the target object in the expanded video based on the target shooting style and the original motion trajectory.

[0015] The third determining module is used to determine the clipping box sequence based on the target motion trajectory and the mapped motion trajectory;

[0016] The cropping module is used to crop the expanded video according to the target image size based on the cropping frame sequence, so as to obtain the target video of the original video under the target shooting style.

[0017] On the one hand, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps included in any of the aforementioned video processing method embodiments.

[0018] On the one hand, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps included in any of the aforementioned video processing method embodiments.

[0019] On the one hand, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps included in any of the aforementioned video processing method embodiments.

[0020] The aforementioned video processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire an original video; track and detect a target object in the original video to obtain the original motion trajectory of the target object in the original video; acquire the target screen size and target shooting style; determine screen expansion parameters based on the original motion trajectory, the target screen size, and the original screen size of the original video; expand the original video according to the screen expansion parameters to obtain an expanded video; map the original motion trajectory onto the expanded video to obtain a mapped motion trajectory; determine the target motion trajectory of the target object in the expanded video according to the target shooting style and the original motion trajectory; determine a cropping box sequence based on the target motion trajectory and the mapped motion trajectory; and crop the expanded video according to the target screen size based on the cropping box sequence to obtain a target video of the original video under the target shooting style. In this embodiment of the invention, the original motion trajectory of the target object is obtained based on subject tracking and used as the basis for the reconstruction map of the generated video. The original motion trajectory is adjusted according to the target shooting style and the target screen size to obtain the target motion trajectory that the target object should have in the screen under the target shooting style. Then, the original video is cropped and completed according to the target motion trajectory to obtain the video after reconstruction. Without changing the content of the original video frames, it can quickly realize the automatic reconstruction of various camera movement styles, which not only preserves the details of the material, but also improves the composition aesthetics. It is particularly suitable for the composition imbalance problem that ordinary users have in short video shooting. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is an application environment diagram of a video processing method in one embodiment;

[0023] Figure 2 This is a flowchart illustrating a video processing method in one embodiment;

[0024] Figure 3 This is a flowchart illustrating a video processing method in another embodiment;

[0025] Figure 4 This is a schematic diagram illustrating the steps of generating fill content based on a diffusion model in another embodiment;

[0026] Figure 5 This is a schematic diagram of the diffusion model in another embodiment;

[0027] Figure 6 This is a structural block diagram of a video processing device in one embodiment;

[0028] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0029] To make the objectives, technical solutions, and beneficial effects of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0030] In the research and application of post-capture video reshooting technologies, the focus is primarily on the video shooting stage. This involves designing hardware devices or intelligent algorithms to provide real-time guidance to the user, ensuring the center of the frame remains focused on the subject. These methods mainly rely on camera posture sensors, viewfinder prompts, real-time face recognition, or target tracking algorithms to provide composition suggestions or automatically correct camera angles during shooting, helping the photographer maintain a reasonable composition. This type of technology is widely used in mobile phone photography, action cameras, intelligent stabilizers, and professional video equipment, effectively reducing subject shift caused by hand shake, perspective shift, or lack of concentration.

[0031] However, these technical solutions all rely on the premise that the shooting process is still in progress, meaning that intervention or optimization occurs "front-end" in the image generation process. For already shot video footage, especially when the subject is significantly off-center, the composition is unbalanced, or even centered, there is currently a lack of effective "post-processing" solutions. In other words, if the original footage is not stably aligned with the subject during the shooting stage, traditional composition guidance mechanisms become ineffective, and existing systems struggle to intelligently reconstruct or reshoot already captured footage. This constitutes a significant technological gap in practical applications, especially in scenarios where users cannot reshoot or where the footage is of high value and irreplaceable quality, making the need for "post-production restoration" of video composition increasingly prominent.

[0032] The video processing method provided in this application embodiment can be applied to, for example, Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another network server. Server 104 can acquire the original video through terminal 102, which may be displayed on terminal 102; track and detect the target object in the original video to obtain the original motion trajectory of the target object in the original video; acquire the target screen size and target shooting style; determine screen expansion parameters based on the original motion trajectory, the target screen size, and the original screen size of the original video; expand the original video according to the screen expansion parameters to obtain an expanded video; map the original motion trajectory to the expanded video to obtain a mapped motion trajectory; determine the target motion trajectory of the target object in the expanded video according to the target shooting style and the original motion trajectory; determine a cropping box sequence based on the target motion trajectory and the mapped motion trajectory; crop the expanded video according to the target screen size based on the cropping box sequence to obtain the target video of the original video under the target shooting style, and return the target video to terminal 102 for display. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0033] In one exemplary embodiment, such as Figure 2 As shown, a video processing method is provided, which is applied to... Figure 1 Taking server 104 as an example, the following steps are included: Step 202: Obtain the original video.

[0034] The original video can be a video file that a user captures and uploads to the processing system using a smartphone, digital camera, or other device. This video typically contains one or more target objects (i.e., subjects) that need to be highlighted and tracked, such as people, pets, or vehicles. The original video has its inherent original screen dimensions, such as 1080p (1920 pixels × 1080 pixels) or 4K (3840 pixels × 2160 pixels), and its aspect ratio is usually 16:9.

[0035] Step 204: Track and detect the target object in the original video to obtain the original motion trajectory of the target object in the original video.

[0036] The target object can be obtained through interactive segmentation of the original video based on user interaction. Specifically, the system receives interactive prompts from the user on the first frame (or any keyframe) of the original video, such as clicking, selecting, or drawing, explicitly specifying the target object to be tracked. Using a preset video segmentation model (e.g., SAM2, or SegmentAnything Model 2), the video sequence is segmented frame by frame according to the interactive prompts, outputting a binary mask of the target object in each frame. In the binary mask, the pixel value of the target object region is 1 (or 255), and the pixel value of the background region is 0. Then, based on the binary mask of the target object, its minimum bounding box is calculated. This bounding box can be represented by the coordinates (x, y) of its top-left vertex, its width w, and its height h, denoted as the bounding box. The center point coordinates (x + w / 2, y + h / 2) of this bounding box represent the center position of the target object in that frame. The frame-by-frame sequence of bounding boxes can then be used to construct the original motion trajectory. Optionally, considering that the original motion trajectory may exhibit jitter and jumps due to segmentation errors or momentary occlusion, temporal filtering can be applied to the original motion trajectory to obtain a smooth and stable trajectory. Temporal filtering can include outlier removal by calculating the displacement of the bounding box center point between consecutive frames. If the displacement of a frame exceeds a preset threshold (e.g., 50% of the diagonal length of the bounding box in the previous frame), the frame is determined to be an outlier, and linear interpolation is used to repair it based on the trajectories of the valid frames before and after it. Temporal filtering can include motion modeling and filtering, using a Kalman filter to smooth the bounding box sequence after outlier removal. The state vector of the Kalman filter can be defined as [x, y, w, h, v_x, v_y], where v_x and v_y are the velocities of the bounding box center in the x and y directions, respectively. Through prediction and update steps, the filter can effectively suppress noise and predict the trajectory during brief occlusion based on the motion model of the target object, outputting a smoothed bounding box sequence as the smoothed and optimized original motion trajectory. By using temporal smoothing filtering and motion model-based prediction, the jitter and jump problems that may be caused by single-frame detection and segmentation are effectively overcome, ensuring the smoothness and stability of the tracking process in the final generated video and improving the quality of the finished product.

[0037] Step 206: Obtain the target image size and target shooting style.

[0038] The target image size refers to the desired output video image size or aspect ratio, such as 1:1 (square), 9:16 (commonly used for vertical short videos), 21:9 (widescreen cinema), etc. This parameter can be actively selected by the user or automatically set according to the target platform (such as TikTok, Instagram). The target shooting style refers to the virtual camera movement style to be simulated, such as "center tracking," "panning follow," "close-up," "wide pan," etc. Each style corresponds to a set of predefined camera movement rules used to determine how the cropping frame moves relative to the target object. For example, the "center tracking" style requires the center of the cropping frame to always be close to the center of the target object; the "panning follow" style may allow the cropping frame to leave more space in front of the target object's direction of movement to create a more natural composition. By parameterizing the target shooting style, the same original video can quickly generate multiple target videos with different camera movement styles (such as center, pan, and close-up), allowing users to freely switch and compare, greatly improving creative efficiency and user experience.

[0039] Step 208: Determine the image expansion parameters based on the original motion trajectory, the target image size, and the original image size of the original video.

[0040] The core purpose of the image expansion parameters is to determine which directions (top, bottom, left, right) of the original video need to be expanded to meet the cropping requirements of the target image size, and the specific expansion width ratio. Image expansion includes the following analysis dimensions: Safety boundary analysis: Traversing each frame of the smoothed original motion trajectory. For each frame, calculating the minimum distance between the bounding box of the target object and the four boundaries of the original image. Setting a safety margin threshold (e.g., 10% of the target image width or height). If the distance from the target object's bounding box to a certain boundary is less than this threshold, it is determined that image expansion is needed in that direction. Furthermore, for global expansion direction and ratio determination, counting the number of frames determined to need expansion in each direction (top, bottom, left, right) throughout the entire video sequence. If the number of frames needing expansion in a certain direction exceeds a certain percentage (e.g., 20%) of the total number of frames, it is finally determined that expansion is needed in that direction. The expansion width ratio is determined based on the maximum or average expansion amount required for all frames in that direction, and comprehensively considering the ratio between the target image size and the original image size, to ensure that the target object does not get too close to the new boundary of the expanded image during cropping.

[0041] Step 210: Expand the original video according to the image expansion parameters to obtain the expanded video.

[0042] The image expansion is used to generate new visual content in a specific direction, providing sufficient canvas space for subsequent virtual camera movements. To intelligently expand the image boundaries without altering the main content of the original video, providing enough pixel space for subsequent virtual camera cropping, visually coherent and reasonable expanded area content is generated based on the spatiotemporal context information of the original video. Specifically, in this embodiment, content completion is performed based on a generative artificial intelligence model. This process includes two stages: an input / output and problem definition stage. The input of this stage includes the image expansion parameters being specified as a spatial expansion definition. This definition clearly indicates which spatial directions (up, down, left, right) of the original video frame need to be expanded, and the pixel width or proportion to be expanded in each direction. Based on this definition, the system generates a mask for each frame, dividing the image into two regions: a reserved area (corresponding to the original video content) and an expanded area (the area where the model needs to generate content). The output of this stage includes an expanded video that is completely synchronized with the original video sequence. The video's aspect ratio is larger than the original video. The newly added extended area content should maintain visual consistency with the retained area content in terms of texture, lighting, and style, thereby avoiding obvious splicing artifacts, blurring, or logical errors. To fill the extended area with content, this embodiment of the invention employs a video sequence-to-sequence generation model. This model can simultaneously consider spatial and temporal information, ensuring that the generated extended content is not only spatially coherent within a single frame but also temporally stable throughout the entire video sequence (e.g., the movement of clouds and the swaying of leaves generated in the extended area are natural and smooth). The model's workflow can be based on a conditional generation paradigm. It uses a masked sequence of original video frames as conditional input and learns the mapping from the conditional input to the complete output image through a deep learning network (e.g., an architecture including an encoder and decoder). During the training phase, this network learns prior knowledge of the image structure in a large amount of video data, enabling it to predict the most reasonable extended area content based on the given retained area content during the inference phase.

[0043] Optionally, the generative model architecture for diffused content can be based on a diffusion model or a Transformer-based generative model. These models, through iterative denoising or autoregressive generation, can produce high-quality, high-resolution video content. During the generation process, the model uses its internal attention mechanism or other long-range dependency modeling capabilities to ensure that detailed information from the preserved area is effectively propagated to the extended area, achieving seamless integration. For example, an image / video inpainting generation model could be based on a diffusion Transformer.

[0044] Expanding an original video using a diffusion-based Transformer model can involve the following process: For each frame of the original video, a mask image is created, where the original video content area is the effective content (mask value 0), and the surrounding area to be expanded is the area to be filled (mask value 1). A 3D VAE encoder is used to perform spatiotemporal feature compression on the original video frame sequence to obtain low-dimensional latent features. The latent features, mask, and random Gaussian noise are concatenated and divided into spatiotemporal feature blocks. These feature blocks are input into the DiT backbone network. This network consists of multiple stacked Transformer modules, each containing layer normalization, a multi-head self-attention mechanism, and a feedforward network. It captures global dependencies through spatiotemporal position encoding (such as 3DRoPE) to complete the denoising and content generation process. The generated feature sequence is then processed by an "Unpatchify" operation and a 3D VAE decoder to reconstruct a complete image sequence, i.e., the expanded video. The aspect ratio of this expanded video is larger than the original video, and the newly added content seamlessly integrates with the original image, resulting in a visually natural and reasonable presentation. Therefore, this embodiment of the invention intelligently expands the original video footage, providing ample compositional freedom for subsequent steps and laying the technical foundation for achieving high-quality reshoot effects.

[0045] Step 212: Map the original motion trajectory onto the expanded video to obtain the mapped motion trajectory.

[0046] Because of the expanded image, the coordinate space of the original video changes in the expanded video. For example, if the expansion is Δw pixels to the right, the coordinates of a point at (x, y) in the original video become (x + Δw, y) in the expanded video (assuming the expansion direction is to the right and the origin is at the top left corner). Therefore, the original motion trajectory obtained in step 204 (i.e., the bounding box coordinates of the target object in each frame) needs to be translated according to the expansion direction and width to ensure it is correctly represented in the coordinate system of the expanded video, thus obtaining the mapped motion trajectory.

[0047] Step 214: Determine the target motion trajectory of the target object in the expanded video based on the target shooting style and the original motion trajectory.

[0048] The target motion trajectory describes the desired motion path of the target object in the final output video (usually represented by its center point). This is determined by the target shooting style. For example, if the style is "center tracking," the target motion trajectory is essentially the same as the mapped trajectory obtained in step 212, ensuring the target object remains centered in the frame. If the style is "translation tracking," the system will preset an offset in front of the target object's motion velocity vector, making the target object slightly behind in the frame, creating a more dynamic tracking effect. This offset can be dynamically calculated using a speed-related function.

[0049] Specifically, multiple candidate shooting style templates can be preset, each corresponding to a set of camera movement parameters. Candidate shooting style templates can include: Centered tracking template: simulating the visual effect of the target object being positioned as centrally as possible in each frame; Panning tracking template: simulating the effect of the camera moving in the same direction as the target object, such as leaving more space in front of the target object's direction of movement to create a "look-ahead" composition; Gradual zoom-in template: simulating the effect of the camera gradually approaching the target object, with the cropping frame size potentially changing over time; Gradual zoom-out template: the opposite of the zoom-in template, simulating the effect of the camera gradually moving away from the target object. Based on the user's selection and specification of the shooting style template, the camera movement template corresponding to the target shooting style is obtained. A time-series analysis is performed on the original motion trajectory to extract a set of motion state parameters describing the motion characteristics of the target object. These parameters can include: instantaneous velocity vector (describing the speed and direction of motion) and instantaneous acceleration vector (describing the trend of motion change). By setting thresholds for these parameters, the motion state can be classified (e.g., stationary, uniform motion, accelerated motion, etc.).

[0050] Each target shooting style corresponds to a set of trajectory transformation functions. The input to this function is the aforementioned motion state parameters, and the output is a trajectory adjustment vector (e.g., an offset used to calibrate the target position). The style rules are specifically reflected in the definition of this function. For example, for the "translation following" style, its transformation function might be defined as: offset vector = -k * normalized velocity vector (where k is the style intensity coefficient), thus achieving lead space reservation. For the "center tracking" style, its transformation function might be a zero vector, i.e., no additional offset is performed. Trajectory synthesis: Each video frame is traversed, and based on the motion state parameters of the target object in that frame, the trajectory adjustment vector for that frame is calculated using the trajectory transformation function corresponding to the current style. The target position on the original motion trajectory of that frame is added to the calculated adjustment vector to obtain the new position on the target motion trajectory of that frame. The sequence of new positions from all frames is concatenated to generate the preliminary target motion trajectory.

[0051] Optionally, classic composition rules can be introduced to fine-tune the target's motion trajectory, making it more visually aesthetically pleasing. For example, based on the rule of thirds, the expanded image can be divided into thirds both horizontally and vertically, forming a tic-tac-toe grid. The target trajectory can then be fine-tuned so that the center point or key parts of the target object (such as the human eye) are as close as possible to the intersections of these grid lines. Alternatively, more space can be reserved in the line of sight: if the target object is a person or animal and its line of sight is clearly defined, more space can be reserved in that direction. In this case, the offset direction of the target trajectory should be determined by both the motion speed and the line of sight, with priority given to the line of sight. Optionally, the original target trajectory point sequence calculated by the above rules can be smoothed and filtered (e.g., using a Kalman filter or a low-pass filter) to eliminate minor jitter caused by speed estimation noise or rule application, ensuring that the final generated virtual camera motion path is smooth, natural, and consistent with the motion characteristics of a real camera.

[0052] Step 216: Determine the clipping box sequence based on the target motion trajectory and the mapped motion trajectory.

[0053] The cropping frame is a rectangular window with the target image size. The position of this cropping frame is determined for each frame of the expanded video, ensuring that cropping according to this sequence simulates the camera movement of the target shooting style. The principle for determining the cropping frame sequence includes placing the center point of the cropping frame at the position defined by the target motion trajectory. Optionally, to ensure smooth and fluid cropping frame movement and avoid abrupt jumps, the cropping frame center point sequence determined by the target motion trajectory can be further low-pass filtered (e.g., Gaussian filtering) or Kalman filtering to generate a final smooth cropping frame center point path. For each frame, a cropping frame is generated centered on the smoothed center point, according to the target image size. The cropping frames of all frames constitute the cropping frame sequence.

[0054] Step 218: According to the cropping frame sequence, crop the expanded video according to the target image size to obtain the target video of the original video under the target shooting style.

[0055] The process involves iterating through each frame of the expanded video and cropping the image using the corresponding cropping frame. All the cropped frames are then reassembled into a new video sequence at the original video's frame rate, resulting in the final target video. This target video has the user's desired screen size and camera movement style, with a prominent subject and aesthetically pleasing composition.

[0056] This invention overcomes the limitation of related technologies that can only guide composition during the shooting stage. It can effectively repair and optimize videos that have already been shot but have compositional defects (such as subject misalignment), filling a technological gap. Furthermore, compared to solutions like TrajectoryCrafter and Vid2Avatar that require complete video reconstruction or complex 3D scene generation, this invention adopts a lightweight "track-expand-crop" path. Its core lies in the intelligent utilization and local repair of existing footage, rather than full-frame generation, greatly reducing computational complexity and processing time, making it easy to achieve real-time or near real-time processing on mobile devices or in the cloud.

[0057] In one embodiment, the image expansion parameters include the image expansion direction and the image expansion amount; the original video includes multiple original video frames; the process of determining the image expansion parameters includes:

[0058] For each of the original video frames, the target image size is compared with the original image size of the original video frame to obtain the size difference between the target image size and the original image size;

[0059] If it is determined that the original video frame needs to be expanded based on the size difference, the distances between the target object and the screen boundaries of the original video frame in multiple directions are determined based on the original motion trajectory.

[0060] The image expansion direction of the original video frame is determined from the plurality of directions based on the minimum value of the distance;

[0061] Based on the first comparison result between the size difference and the original image size, the image expansion amount in the image expansion direction of the original video frame is determined;

[0062] The image expansion direction and image expansion amount of all the original video frames are integrated to obtain the target expansion direction and the target expansion amount of the original video.

[0063] To ensure that the expansion decision considers both the real-time compositional safety requirements of each frame and generates a consistent expansion strategy applicable to the entire video, thereby guaranteeing the visual coherence of the final video, the image expansion parameters include the image expansion direction and the image expansion amount. The original video includes multiple original video frames. The process of determining the image expansion parameters includes the following: For each original video frame, to determine its corresponding single-frame expansion parameters, the following analysis can be performed: 1. Size comparison: Compare the target image size (e.g., aspect ratio of 9:16) with the original image size of the original video frame (e.g., aspect ratio of 16:9), and calculate the size difference between the two. This difference is mainly reflected in the difference in aspect ratio. For example, if the target is a portrait screen and the original is a landscape screen, then expansion is needed in the vertical direction (up and down) or the horizontal direction (left and right) to adapt to the target size. 2. Expansion necessity determination: Determine whether image expansion is needed for the current original video frame based on the size difference. If the aspect ratio of the target image size is different from that of the original image size, then expansion is required. Even with the same aspect ratio, if subsequent steps reveal insufficient safety margin in the composition, expansion may be necessary. 3. Safety Distance Calculation: If expansion is deemed necessary, the distances (d_{top}, d_{bottom}, d_{left}, d_{right}) between the target object and the frame boundaries in the four directions (top, bottom, left, right) of the original video frame are calculated based on the original motion trajectory (i.e., the bounding box of the target object in the current frame). This distance is typically defined as the shortest pixel distance from the outer edge of the target object's bounding box to the corresponding frame boundary.

[0064] In determining which frames need to be expanded, the expansion direction of a single frame must first be determined: based on the four distances (d_{top}, d_{bottom}, d_{left}, d_{right}) calculated in the previous steps, the minimum value (d_{min}) is found. The direction indicated by this minimum value is the position of the target object closest to the frame boundary, and therefore the direction that most needs to be expanded to reserve composition space. Based on this, the expansion direction of the original video frame is determined. For example, if (d_{left}) is the smallest, then expansion to the left is initially determined. Then, based on the first comparison result between the size difference and the original frame size, the amount of frame expansion in the expansion direction of that frame is determined. The first comparison result can be understood as the minimum expansion amount required to meet the target aspect ratio. Specifically, this includes: firstly, based on the target aspect ratio and the original frame size, calculating the pixel width or height ((Delta W) or (Delta H)) that needs to be increased in a specific direction to achieve the target ratio. Secondly, to ensure compositional safety, the expansion amount must be large enough that the distance between the expanded target object and the new boundary is greater than a preset safety threshold (e.g., 5% of the target image height). Therefore, the final expansion amount of this frame is the larger of the calculated value (based on the size difference) and the value based on the safety distance requirement.

[0065] Finally, the preliminarily determined image expansion directions and image expansion amounts for all the original video frames are integrated and analyzed to obtain a unified target expansion direction and target expansion amount applicable to the entire original video. The process of determining the global expansion parameters may include: direction integration, which involves statistically analyzing the frequency with which each direction (up, down, left, right) is determined to require expansion in all frames. A voting or threshold determination mechanism is used: if a direction is required in more than a certain proportion (e.g., 20%) of the frames, then that direction is included in the final target expansion direction set (D^*) (e.g., (D^* = {Left, Top})). This avoids global expansion for a few special cases, improving processing efficiency. If there are relative directions (e.g., both upward and downward expansion are required), the "nearest edge priority" principle can be used to retain only the one with more urgent needs, or the "symmetric compensation" principle can be used to merge the bidirectional expansion amounts to generate a more balanced image. Delta expansion integration: For each direction in the final set of target expansion directions (D^*), iterate through all video frames and take the maximum value of the expansion amount across all frames in that direction as the target expansion amount for that direction. That is, (Delta_{final} = max(Delta_1, Delta_2, ..., Delta_N)). This ensures that during the entire video playback process, the target object will not be too close to the edge of the frame in any frame due to insufficient expansion amount, thus guaranteeing global compositional safety.

[0066] This invention accurately captures the spatial relationship between the target object and the frame boundary at every moment in the video through frame-by-frame analysis, ensuring that expansion decisions are based on precise data. Furthermore, through a global integration mechanism, potentially inconsistent frame-by-frame expansion requirements are merged into a stable and unified global expansion strategy. This method effectively balances the two major requirements of local composition optimization and global visual coherence, avoiding image jitter or flickering caused by frequent changes in the expansion strategy between frames. Simultaneously, the principle of determining the expansion amount based on the maximum value provides the most reliable compositional safety guarantee for the entire video sequence, forming a crucial foundation for generating high-quality, professional-grade reshoot videos.

[0067] In one embodiment, the process of determining the expanded video includes:

[0068] The extended area corresponding to the original video is determined based on the image extension parameters;

[0069] For each original video frame in the original video, a preset generation model is used to generate content based on the current video frame and the historical video frames preceding the current video frame, to obtain the filling video corresponding to the extended area.

[0070] The filled video is combined with the original video to obtain the expanded video.

[0071] To ensure the naturalness and credibility of the generated filled video, in this embodiment of the invention, when generating the expanded content of any frame, only the information of the current frame and previous frames is relied upon, without using future frames. Specifically, the process of determining the expanded video includes the following sub-steps:

[0072] The expansion region corresponding to the original video is determined based on the image expansion parameters, aiming to clarify the area where content needs to be generated. Based on the image expansion direction (such as one or more of top, bottom, left, and right) and the corresponding expansion amount in the expansion direction determined in step 208, an expansion region template is defined for the entire video sequence. For each frame, this template identifies which pixel areas belong to the expansion region that needs to be filled and which belong to the original content area that needs to be retained.

[0073] For each original video frame in the original video, a preset generation model is used to generate content based on the current video frame and the historical video frames preceding the current video frame, resulting in the filled video corresponding to the extended region.

[0074] Considering that traditional video generation models may use information from future frames such as t+1 and t+2 when processing frame t, but in actual video content generation scenarios, such as background expansion in real-time video calls and real-time composition correction in live streaming, future frames do not exist when the current frame data is generated, the content of the current frame does not depend on future frames. Therefore, to ensure the temporal causality of the filling content generation, this embodiment of the invention employs a causal generation model to generate filling video frame by frame. Specifically, when processing frame t, the inputs of the generation model include:

[0075] Current video frame I_t: the t-th frame of the original video. Historical video frames {I_1, I_2, ..., I_{t-1}}: all or some frames before the current frame (e.g., frames within a sliding window of length N). It should be noted that the input to the generative model in this embodiment does not include any future frames I_{t+1}, I_{t+2}, .... Extended region mask M_t of the current frame: identifies the region in the t-th frame that needs to be filled. Optionally, the generative model can employ an architecture combining causal 3D VAE and causal Transformer. In the causal 3D VAE, the encoder and decoder employ Temporally Causal Convolution in the temporal dimension. Specifically, in the temporal convolution operation, all padding is applied only at the beginning of the sequence, ensuring that the computation of the convolution kernel at any time t depends only on the input at the current time t and previous times, and cannot "see" future information.

[0076] In the DiT backbone network, the causal Transformer's self-attention layer can be configured as a causal attention mask. This means that when calculating the attention of the token at position t, it can only focus on the tokens from position 1 to position t in the sequence, and cannot focus on tokens from position t+1 onwards. This also ensures the temporal causality of the generation process. For frame t, the generation model uses the current frame I_t, historical frame information, and the mask M_t to predict the content that should appear in the expanded region of the frame. The content generated by the generation model is based on the best estimate of the historical context and understanding of the current frame, which is consistent with the process of real video recording, thus ensuring the authenticity and naturalness of the content filled in the image expansion. Through frame-by-frame processing, the filled video V_fill for all expanded regions of the entire video is finally obtained. Finally, the filled video is combined with the original video to obtain the expanded video, which may include: for each frame: using the t-th frame I_t of the original video V_orig as the background. The pixels in the t-th frame of the filled video V_fill corresponding to the expanded region are superimposed or replaced at the corresponding positions of I_t. Specifically, M_t can be used as a mask: I_{exp,t} = I_t * (1 - M_t) + I_{fill,t} *M_t, where I_{exp,t} is the t-th frame after combination, and I_{fill,t} is the t-th frame of V_fill. Finally, arranging all the combined frames in order yields the complete expanded video V_exp.

[0077] In one embodiment, the generation model includes a diffusion model; the process of determining the filled video includes:

[0078] The original video frame sequence is subjected to spatiotemporal joint compression coding to obtain the image latent feature sequence;

[0079] The binary mask sequence identifying the extended region is downsampled to obtain a mask feature sequence; the resolution of the mask feature sequence matches the size of the latent feature sequence of the image.

[0080] The image latent feature sequence, the mask feature sequence, and the noise sequence are fused to form a fused feature sequence;

[0081] The fused feature sequence is input into the diffusion model for denoising to obtain a denoised latent feature sequence; wherein, the diffusion model models the temporal and spatial dependencies of the fused feature sequence based on a self-attention mechanism;

[0082] The latent feature sequence is decoded to obtain the filled video corresponding to the extended region.

[0083] When performing video reconstruction, to achieve lightweight reconstruction rather than full-frame regeneration as in traditional methods, the generation model includes a diffusion model. The process of determining the video filler includes:

[0084] First, the original video frame sequence is subjected to spatiotemporal joint compression coding, compressing the high-dimensional video pixel data into a low-dimensional, semantically rich latent space to significantly reduce the computational complexity of subsequent models. The encoder can be a 3D VAE encoder. This encoder is designed for video data and can simultaneously capture information in both spatial (within a single frame) and temporal (between frames) dimensions. The encoder's processing includes: the original video frame sequence V_orig (shape T × H × W × 3) is input into the 3D VAE encoder. The encoder downsamples in the spatial dimension using its internal 3D convolutional layers and residual network blocks, and uses a causal convolutional structure in the temporal dimension to maintain temporal causality (ensuring that the generation of frame t does not depend on information from future frames). Finally, a low-dimensional image latent feature sequence Z_vid (shape T' × H_l × W_l × C) is output, where T', H_l, and W_l are the compressed temporal, height, and width dimensions, and C is the number of latent channels.

[0085] The binary mask sequence identifying the expanded region is downsampled to obtain a mask feature sequence; wherein the resolution of the mask feature sequence matches the size of the image latent feature sequence. To provide spatial guidance for the content generation process of the diffusion model, a binary mask M is generated for each frame based on the image expansion parameters, where the original content region is 0 (black, reserved area) and the region to be expanded is 1 (white, expanded area), thus providing spatial guidance for the generation process of the filling content corresponding to the expanded region. Specifically, the binary mask sequence M (shape T × H × W × 1) can be downsampled (e.g., using average pooling or nearest neighbor interpolation) to make its spatial resolution the same as H_l and W_l of the aforementioned image latent feature sequence Z_vid, thus obtaining the mask feature sequence Z_mask. This ensures that the mask information can be precisely aligned spatially with the image latent features, providing pixel-level guidance for subsequent fusion and generation. Through the precise guidance of the binary mask, the model is explicitly told the region to be filled, achieving precise spatial control of the generation process and ensuring the connection between the expanded content and the original image.

[0086] Further, a set of random Gaussian noise sequences Z_noise with the same size as Z_vid is sampled. During the forward pass of the diffusion model, noise is progressively added to the data; during the backward pass (denoising), the model's learning objective is to recover clean data from the noise. The latent feature sequence Z_vid, the mask feature sequence Z_mask, and the noise sequence Z_noise are concatenated along the channel dimension to form a fused feature sequence Z_fused, which serves as the input to the diffusion model. The diffusion model uses a diffusion Transformer as its backbone network. Before being input into DiT, the fused feature sequence Z_fused is first patchified (partitioned into feature blocks). That is, it is divided into small blocks in the spatiotemporal dimension and flattened, converting it into a series of token sequences to adapt to the Transformer's processing method. This token sequence is then input into the DiT network. DiT consists of multiple identical Transformer modules stacked together. The core of each module is a multi-head self-attention mechanism and a feedforward network. Through the self-attention mechanism, the model enables each image patch to focus on other parts of the entire image, thereby generating globally coordinated extended content. Optionally, by introducing technologies such as 3D rotational position encoding, tokens are given spatiotemporal position information, enabling the model to understand the relationship between frames, thereby ensuring that the generated extended area content (such as flowing clouds and swaying leaves) is temporally coherent and transitions naturally, avoiding flickering and jumps.

[0087] In this embodiment of the invention, the diffusion model gradually removes noise from Z_fused through multiple time steps, ultimately outputting a denoised latent feature sequence Z_denoised. Finally, Z_denoised is unpatchified to restore its original (T', H_l, W_l, C) dimensions. To convert the generated high-quality latent features back to pixel space, a 3D VAE decoder symmetric to the encoder is used, inputting the denoised latent feature sequence Z_denoised into the 3D VAE decoder. The decoder reconstructs the spatiotemporal dimension through 3D deconvolution layers and upsampling operations, ultimately outputting an expanded video V_exp (i.e., a filled video). V_exp has an expanded frame size; the content of its expanded region is generated by the model and seamlessly blends with the original region, resulting in visual realism and temporal consistency.

[0088] This invention utilizes the temporal self-attention mechanism in 3D VAE and DiT to ensure the smoothness and consistency of the generated video content in the temporal dimension when the diffusion model generates the expanded region of each frame, effectively avoiding inter-frame flickering, jitter, or content logic conflicts. Furthermore, the entire process is performed in a low-dimensional latent space, avoiding repeated iterations in ultra-high-dimensional pixel space. Compared to methods such as TrajectoryCrafter, which require full-frame generation in pixel space, this significantly reduces the computational overhead and generation time of video processing in this invention.

[0089] In one embodiment, the process of determining the target motion trajectory includes:

[0090] The original motion trajectory is analyzed to obtain the motion state labels of the target object in each video frame of the original video;

[0091] Based on the target shooting style, determine the trajectory adjustment parameters corresponding to multiple candidate motion states;

[0092] The original motion trajectory is smoothed and offset calibrated based on the trajectory adjustment parameters corresponding to the motion state labels in each video frame of the original video to obtain the target motion trajectory.

[0093] To intelligently and systematically transform abstract shooting styles into concrete trajectory parameters, this invention analyzes the target's motion state and dynamically adjusts the trajectory based on style rules, thereby simulating virtual camera movement effects with different artistic intentions. Specifically, to quantify the dynamic behavior of the target object, a time-series analysis is performed on the smoothed original motion trajectory Traj_smooth, which may include:

[0094] Motion parameter extraction: Calculate the instantaneous motion vector of the target object in each frame. The instantaneous motion vector can include: instantaneous velocity vector v_t = (vx_t, vy_t), which can be obtained by inter-frame difference and smoothing of the center point coordinates (cx_t, cy_t). Its magnitude ||v_t|| represents the motion speed.

[0095] The instantaneous acceleration vector a_t can be obtained by the inter-frame difference of the velocity vector.

[0096] Based on the instantaneous motion vectors and preset motion vector thresholds, a motion state label L_t is assigned to each frame. For example: if ||v_t|| < θ_static, then L_t = "stationary". If θ_static ≤ ||v_t|| < θ_slow, then L_t = "slow movement". If ||v_t|| ≥ θ_slow, then L_t = "fast movement". Further, "uniform motion" and "accelerated / decelerated motion" can be distinguished based on the acceleration magnitude ||a_t||. To achieve the mapping between shooting style and motion trajectory, this embodiment of the invention can preset a style-rule mapping table, where each target shooting style (such as "center tracking" or "panning follow") corresponds to a set of trajectory adjustment functions or parameter sets in this table. The trajectory adjustment parameters collectively define a mapping function F_style(L_t, v_t,...) from "motion state" to "spatial offset". The function takes the motion state of the current frame as input (including the label L_t, velocity vector v_t, etc.) and outputs a trajectory adjustment vector Δ_t = (Δx_t, Δy_t). This adjustment vector specifies how far and in which direction the center of the cropping box should be offset relative to the actual position of the target object (i.e., the mapped motion trajectory) in order to conform to the shooting style.

[0097] Specifically, the trajectory adjustment parameters may include: a base offset (B): a fixed offset value; a velocity gain coefficient (K): a coefficient used to map the velocity magnitude to the offset; and an offset direction vector (D): a unit vector that defines the basic direction of the offset. The process of applying the mapping rule between shooting style and motion state to the entire video sequence may include: first, performing frame-by-frame offset calculation: for each frame, based on its motion state label L_t, selecting the corresponding parameters from the parameter set determined in step S2142, and calculating the trajectory adjustment vector Δ_t for that frame as follows:

[0098] Δ_t = F_style(L_t, v_t, ...);

[0099] Then, trajectory synthesis is performed. The calculated adjustment vector Δ_t is added to the target position on the mapped motion trajectory Traj_mapped(t) for that frame to obtain the new position on the target motion trajectory for that frame, as follows:

[0100] Traj_target(t) = Traj_mapped(t) + Δ_t.

[0101] Optionally, the initially calculated Traj_target sequence can be smoothed using a filter (such as a low-pass filter or a Kalman filter). This smoothing filter is used to offset and calibrate the stylized motion trajectory, ensuring that the final virtual camera path remains smooth, stable, and conforms to the laws of physical motion after applying the stylized offset. This avoids image jitter caused by sudden parameter changes or abrupt changes in motion state. For example, if the target shooting style is a center-tracking style, considering that the core of this style is to ignore motion state and always keep the target in the center, the trajectory adjustment function is: F_center(...) = (0,0). That is, no matter how the target moves, the adjustment vector is always a zero vector. Since Traj_target is almost equal to Traj_mapped, the center of the clipping box is always focused on the target object, thus achieving a stable focus visual effect.

[0102] Correspondingly, if the target shooting style is a panning follow style, the core of this style is to reserve leading space in the direction of movement. Therefore, its corresponding trajectory adjustment function is: F_follow(L_t, v_t) = -K * (v_t / ||v_t||) (when ||v_t|| > 0). Here, K is the velocity gain coefficient, which can be adjusted according to L_t (K is larger for fast movement); (v_t / ||v_t||) is the unit vector in the velocity direction. The negative sign - indicates that the offset direction is opposite to the direction of movement; in this style, D = -normalize(v_t). The target object's position in the frame will be biased towards the rear of the direction of movement, while a dynamic negative space appears in front of it, simulating the classic composition of a photographer shooting from the side.

[0103] This invention establishes a clear chain of "motion state - style parameters - trajectory adjustment," achieving a highly controllable and interpretable automated virtual camera movement method. This transforms the difficult-to-quantify "artistic style" into calculable mathematical rules, enabling the same original video to quickly generate various professional-grade camera movement effects by switching different rule sets. This parameterized method not only greatly improves creative flexibility and efficiency but also, due to the clear rules, generates highly consistent trajectories, avoiding the unstable results that may arise from end-to-end black-box models. Furthermore, through final global smoothing calibration, it ensures that regardless of the style rules applied, the final output video has a broadcast-quality smoothness.

[0104] In one embodiment, the target motion trajectory is used to characterize the desired position of the target object in each frame of the target video; the process of determining the image cropping includes:

[0105] Determine the cropping frame size based on the target image size;

[0106] The trajectory offset vector of the clipping box is determined based on the distance between the mapped motion trajectory and the target motion trajectory.

[0107] The cropping box sequence is determined based on the size of the cropping box and the trajectory offset vector;

[0108] The expanded video is cropped frame by frame using the cropping box sequence to obtain the target video.

[0109] In order to accurately convert the planned target motion trajectory (ideal virtual camera focus path) into a series of executable cropping frames and ultimately generate the target video, thereby ensuring that the intention of virtual camera movement is accurately and smoothly reflected in the final image, in this embodiment of the invention, the cropping frame is a rectangular window whose size is determined by the previously acquired target image size. Specifically, the width W_crop and height H_crop of the cropping frame are equal to the width W_target and height H_target of the target image, respectively. That is: W_crop = W_target; H_crop = H_target. This fixed size ensures that the final output target video has the consistent aspect ratio (e.g., 9:16) expected by the user.

[0110] Considering that there are two key trajectories in the coordinate system of the expanded video: the mapped motion trajectory Traj_mapped, which represents the actual position sequence of the target object, and the target motion trajectory Traj_target, which is the position sequence of the target object expected to appear in the final video, calculated according to the camera movement style. To "move" the target object from its actual position to the desired position, the center of the cropping box needs to be offset accordingly. Specifically, the process of offsetting the center of the cropping box can include calculating the trajectory offset vector. For frame t of the video, the offset vector Offset_t is calculated as follows: Offset_t = Traj_target(t) - Traj_mapped(t); where Traj_target(t) and Traj_mapped(t) represent the center point coordinates (x, y) of the target object on the target motion trajectory and the mapped motion trajectory, respectively, at frame t. This vector Offset_t explicitly indicates the direction and distance that the center of the cropping box needs to move relative to the actual position of the target object in order to achieve the desired camera movement effect. For example, for the "pan follow" style, Offset_t might be a vector pointing in the opposite direction of the target object's movement, thus reserving space in front of it in the image.

[0111] The final cropping box is constructed for each frame and formed into a sequence, which may include the following processes: 1. Crop Box Centering: The center point Center_t of the cropping box in frame t is determined by the following formula: Center_t = Traj_mapped(t) + Offset_t, which is equivalent to Center_t = Traj_target(t). This means that the center of the cropping box is placed at the desired position defined by the target motion trajectory. This is the core step in realizing various camera movement styles. 2. Trajectory Smoothing Optimization: Considering that the Center_t sequence may have slight jitter due to the computational noise of Traj_target, the Center_t sequence can be smoothed, for example, by using a Savitzky-Golay filter or cubic spline interpolation. This ensures that the movement path of the virtual camera is smooth and without abrupt changes, conforming to the mechanical motion characteristics of professional photography. 3. Sequence Generation: Traverse all frames, using the smoothed (or unsmoothed) Center_t as the center and (W_crop, H_crop) as the size, generate the cropping box Crop_t for each frame. The cropping boxes {Crop_1, Crop_2, ..., Crop_T} of all frames constitute the cropping box sequence.

[0112] Iterate through each frame of the expanded video V_exp: For frame t, extract the pixels within the rectangular area defined by Crop_t from V_exp. Reassemble and encode the extracted frames into a new video file according to the original video's frame rate (FPS) and encoding format. This new video is the final target video. In this video, the target object will appear on the screen according to the path planned by Traj_target, thus accurately achieving the selected target shooting style (such as centering, panning, etc.).

[0113] The cropping scheme refined in this embodiment of the invention, by introducing the concept of trajectory offset vector, reveals the transformation logic from the "actual position of the target" to the "desired composition of the image," thereby making the control of camera movement style precise and calculable. By binding the center of the cropping frame to the target's motion trajectory and supplementing it with path smoothing, a highly controllable and smooth professional-grade virtual camera movement effect is achieved. The video processing method of this embodiment of the invention is computationally lightweight, requiring no complex global generation or 3D reconstruction of the video content, thus ensuring processing efficiency.

[0114] In one embodiment, the original video includes multiple original video frames; the process of determining the target object includes determination based on object selection operations; the process of determining the original motion trajectory includes:

[0115] The position of the target object in each of the original video frames is detected to obtain the original position sequence corresponding to the target object;

[0116] For the position changes of adjacent frames in the original position sequence, outlier filtering is performed on the original position sequence to obtain the filtered position sequence;

[0117] The filtered position sequence is smoothed to obtain the original motion trajectory.

[0118] The original video includes multiple original video frames. The process of determining the target object includes determination based on object selection operations. The process of determining the original motion trajectory may specifically include the following steps:

[0119] First, target location sequence detection is based on interactive segmentation: The system receives object selection operations provided by the user on the initial frame of the original video through clicking, box selection, or drawing, specifying the target object to be tracked, such as a person in the scene. Using a video object segmentation model (such as SAM2), the target object is identified and segmented in each frame of the original video according to the object selection operation, and a binary mask of the target object in each frame is output. Finally, based on the binary mask of each frame, the spatial location representation of the target object is calculated, thereby obtaining the original location sequence corresponding to the target object. Specifically, the spatial location representation can be represented by the bounding box of the target object. This bounding box is its smallest bounding rectangle, which can be described by a quadruple (x_t, y_t, w_t, h_t), where (x_t, y_t) are the pixel coordinates of the upper left corner of the bounding box in frame t, w_t is the width, and h_t is the height. The center coordinates of the bounding box (cx_t, cy_t) = (x_t + w_t / 2, y_t + h_t / 2) can be used as the center position of the target object in that frame. Connecting the bounding boxes corresponding to all frames in chronological order forms the original, unprocessed raw position sequence Traj_raw = { (x_t, y_t, w_t, h_t) | t = 1, 2, ..., T}, where T is the total number of frames.

[0120] Considering that video segmentation models may produce incorrect segmentation in some frames due to occlusion, motion blur, or sudden changes in illumination, resulting in outliers (or outliers) that significantly deviate from the true trajectory in the original position sequence. These outliers manifest as unreasonable and drastic jumps in target position between adjacent frames. Therefore, in this embodiment of the invention, the following outlier filtering process can be performed: First, for frame t (t >= 2), calculate the displacement Δd_t = sqrt( (cx_t - cx_{t-1})^2 + (cy_t - cy_{t-1})^2 ) between it and the center point of the target bounding box in frame t-1. Then, set a dynamic outlier threshold θ_outlier. This threshold can be related to the size of the target, for example, set as a certain proportion (e.g., 50%) of the diagonal length of the bounding box in the previous frame (frame t-1), i.e., θ_outlier = k * sqrt(w_{t-1}^2 + h_{t-1}^2), where k is a proportionality coefficient (e.g., 0.5). If Δd_t > θ_outlier, then the position data (x_t, y_t, w_t, h_t) of frame t is determined to be an outlier. Therefore, frames identified as outliers are removed from the sequence. Alternatively, linear interpolation can be used to repair the outlier using the position data of the preceding and following normal frames.

[0121] For example, if frame t is an outlier, the estimated position of frame t is calculated by interpolation using normal data from frames (t-1) and (t+1), and this estimated value replaces the original outlier. This results in a more continuous filtered position sequence, Traj_filtered, with obvious abrupt changes removed. Considering that even after removing obvious outliers, the filtered position sequence may still contain high-frequency jitter due to inherent noise in the segmentation process, using this sequence would lead to an unstable motion path for the final virtual camera. Therefore, in another embodiment, the filtered position sequence can be smoothed to predict target motion trends and suppress noise. Specifically, a two-stage smoothing strategy combining a Kalman filter and an exponential moving average can be used: the state vector definition process in the Kalman filter includes defining the target's state as X_t = [x_t, y_t, w_t, h_t, vx_t, vy_t]^T, where vx_t and vy_t represent the motion velocities in the x and y directions, respectively. This definition uses a constant velocity motion model. In the specific filtering process, the Kalman filter includes a prediction step and an update step. The prediction step involves predicting the state X_t^- of frame t and the error covariance based on the state of frame t-1 using the state transition matrix F. The update step involves comparing the actual observations of frame t (i.e., (x_t, y_t, w_t, h_t) of frame t in Traj_filtered) with the predicted values, calculating the Kalman gain K_t, and then using the gain to update the state estimate X_t and covariance. It should be noted that if a frame is identified as an outlier and removed, there will be no valid observations for that frame in the update step. In this case, the Kalman filter performs a "no update," that is, adopts the result of the prediction step as the final state estimate for that frame. This mechanism allows the filter to still provide reasonable trajectory predictions and maintain trajectory continuity even when the target is briefly occluded or segmentation completely fails.

[0122] A low-pass filter is used to apply a first-order exponential moving average filter to the state sequence output by the Kalman filter (mainly the center point coordinates (cx_t, cy_t) and dimensions (w_t, h_t)). The formula is: x_t = α * cx_t + (1-α) * x_{t-1}, where x_t is the x-coordinate of the smoothed center point, and α is the smoothing factor (usually around 0.2; the smaller the α, the stronger the smoothing effect). The same operation is performed on cy_t, w_t, and h_t. Low-pass filtering effectively reduces any high-frequency jitter that may remain after Kalman filtering, making the trajectory smoother. After the above smoothing process, a stable and smooth original motion trajectory Traj_smooth is finally obtained. The trajectory acquisition method refined in this embodiment of the invention effectively avoids the destructive impact of severe single-frame segmentation errors on the entire trajectory through an outlier filtering mechanism, improving the robustness of the algorithm. Kalman filtering not only smooths the noise but also introduces a kinematic model, enabling reasonable prediction of the trajectory based on historical motion trends when the target is briefly occluded, ensuring temporal continuity. The final low-pass filtering further improves the smoothness of the trajectory. This multi-level processing flow ensures that the final obtained original motion trajectory is stable, reliable, and predictable, laying a crucial foundation for generating smooth and professional virtual tracking effects in subsequent steps. It is the key technical guarantee for whether the embodiments of this invention can achieve practical-grade effects.

[0123] In one exemplary embodiment, such as Figure 3 As shown, the video processing can be broken down into three core processing modules and one decision-making unit. These modules work collaboratively in sequence according to the data flow, ultimately outputting a reconstructed image. The inputs, outputs, and functions of each module are described below:

[0124] 1. The inputs to the interactive video subject segmentation module include:

[0125] The original input video frame sequence may also include user-provided interactive prompts (click, selection box, or drawing) in the first frame to display a specified subject of interest. The output of the interactive video subject segmentation module includes: a binary mask for each subject; and the minimum bounding box (bbox) corresponding to the mask. The interactive video subject segmentation module tracks the user-selected subject throughout the video using the semi-automatic video segmentation algorithm SAM2. For complex occlusion or fast-moving scenes, it adaptively adjusts the mask confidence to ensure a coherent and clean subject region, and generates bboxes in real-time, providing a precise spatial reference for subsequent cropping window planning.

[0126] 2. The input to the main bounding box timing judgment & smoothing module includes: a frame-level bounding box sequence from module 1.

[0127] The output of Module 2 includes: a smoothed subject trajectory (a temporally continuous bounding box sequence) and motion state labels (stationary / slowly moving / rapidly moving), which can be used as a reference for adjusting the camera movement ratio in candidate steps. This module is used to perform outlier removal and Kalman filtering on the original bounding box sequence to eliminate jitter, and to calculate inter-frame displacement, velocity, and acceleration to determine the subject's motion mode. It also provides the decision unit with a more stable and predictable subject center trajectory, avoiding abrupt cuts in the cropping window.

[0128] 3. Decision Unit, used to generate expansion and cropping parameters. Its inputs include: smoothed subject trajectory (obtained from module 2), original frame size and frame margin, and preset target aspect ratio and camera movement style template. Its outputs include: optimal expansion direction (up / down / left / right or combination) and expansion ratio, cropping frame time series position (i.e., virtual camera path), and camera movement ratio parameters (motion speed easing curve).

[0129] This unit is used to evaluate the minimum safety margin of the target object as the moving subject at the distance edge of each frame, determine the direction and width that must be expanded; and automatically allocate screen white space and calculate the most suitable lens movement range while meeting the target aspect ratio; finally, this unit can generate the cropping window trajectory that can drive module 3 to ensure that the subject is always in the visual focus and the rhythm of the picture is natural and smooth.

[0130] The video frame expansion model takes the following inputs: the original video frames, the expansion direction output by the decision unit, the expansion ratio, and the temporal position of the cropping box within the expanded frame. The outputs include: the expanded video frames (wider frames with filled edges) and the final viewpoint sequence cropped along a trajectory, i.e., the reshoot result. This module can seamlessly "widen" the canvas in a specified direction using diffusion-style inpainting / reference completion techniques based on expansion instructions, and dynamically crop proportional windows according to a smooth cropping trajectory, simulating camera movements such as translation, zooming, and panning.

[0131] In such Figure 3 After the video processing workflow shown is completed in a closed loop, this embodiment of the invention can achieve composition correction and camera movement reconstruction of the captured video without destroying the original subject details. The final output video has the advantages of a stable central subject, reasonable white space, smooth motion path, and low computational cost, and can be used for short video publishing or professional post-production.

[0132] The schematic diagram of the DiT screen expansion framework used in generating fill content in this embodiment of the invention can be referred to as follows. Figure 4This framework aims to achieve spatial expansion of video frames, that is, to generate and fill the peripheral areas of the video based on known video content, making the image more complete and coherent. The overall method is mainly implemented through a Diffusion Transformer (DiT)-based model and involves several specific processing stages, including image encoding, masking, latent feature extraction, diffusion model generation, and finally video decoding and reconstruction. Figure 4 As shown, the framework first takes a series of masked image canvas sequences as input, where the central area of ​​the canvas represents the original video content, and the surrounding areas are represented by black to indicate the areas to be expanded. Additionally, another input is a mask sequence for image filling, which defines, in binary form, the effective area of ​​the video content (the center being the original video area, represented by black) and the area to be filled (the outer area, represented by white). This mask sequence explicitly identifies the spatial locations that need to be expanded and filled. Subsequently, the masked image sequence is input into a 3D VAE Encoder for joint spatiotemporal compression encoding. Specifically, 3D causal structure self-attention is used to perform dimensionality reduction compression on the video frames, converting high-dimensional image sequence data into low-dimensional latent spatial feature representations to reduce the computational complexity of subsequent processing while maintaining effective information representation capabilities.

[0133] like Figure 4 As shown, on another parallel processing path, the binary mask sequence undergoes downsampling to ensure that the mask resolution matches the size of the latent feature sequence of the image after 3D VAE encoding, maintaining consistency in both spatial and temporal dimensions and providing precise guidance for subsequent feature fusion. (Continue to refer to...) Figure 4Next, the two processed data streams undergo a "spatiotemporal feature block partitioning" operation, transforming the data into a sequence feature format suitable for the Transformer model. Simultaneously, a set of random Gaussian noise image sequences is introduced into the other input stream to aid in denoising training during the diffusion process, ensuring the diversity and realism of the generated content. These three sets of processed feature sequences (latest image features, mask features, and noise) are concatenated via a Concat operation to form a unified Transformer input sequence. The fused feature sequence is then input into the DiT backbone network, which consists of multiple Transformer modules. Each Transformer module contains a pre-normalized (Pre-Norm) multi-head self-attention structure and a feedforward network. Through self-attention mechanisms and spatiotemporal location encoding (e.g., 3D RoPE), global dependencies between video frames are captured, ensuring the consistency and coherence of temporal and spatial information during the diffusion generation process. Subsequently, the Transformer-processed feature sequence enters the spatiotemporal block decompression module to restore and reassemble the spatiotemporal dimensions. This process uses linear projection and "Unpatchify" operations to restore the sequence of feature data to its initial 3D spatiotemporal feature dimensions, ensuring that the feature identifiers can be accurately interpreted by subsequent decoding modules. Finally, the 3D VAEDecoder decodes the spatiotemporal feature sequence into a generated image sequence, restoring the video to its original spatial resolution. This step utilizes the decoding capabilities of 3D VAE to achieve the final reconstruction of the extended video content, generating a complete video frame sequence.

[0134] The above describes the DiT diffusion architecture used in the generative model of this invention. Below is the VAE / DiT backbone network and the model network construction for dividing feature blocks and decomposing sub-feature blocks: a) 3D Causal VAE; The 3D Causal VAE consists of an encoder, a decoder, and a latent space regularizer. The encoder and decoder employ a symmetrical multi-stage structure. The encoder contains multiple stages, each consisting of stacked ResNet blocks. Some of these blocks perform 3D downsampling, while others only perform 2D downsampling. By alternating between these two types of blocks, the encoder can compress the input video in both spatial and temporal dimensions. The decoder has a similar structure to the encoder but performs 3D upsampling and 2D upsampling operations for video reconstruction. In the temporal dimension, the 3D Causal VAE employs temporally causal convolution. Specifically, it places all padding at the beginning of the convolutional space to ensure that the current prediction is not affected by future information. Meanwhile, since 3D convolutions consume a significant amount of GPU memory when processing long videos, the authors applied context parallelism in the temporal dimension to distribute computation across multiple devices. Temporally causal convolution is a special type of 3D convolution used to maintain causality in the temporal dimension, ensuring that current predictions are not influenced by future information. This is particularly important in video generation tasks, as the model should not access future frames when generating the current frame. Considering that in standard 3D convolutions, the kernel is bidirectional in the temporal dimension, meaning the output at the center is influenced by both past and future time steps, in temporally causal convolution, the kernel is unidirectional in the temporal dimension, only allowing access to information at the current time step and earlier. Specifically, temporally causal convolution employs a special padding method in the temporal dimension. Traditional 3D convolutions typically use "same" padding in the temporal dimension, adding zeros symmetrically at both ends of the input sequence. Temporally causal convolution uses a padding method where (kernel_size - 1) zeros are added at the beginning of the input sequence, but no padding is added at the end. This means that the convolutional kernel can only access information from the current time step and earlier in each time step's computation, and cannot access future information.Taking a convolutional kernel of size 3 in the time dimension as an example, traditional "same" padding adds a zero padding at both ends of the input sequence. Temporally causal convolution, however, adds two zero paddings at the beginning of the input sequence but no padding at the end. This padding method ensures causality in the time dimension.

[0135] Optionally, the DiT backbone network used in the diffusion model can refer to... Figure 5 ,like Figure 5 As shown, the DiT backbone network is composed of multiple Transformer Blocks, such as... Figure 5 As shown, the DiT backbone network consists of the following modules connected in sequence: 1. Input Tokens; 2. First Layer Norm (normalization layer); 3. Self-Attention layer; 4. Second Layer Norm (normalization layer); 5. Cross-Attention layer; 6. Third Layer Norm (normalization layer); 7. Feed-Forward Network (FFN); 8. Output Tokens.

[0136] The input token first passes through a Layer Norm, then sequentially through Self-Attention, Layer Norm, Cross-Attention, and FFN, finally outputting the result. The Self-Attention layer focuses each element of the input token on the entire sequence, capturing long-range dependencies within the input sequence; the Cross-Attention layer focuses the input token on another sequence (usually the encoder's output), capturing relationships between different sequences. Normalization is used to accelerate training and improve model robustness. FFN is typically a two-layer MLP used for non-linear transformations of the token. This alternating stacking of Self-Attention, Cross-Attention, and FFN is widely used in various Transformer models, efficiently modeling internal and external dependencies of the input data. c, patchify and unpatchify: patchify and unpatchify are two operations used to process the video latency of the 3DCausal VAE output. Patchify is used to divide a video latent (of shape T × H × W × C) obtained from 3D Causal VAE encoding into multiple patches, and flatten these patches into a two-dimensional sequence (of shape (T / q · H / p · W / p) × (p · p · q · C)). Here, T represents the number of frames, H and W represent the height and width of each frame, and C represents the number of channels. p and q are the patch sizes in the spatial and temporal dimensions, respectively.

[0137] Specifically, the sequence can be divided into groups of q frames in the time dimension and into patches of p×p in the spatial dimension. Each patch is then flattened and concatenated into a long sequence. If q > 1, the first frame is repeated at the beginning of the video (or image) sequence to ensure consistent processing of images and videos. Unpatchify is used to restore the 2D sequence processed by the Transformer to its original video latent shape, i.e., (T / q · H / p · W / p) × (p · p · q · C) → T × H × W × C. Specifically, the 2D sequence output by the Transformer can be first divided into a 3D tensor of (T / q, H / p, W / p), then the flattened dimensions of each patch are restored to (p, p, q, C), and finally the dimensional order is adjusted to obtain a video latent of (T, H, W, C). In this embodiment of the invention, the patchify operation can convert the video latent into a sequence format that is easily processed by the Transformer, and achieve unified processing of images and videos by repeating frames at the beginning. The unpatchify operation is the inverse process of patchify, used to restore the Transformer output to the original video latent for subsequent decoding and reconstruction. Specifically, patchify and unpatchify are typically implemented efficiently using tensor operations such as reshape and transpose, without explicitly moving data. These two operations together form a bridge between the video latent in CogVideoX and the 3DCausal VAE and the Transformer, enabling the two modules to seamlessly connect and efficiently complete the video encoding, transformation, and decoding process.

[0138] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. Based on the same inventive concept, this application also provides a video processing apparatus for implementing the video processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method. Therefore, the specific limitations in one or more video processing apparatus embodiments provided below can be found in the limitations of the video processing method above, and will not be repeated here.

[0139] In one exemplary embodiment, such as Figure 6 As shown, a video processing apparatus is provided, comprising: a first acquisition module for acquiring an original video; a detection module for tracking and detecting a target object in the original video to obtain the original motion trajectory of the target object in the original video; a second acquisition module for acquiring a target screen size and a target shooting style; a first determination module for determining screen expansion parameters based on the original motion trajectory, the target screen size, and the original screen size of the original video; an expansion module for expanding the original video based on the screen expansion parameters to obtain an expanded video; a mapping module for mapping the original motion trajectory to the expanded video to obtain a mapped motion trajectory; a second determination module for determining the target motion trajectory of the target object in the expanded video based on the target shooting style and the original motion trajectory; a third determination module for determining a cropping box sequence based on the target motion trajectory and the mapped motion trajectory; and a cropping module for cropping the expanded video according to the target screen size based on the cropping box sequence to obtain a target video of the original video under the target shooting style.

[0140] Each module in the aforementioned video processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0141] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a video processing method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0142] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0143] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps included in any of the foregoing video processing method embodiments.

[0144] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps included in any of the foregoing video processing method embodiments.

[0145] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps included in any of the foregoing video processing method embodiments.

[0146] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0147] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0148] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0149] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A video processing method, characterized in that, The method includes: Obtain the original video; The target object in the original video is tracked and detected to obtain the original motion trajectory of the target object in the original video. Obtain the target image size and target shooting style; the target shooting style refers to the desired virtual camera movement method to be simulated. Based on the original motion trajectory, the target image size, and the original image size of the original video, determine the image expansion parameters; Based on the aforementioned image expansion parameters, the original video is expanded to obtain the expanded video. The original motion trajectory is mapped onto the expanded video to obtain the mapped motion trajectory; Based on the target shooting style and the original motion trajectory, determine the target motion trajectory of the target object in the expanded video; Based on the target motion trajectory and the mapped motion trajectory, determine the clipping box sequence; Based on the cropping frame sequence, the expanded video is cropped according to the target image size to obtain the target video of the original video under the target shooting style.

2. The method according to claim 1, characterized in that, The image expansion parameters include the image expansion direction and the image expansion amount; the original video includes multiple original video frames; the process of determining the image expansion parameters includes: For each of the original video frames, the target image size is compared with the original image size of the original video frame to obtain the size difference between the target image size and the original image size; If it is determined that the original video frame needs to be expanded based on the size difference, the distances between the target object and the screen boundaries of the original video frame in multiple directions are determined based on the original motion trajectory. The image expansion direction of the original video frame is determined from the plurality of directions based on the minimum value of the distance; Based on the first comparison result between the size difference and the original image size, the image expansion amount in the image expansion direction of the original video frame is determined; The image expansion direction and image expansion amount of all the original video frames are integrated to obtain the target expansion direction and the target expansion amount of the original video.

3. The method according to claim 1, characterized in that, The process of determining the expanded video includes: The extended area corresponding to the original video is determined based on the image extension parameters; For each original video frame in the original video, a preset generation model is used to generate content based on the current video frame and the historical video frames preceding the current video frame, to obtain the filling video corresponding to the extended area. The filled video is combined with the original video to obtain the expanded video.

4. The method according to claim 3, characterized in that, The generation model includes a diffusion model; the process of determining the filled video includes: The original video frame sequence is subjected to spatiotemporal joint compression coding to obtain the image latent feature sequence; The binary mask sequence identifying the extended region is downsampled to obtain a mask feature sequence; the resolution of the mask feature sequence matches the size of the latent feature sequence of the image. The image latent feature sequence, the mask feature sequence, and the noise sequence are fused to form a fused feature sequence; The fused feature sequence is input into the diffusion model for denoising to obtain a denoised latent feature sequence; wherein, the diffusion model models the temporal and spatial dependencies of the fused feature sequence based on a self-attention mechanism; The latent feature sequence is decoded to obtain the filled video corresponding to the extended region.

5. The method according to claim 1, characterized in that, The process of determining the target's trajectory includes: The original motion trajectory is analyzed to obtain the motion state labels of the target object in each video frame of the original video; Based on the target shooting style, determine the trajectory adjustment parameters corresponding to multiple candidate motion states; The original motion trajectory is smoothed and offset calibrated based on the trajectory adjustment parameters corresponding to the motion state labels in each video frame of the original video to obtain the target motion trajectory.

6. The method according to claim 1, characterized in that, The target motion trajectory is used to characterize the desired position of the target object in each frame of the target video; the process of determining the image cropping includes: Determine the cropping frame size based on the target image size; The trajectory offset vector of the clipping box is determined based on the distance between the mapped motion trajectory and the target motion trajectory. The cropping box sequence is determined based on the size of the cropping box and the trajectory offset vector; The expanded video is cropped frame by frame using the cropping box sequence to obtain the target video.

7. The method according to claim 1, characterized in that, The original video includes multiple original video frames; the process of determining the target object includes determination based on object selection operations; the process of determining the original motion trajectory includes: The position of the target object in each of the original video frames is detected to obtain the original position sequence corresponding to the target object; For the position changes of adjacent frames in the original position sequence, outlier filtering is performed on the original position sequence to obtain the filtered position sequence; The filtered position sequence is smoothed to obtain the original motion trajectory.

8. A video processing apparatus, characterized in that, The device includes: The first acquisition module is used to acquire the original video; The detection module is used to track and detect target objects in the original video to obtain the original motion trajectory of the target objects in the original video; The second acquisition module is used to acquire the target image size and the target shooting style; the target shooting style refers to the camera movement method of the virtual camera that is to be simulated. The first determining module is used to determine the image expansion parameters based on the original motion trajectory, the target image size, and the original image size of the original video. An expansion module is used to expand the original video according to the image expansion parameters to obtain an expanded video; A mapping module is used to map the original motion trajectory onto the expanded video to obtain the mapped motion trajectory; The second determining module is used to determine the target motion trajectory of the target object in the expanded video based on the target shooting style and the original motion trajectory. The third determining module is used to determine the clipping box sequence based on the target motion trajectory and the mapped motion trajectory; The cropping module is used to crop the expanded video according to the target image size based on the cropping frame sequence, so as to obtain the target video of the original video under the target shooting style.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video processing method and device and electronic device

    CN110189378A

  • Video generation method and device and electronic device

    CN112019768A