A lightweight flag flying video generation method and system
By analyzing and segmenting a single static image to obtain an optical flow vector and generating a motion candidate field that conforms to the motion law of the target, and combining it with user control signals for weighted synthesis, the problem of stable deployment of image-to-video generation on resource-constrained devices is solved. This enables controllable dynamics of key elements such as flags, meeting the needs of popular science display and task simulation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH OF CHINA
- Filing Date
- 2026-04-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing image-to-video generation methods are difficult to deploy stably on resource-constrained terminal or payload devices, and the generation process is complex, making it difficult to meet the requirements of low-power, interpretable, and verifiable science popularization exhibitions and task simulation scenarios.
A lightweight method is adopted to obtain optical flow vectors by parsing and segmenting the target region of a single static image, generating motion candidate fields that conform to the motion law of the target, and using user control signals for weighted synthesis to generate continuous frames and output them as target video.
Under low computing resource conditions, it achieves controllable dynamics of key elements such as flags, providing intuitive explanations of physical phenomena and interactive visual feedback, and is suitable for resource-constrained science popularization and task simulation scenarios.
Smart Images

Figure CN122120577A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision and image processing technology, specifically to a lightweight method and system for generating videos of flags waving. Background Technology
[0002] In aerospace science popularization, museum exhibitions, teaching demonstrations, and virtual simulations, it is often necessary to animate key elements (such as flags, ribbons, banners, etc.) in historical images or single-frame photographs in a visually appealing way to enhance the audience's understanding and memory of the scene's physical environment. Taking lunar scenes as an example, the moon has almost no perceptible atmosphere, and therefore lacks wind and weather processes similar to those on Earth. In this environment, flags would not flutter in the wind as they do on Earth, but would likely remain stationary or only sway briefly after external disturbances. For ease of filming and display, the flag installations in the Apollo missions employed a structure with a crossbar to unfold the flag, ensuring its clear visibility even in windless conditions. These objective physical facts and the static nature of authentic historical images necessitate the reasonable and controllable animate manipulation of static images in science popularization and interactive displays, providing visual motion feedback during interactive elements.
[0003] Existing image / text-to-video generation methods in recent years have largely relied on diffusion-based generative models or large-scale spatiotemporal generative networks. While these methods can generate strong visual effects, they typically suffer from long inference times, high peak memory usage, and strong dependence on high-performance GPUs / computing platforms. Furthermore, their motion control often relies on complex conditional prompts or additional model modules, making stable deployment on resource-constrained terminal or payload devices difficult. For science popularization exhibitions and task simulation scenarios requiring low-power, interpretable, and verifiable output, directly adopting large-scale generative models leads to high deployment costs, high system complexity, and difficulties in real-time interaction.
[0004] The invention patent with patent application publication number CN120852780A discloses a method for generating high-fidelity dynamic scene videos based on a single static image. The paper uses a deep neural network to perform multi-level deep feature extraction on the input single static RGB image, identify and segment key elements in the image, including foreground objects, background environment, texture regions, lighting information, spatial layout and semantic tags, etc.; based on scene semantics, combined with a prior knowledge base or physical model, it infers the potential motion patterns or change trends that may exist in the scene. However, this patent is based on neural network processing to predict optical flow.
[0005] Therefore, it is necessary to propose a lightweight, controllable, and engineering-deployable solution: given only a single static image, a small number of interpretable motion primitives are obtained through optical flow and vector field decomposition, and then continuous frames are generated by weighted synthesis with weights changing over time, thereby realizing the generation of flag motion videos with low computational resources. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide an image-to-video generation method that can be stably deployed on resource-constrained terminal or payload devices.
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0008] A lightweight method for generating videos of flags waving includes: A single static image is parsed and the target region is segmented to obtain a binary mask of the target region; the optical flow vector corresponding to the pixels in the source frame is identified, and the target frame is generated. Based on optical flow-driven image deformation, obtain motion candidate fields that conform to the motion law of the target; The user-output control signals are mapped to weights, and the motion candidate fields are weighted to construct a controlled motion field; Based on the controlled motion field, source frame, target region binary mask, and intermediate target frame, image deformation and fusion are performed on the target region to generate continuous frames and output as the target video.
[0009] In this embodiment, inverse sampling is used to generate intermediate target frames, and the result is expressed by the following formula: ; make ,but: ; ; In the formula, To represent as source frame The intermediate target frame obtained by deformation Represented as any pixel position, For a moment At pixel position The optical flow vector at that location, For the source frame In continuous coordinates Sampling results at the location, The coordinates are the integer pixel coordinates of the neighborhood. for The set of integer pixels in the four surrounding neighborhoods. For the reason The interpolation weights are determined by the fractional part of the result.
[0010] In this embodiment, obtaining candidate motion fields that conform to the target motion law includes: A set of controllable "target swing" motion fields is generated by using a parametric analytical optical flow generator, and different parameter configurations are regarded as different candidates to obtain motion candidate fields.
[0011] In this embodiment, the method for obtaining the parameterized analytical optical flow generator includes: Introduce pixel-normalized coordinates and construct a generator in this coordinate system; The generator analytically constructs the displacement field in a normalized coordinate system; For the analytically constructed displacement field, we define an enhanced basis function term and obtain the eigenvectors of the displacement field; For the feature vector, a configurable weight vector is introduced to obtain the normalized displacement field; and the normalized displacement field is driven by the parameter vector.
[0012] In this embodiment, the eigenvector of the displacement field is represented by the following formula: ; ; In the formula, , These are the eigenvectors used to generate the normalized displacement field. 、 respectively along the horizontal direction The dominant periodic modulation term constructed, 、 respectively along the vertical direction The constructed oscillation velocity basis function components, 、 These are the shear and oscillation coupling terms and their spatial amplitude envelope weights, respectively. 、 They are respectively Directional texture noise items It is the transpose symbol. For normalized sampling coordinates, For a moment.
[0013] In this embodiment, the normalized displacement field is driven by a parameter vector, and is represented as follows: ; In the formula, For the normalized displacement field, For generator, For normalized sampling coordinates, For a moment, For displacement amplitude coefficient, For spatial frequency, This is the time phase scaling factor. For the weight of the vertical swing term, This is the noise intensity parameter.
[0014] In this embodiment, constructing a controlled motion field includes: Suppose the user is The control signals that constantly provide the direction and intensity of the sports field are: ; Modulation of the UV direction of the sports field is achieved by using a component-weighted method: ; In the formula, for Control signals at all times Let be the direction vector, where ; for Always Displacement in directional component for Always Displacement in directional component For strength scalar, A controlled motion field modulated by user control signals. For pixel coordinates, , In order to improve the sports field , Modulation coefficients on the directional components, Let the initial motion field vector be... This is the gain coefficient, used to control the degree to which user input affects the amplitude of motion.
[0015] In this embodiment, generating consecutive frames and outputting them as the target video includes: Using a bilinear sampling operator, the input static image is deformed according to a controlled motion field to obtain a multi-frame deformation sequence; The deformation results are written only within the target area to obtain the final output frame, which is then used to output the video sequence.
[0016] In this embodiment, the final output frame is represented by the following formula: ; The output video sequence is represented by the following formula: ; In the formula, For the final output frame, For the source frame, For element-wise multiplication, For the target region, a binary mask is used. For intermediate target frames, Background image, for Frame morphing sequence, For a moment, It is a video sequence.
[0017] The present invention also provides a system for generating lightweight flag-waving videos using the above-described method, comprising: The target region segmentation and optical flow-driven deformation module is used to parse a single static image and segment the target region to obtain a binary mask of the target region; it also identifies the optical flow vectors corresponding to the pixels in the source frame and generates the target frame. The motion candidate set construction module is used to obtain motion candidate fields that conform to the motion law of the target based on optical flow-driven image deformation. The user control signal modeling and weighted driving module is used to map the user-output control signals to weights, weight the motion candidate fields, and construct a controlled motion field. The frame synthesis and video output module is used to perform image deformation and fusion on the target region based on the controlled motion field, source frame, target region binary mask, and intermediate target frame, to generate continuous frames and output them as the target video.
[0018] Compared with the prior art, the beneficial effects of the present invention are: This invention can be widely applied to science popularization demonstrations and mission simulation scenarios related to real flight missions. It is particularly suitable for real mission environments where computing resources are limited and power consumption and size are constrained. It enables the "controllable dynamization" of key elements such as flags in single-frame historical images or static images of the mission site, providing viewers with intuitive explanations of physical phenomena and interactive visual feedback. Because it does not rely on diffusion-type or large-scale spatiotemporal generation networks, this invention can operate stably under low computing power conditions, meeting the science popularization demonstration needs of "low resources, controllability, interpretability, and deployability" in real mission scenarios.
[0019] This invention proposes a lightweight method for generating controllable motion videos of flags from static images, consisting of "region analysis—motion primitive construction—weighted synthesis—video output." The static input image first enters the target region segmentation and optical flow-driven deformation module, which automatically extracts the flag's foreground mask and boundary contours, establishing their hierarchical relationship with the background. This provides spatial constraints for subsequent motion to act only on the flag region. The core generation mechanism uses "optical flow (motion field)-driven image deformation" to generate the target frame. Subsequently, the motion candidate set construction module generates a dense displacement / optical flow candidate field that conforms to the flag's swing pattern, forming a set of selectable and combinable motion primitives. The candidate motion fields are then fed into the user control signal modeling and weighted driving module. The system maps user-input control signals such as direction, intensity, and rhythm into weights, weighting and combining the candidate motions to obtain the target motion field. The target motion field is then input into the frame synthesis and video output module, where image deformation and fusion are performed on the flag region to generate continuous frames and output as the target video. Attached Figure Description
[0020] Figure 1 This is a flowchart of a lightweight flag-waving video generation method according to an embodiment of the present invention.
[0021] Figure 2 This is a flowchart of an embodiment of the present invention.
[0022] Figure 3 This is a schematic diagram of the generation result in an embodiment of the present invention.
[0023] Figure 4 This is a block diagram of a lightweight flag-waving video generation system according to an embodiment of the present invention. Detailed Implementation
[0024] To facilitate understanding of the technical solution of the present invention by those skilled in the art, the technical solution of the present invention will now be further described in conjunction with the accompanying drawings.
[0025] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0026] Please see Figure 1 , Figure 2 As shown, the present invention provides a lightweight method for generating flag-waving videos, comprising: S10: Analyze and segment the target region of a single static image to obtain a binary mask of the target region; confirm the optical flow vector corresponding to the pixels in the source frame and generate an intermediate target frame.
[0027] In this embodiment, the flag in a single static image is taken as the target area for detailed explanation.
[0028] In this embodiment, the target region segmentation and optical flow-driven deformation module takes a single static image as input. To decouple the flag from the background, a basic visual segmentation model is used to obtain a binary mask for the flag region. This invention does not specifically limit the basic visual segmentation model; general segmentation networks or interactive segmentation tools can be used. For real numbers, The height of the static image. The width of the static image.
[0029] Define the flag foreground as: ; In the formula, Foreground image of the flag (target area), For element-wise multiplication, This is the source frame, which is the input static image. For the target region, a binary mask is used. This is the background image. The binary mask is used to clearly distinguish between the "foreground" (the part of interest) and the "background" (the rest of the image).
[0030] In this embodiment, the input is determined and the sampling algorithm is constructed: Let the pixel coordinates be... The corresponding optical flow vector is: ; In the formula, For a moment At pixel position The optical flow vector (two-dimensional displacement vector / motion field vector) at a given location represents the distance from the reference frame to the next location. The displacement of the frame on the image plane. Optical flow vector exist Displacement components in the direction, Optical flow vector exist Displacement components in the direction, This is the transpose symbol, used to convert between row vectors and column vectors. The physical meaning is in the time parameter Below, source frame Pixels Two-dimensional displacement that occurs in the target frame.
[0031] In this embodiment, the present invention uses reverse sampling to generate the target frame to avoid holes and coverage conflicts caused by forward pushing. Indicates by The intermediate target frame obtained by deformation, then for any pixel position ,have: ; In the formula, For intermediate target frames, This is a sampling operator for continuous coordinates. In engineering implementation, bilinear interpolation is preferred, i.e., sampling continuous coordinates... The pixel value is represented as the weighted sum of its four neighboring integer pixels: ; ; In the formula, for The set of integer pixels in the four surrounding neighborhoods. For the reason The interpolation weights are determined by the fractional part, thereby achieving sub-pixel-level smooth displacement and ensuring temporal continuity. The coordinates are the integer pixel coordinates of the neighborhood. For the input image In continuous coordinates The sampling (interpolation) results at the location.
[0032] S20: Based on optical flow-driven image deformation, obtain motion candidate fields that conform to the motion law of the target.
[0033] In this embodiment, pixel coordinate normalization and basic mesh generation: Regarding the construction of the candidate motion field set, this invention employs a parametric analytical optical flow generator to produce a set of controllable "target oscillation" motion fields, treating different parameter configurations as different candidates to obtain the motion candidate fields. Specifically, this includes: S21 introduces pixel normalized coordinates and constructs a generator in this coordinate system.
[0034] In this embodiment, normalized sampling coordinates are introduced to ensure consistency with engineering implementation. Let the pixel coordinates take the top left corner as the origin. , The convention is to define the mapping from pixel coordinates to normalized coordinates. : ; ; In the formula, For normalized sampling coordinates, and Let be the mapping function from pixel coordinates to normalized coordinates and its inverse mapping function. For pixel coordinates, Normalized coordinates The components are represented as follows: Where represents the normalized coordinate components in the horizontal direction. For the vertically normalized coordinate components, , These represent the width and height of the static image, respectively.
[0035] Construct the basic grid in this coordinate system And define the normalized displacement field: Then the first The sampling grid at time (which can also be considered as a normalized optical flow grid) is ; In the formula, for The sampling grid at time t, that is... Time generator For the normalized displacement field, , These are the normalized displacement fields at... Displacement components in the y and x directions, Based on the generator.
[0036] S22, the generator analytically constructs the displacement field in the normalized coordinate system.
[0037] The generator in the normalized coordinate system Upper analytical construction of displacement field ,in To facilitate the interpretation of waveform representation, auxiliary coordinates are introduced: ; In the formula, For normalized coordinates Auxiliary coordinates per unit interval obtained by linear mapping for The truncated valid coordinates are used to ensure that the coordinates fall within the range of [0,1].
[0038] as well as, For truncation operations, z represents any scalar input, such as , .
[0039] S23. For the analytically constructed displacement field, define an enhanced basis function term and obtain the eigenvectors of the displacement field.
[0040] In this embodiment, the displacement field and sampling mesh are analyzed: The preferred time discretization method is: ; In the formula, For the first The discrete-time parameters corresponding to the frame This is the phase scaling factor. The total number of frames in the target sequence. Pi is the mathematical constant of a circle.
[0041] First, define a set of enhanced basis function terms: ; ; ; ; ; In the formula, , respectively along the horizontal direction The dominant periodic modulation term constructed, For spatial frequency, , These are the shear / swing coupling term and its spatial amplitude envelope weights, respectively. , respectively Texture noise term in the direction, , respectively along the vertical direction The constructed oscillation velocity basis function is divided into... It is a sine function. Let be a cosine function. And let them form an eigenvector: ; ; In the formula, , These are the eigenvectors used to generate the normalized displacement field.
[0042] S24. For the feature vector, a configurable weight vector is introduced to obtain the normalized displacement field; and the normalized displacement field is driven by the parameter vector.
[0043] In this embodiment, a configurable weight vector is introduced based on the obtained feature vector. , This allows the basic displacement components to be uniformly written as: ; In the formula, , These are the basic normalized displacement components, i.e., the first component in the normalized coordinate system. Frame pixel position of direction and Directional displacement of the base. , They are respectively in direction and Configurable weight vectors introduced in the direction.
[0044] Normalized displacement field From parameter vector Driver generation: ; In the formula, For the normalized displacement field, For generator, For normalized sampling coordinates, For a moment, For displacement amplitude coefficient, For spatial frequency, This is the time phase scaling factor. For the weight of the vertical swing term, This is the noise intensity parameter.
[0045] S30 maps the control signals output by the user to weights, weights the motion candidate fields, and constructs a controlled motion field.
[0046] In this embodiment, to support interactive and controllable generation, the present invention models user input as a response to the sports field. Direction and intensity modulation. Specifically, the user control signal modeling and weighted driving module includes: S31, User Control Signal Acquisition and Analysis: Assume the user is... The control signals that constantly provide the direction and intensity of the sports field are: ; in, for Control signals at all times It is a direction vector. Let be the direction vector, and be normalizable as , It is represented as the Euclidean norm. It is a scalar of intensity, such as the slider / drag amplitude / speed mapping.
[0047] S32, Component modulation coefficient calculation: Modulation of the UV direction of the motion field is achieved by using a component-weighted method. ; In the formula, for Always Displacement in directional component for Always Displacement in directional component A controlled motion field modulated by user control signals. , In order to improve the sports field , Modulation coefficients on the directional components, Let the initial motion field vector be... This is the gain coefficient, which is greater than zero, and is used to control the degree to which user input affects the amplitude of motion.
[0048] when The time indicates emphasis Directional component, when Time indicates inhibition or reversal The directional component enables controllable oscillation along a user-specified direction.
[0049] S40 performs image deformation and fusion on the target region based on the controlled motion field, source frame, target region mask, and intermediate target frame, generating continuous frames and outputting them as the target video.
[0050] In this embodiment, the frame synthesis and video output module: after obtaining the sampling grid... (or controlled sports field) After that, this module completes frame-by-frame synthesis and outputs the video sequence. To ensure consistency with engineering implementation, this invention preferably employs the following two steps: S41, Full-Image Inverse Deformation: Using a bilinear sampling operator, the input static image is deformed according to a controlled motion field to obtain a multi-frame deformation sequence.
[0051] This invention does not limit itself to a specific bilinear sampling operator, such as the grid sampling operator in deep learning frameworks. The input static image... According to the controlled sports field By deforming, we obtain Frame Deformation Sequence For sampling locations that exceed the boundary, boundary duplication or constant padding strategies can be used to reduce edge artifacts.
[0052] S42, Mask Fusion Output: The deformation result is written only in the target area to obtain the final output frame, which is then used to output the video sequence.
[0053] To maintain a stable and clear background, the deformation result is written back only within the flag area, resulting in the final output frame: ; This yields the output video sequence: ; In the formula, For the final output frame, For the source frame, For element-wise multiplication, For the target region, a binary mask is used. For intermediate target frames, Background image, for Frame morphing sequence, For a moment, It is a video sequence.
[0054] This invention utilizes a basic visual segmentation model to obtain the flag region, combines optical flow information and vector decomposition to obtain a low-dimensional motion representation, and introduces user-input control signals to weighted drive the motion primitives. Under conditions of limited computational resources, it generates a temporally consistent and controllable flag motion video from a single static image, meeting the application needs of aerospace science popularization, interactive teaching, and virtual simulation. Please refer to [link / reference]. Figure 3 The schematic diagram of the generated result shows that the present invention can "controllably animate" key elements such as flags in a single frame of historical video or static image of a mission site, providing viewers with an intuitive explanation of physical phenomena and interactive visual feedback.
[0055] Please see Figure 4 As shown, the present invention also provides a system for generating lightweight flag-waving videos using the aforementioned method, comprising: The target region segmentation and optical flow-driven deformation module is used to parse and segment a single static image to obtain a binary mask of the target region; it also identifies the optical flow vectors corresponding to the pixels in the source frame and generates the target frame.
[0056] The motion candidate set construction module is used to obtain motion candidate fields that conform to the motion law of the target based on optical flow-driven image deformation.
[0057] The user control signal modeling and weighted driving module is used to map the control signals output by the user to weights, weight the motion candidate fields, and construct a controlled motion field.
[0058] The frame synthesis and video output module is used to perform image deformation and fusion on the target region based on the controlled motion field, source frame, target region binary mask, and intermediate target frame, to generate continuous frames and output them as the target video.
[0059] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.
[0060] The above embodiments are merely examples of implementation methods of the invention. The scope of protection of the present invention is not limited to the above embodiments. For those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention.
Claims
1. A lightweight method for generating videos of flags waving, characterized in that, include: A single static image is parsed and the target region is segmented to obtain a binary mask of the target region; Identify the optical flow vectors corresponding to the pixels in the source frame and generate the target frame; Based on optical flow-driven image deformation, obtain motion candidate fields that conform to the motion law of the target; The user-output control signals are mapped to weights, and the motion candidate fields are weighted to construct a controlled motion field; Based on the controlled motion field, source frame, target region binary mask, and intermediate target frame, image deformation and fusion are performed on the target region to generate continuous frames and output as the target video.
2. The lightweight flag-waving video generation method according to claim 1, characterized in that, The intermediate target frame is generated using inverse sampling, and is expressed by the following formula: ; make ,but: ; ; In the formula, To represent as source frame The intermediate target frame obtained by deformation Represented as any pixel position, For a moment At pixel position The optical flow vector at that location, For the source frame In continuous coordinates Sampling results at the location, The coordinates are the integer pixel coordinates of the neighborhood. for The set of integer pixels in the four surrounding neighborhoods. For the reason The interpolation weights are determined by the fractional part of the result.
3. The lightweight flag-waving video generation method according to claim 1, characterized in that, Obtain candidate motion fields that conform to the target motion pattern, including: A set of controllable "target swing" motion fields is generated by using a parametric analytical optical flow generator, and different parameter configurations are regarded as different candidates to obtain motion candidate fields.
4. The lightweight flag-waving video generation method according to claim 3, characterized in that, The methods for obtaining a parameterized analytical optical flow generator include: Introduce pixel-normalized coordinates and construct a generator in this coordinate system; The generator analytically constructs the displacement field in a normalized coordinate system; For the analytically constructed displacement field, we define an enhanced basis function term and obtain the eigenvectors of the displacement field; For the feature vector, a configurable weight vector is introduced to obtain the normalized displacement field; and the normalized displacement field is driven by the parameter vector.
5. The lightweight flag-waving video generation method according to claim 4, characterized in that, The eigenvectors of the displacement field are expressed by the following formula: ; ; In the formula, , These are the eigenvectors used to generate the normalized displacement field. 、 respectively along the horizontal direction The dominant periodic modulation term constructed, 、 respectively along the vertical direction The constructed oscillation velocity basis function components, 、 These are the shear and oscillation coupling terms and their spatial amplitude envelope weights, respectively. 、 They are respectively Directional texture noise items It is the transpose symbol. For normalized sampling coordinates, For a moment.
6. The lightweight flag-waving video generation method according to claim 4, characterized in that, The normalized displacement field driven by the parameter vector can be represented as follows: ; In the formula, For the normalized displacement field, For generator, For normalized sampling coordinates, For a moment, For displacement amplitude coefficient, For spatial frequency, This is the time phase scaling factor. For the weight of the vertical swing term, This is the noise intensity parameter.
7. The lightweight flag-waving video generation method according to claim 1, characterized in that, Constructing a controlled sports field includes: Suppose the user is The control signals that constantly provide the direction and intensity of the sports field are: ; Modulation of the UV direction of the sports field is achieved by using a component-weighted method: ; In the formula, for Control signals at all times Let be the direction vector, where ; for Always Displacement in directional component for Always Displacement in directional component For strength scalar, A controlled motion field modulated by user control signals. For pixel coordinates, , In order to improve the sports field , Modulation coefficients on the directional components, Let the initial motion field vector be... This is the gain coefficient, used to control the degree to which user input affects the amplitude of motion.
8. The lightweight flag-waving video generation method according to claim 1, characterized in that, Generate consecutive frames and output them as the target video, including: Using a bilinear sampling operator, the input static image is deformed according to a controlled motion field to obtain a multi-frame deformation sequence; The deformation results are written only within the target area to obtain the final output frame, which is then used to output the video sequence.
9. The lightweight flag-waving video generation method according to claim 8, characterized in that, The final output frame is represented by the following formula: ; The output video sequence is represented by the following formula: ; In the formula, For the final output frame, For the source frame, For element-wise multiplication, For the target region, a binary mask is used. For intermediate target frames, Background image, for Frame morphing sequence, For a moment, It is a video sequence.
10. A system for generating a lightweight flag-waving video using any one of claims 1-9, characterized in that, include: The target region segmentation and optical flow-driven deformation module is used to analyze a single static image and segment the target region to obtain a binary mask of the target region. Identify the optical flow vectors corresponding to the pixels in the source frame and generate the target frame; The motion candidate set construction module is used to obtain motion candidate fields that conform to the motion law of the target based on optical flow-driven image deformation. The user control signal modeling and weighted driving module is used to map the user-output control signals to weights, weight the motion candidate fields, and construct a controlled motion field. The frame synthesis and video output module is used to perform image deformation and fusion on the target region based on the controlled motion field, source frame, target region binary mask, and intermediate target frame, to generate continuous frames and output them as the target video.