Fluid animation generation method and device combining physics and data drive

By combining physical models and data-driven methods, self-supervised motion optical flow capture and self-attention-driven texture feature learning networks are used to solve the problems of inaccurate fluid animation generation and insufficient texture feature acquisition in the prior art, and high-quality fluid animation generation is achieved.

CN119152094BActive Publication Date: 2025-05-09CAPITAL NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411265161.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2025-05-09
Estimated Expiration
2044-09-10

AI Technical Summary

Technical Problem

In the prior art, when generating natural fluid animation, it is difficult to accurately simulate the fluid characteristics in complex fluid scenes, resulting in poor animation effects and problems of hollowness and insufficient clarity in texture feature acquisition.

Method used

Using a fluid animation generation method combined with physics and data drive, a self-supervised motion optical flow capture network of two-way sequences and a multi-scale dual-flow texture feature learning network driven by self-attention, combined with the intelligent physical model selection and motion field simulation module driven by fluid scene perception, generate more delicate and realistic fluid animation.

Benefits of technology

It effectively alleviates the hollowness and blur problems in fluid animation, improves the delicateness and authenticity of the animation, and is suitable for a variety of fluid scenes, such as rivers, waterfalls, mountain springs, etc.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119152094B_ABST
    Figure CN119152094B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and device for generating fluid animations combining physics and data drive, and the method comprises: step S1: constructing a transparent fluid motion field stage, applying consistency constraints to the optical flow of a bidirectional video sequence, using real fluid video data in an original video data set S for self-supervised training, generating average optical flow data F and constructing training data; step S2: animation generation model training stage, inputting the training data into a multi-scale dual-stream texture feature learning network based on self-attention drive, and training in combination with the original video data set S and the average optical flow data F; step S3: fluid animation generation stage, using a motion field M obtained by simulating an input image using a physical model to guide the generation of actual natural fluid animations. The method provided by the present invention improves the physical authenticity of fluid animation generation and provides a new solution for the simulation and visualization of complex fluid phenomena.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision, deep learning and computer graphics, and specifically relates to a method and device for generating fluid animation by combining physics and data drive. Background Art

[0002] Converting static images into animations represents a compelling and rapidly developing field that is essential for advances in multimedia editing. The process aims to enhance visual imagination beyond the texture effects presented by a single image. Diffusion models and advanced generative models such as Open AI's Sora and Runway's Gen-2 model have played an important role in achieving realistic and high-definition animations through a comprehensive end-to-end learning process. However, accurately simulating the complexity of natural phenomena, especially fluid dynamics in the real world, remains a huge challenge without directly applying the laws of physics.

[0003] In order to generate natural fluid animations, effective strategies to address these challenges include extracting and estimating motion fields from static images. Previous related studies have mentioned that by introducing a stage dedicated to motion field estimation, deep features are extracted to introduce dynamic characteristics into static images. These methods allow static images to present lifelike motion effects, increasing the realism and viewing experience of the images.

[0004] In animation generation networks, an innovative approach is to introduce physics solvers to enhance the realism of animations by simulating motion fields. This approach can more accurately reproduce physical phenomena in nature by modeling the scene in the image. However, these methods usually adopt a unified physics solver in various scenes without fully considering the specific conditions under which the solver is applicable. This oversight may lead to insufficient perception of fluid properties in different scenes, thereby affecting the accuracy of the solution. For example, when simulating a river and a waterfall, a unified physics solver may not accurately capture the different characteristics of the fluids in these two scenes, resulting in poor animation effects.

[0005] At the same time, in terms of texture feature acquisition, these methods enhance animation effects by expanding the parameters of neural networks, but this is accompanied by a significant increase in training costs. Extended parameterization means that more data and computing resources are required, which places high demands on hardware in practical applications and increases the difficulty of implementation. In addition, due to inherent design limitations, they fail to fully address the problems of fine feature extraction and effective texture mapping, resulting in consistency defects such as texture holes and lack of clarity. These defects are particularly evident in high-resolution animations, which directly affect the audience's visual experience. Summary of the invention

[0006] In order to solve the above technical problems, the present invention provides a fluid animation generation method and device combining physics and data drive, which effectively alleviates problems such as blur and holes, and performs well in complex water flow scenes.

[0007] To achieve the above object, the technical solution adopted by the present invention is:

[0008] A method for generating fluid animation by combining physics and data drive, comprising:

[0009] Step S1: In the stage of constructing the transparent fluid motion field, a self-supervised motion optical flow capture network based on the bidirectional sequence is used to impose consistency constraints on the optical flow of the bidirectional video sequence. The real fluid video data in the original video dataset S is used for self-supervised training. The trained motion optical flow capture network is used to generate average optical flow data F and construct training data.

[0010] Step S2: in the animation generation model training phase, the training data is input into a multi-scale two-stream texture feature learning network driven by self-attention, the relevance of texture features is enhanced through a multi-scale deformation structure, and training is performed by combining the original video data set S with the average optical flow data F;

[0011] Step S3: In the fluid animation generation stage, the intelligent physical model selection and motion field simulation module driven by fluid scene perception performs fluid scene perception on the input image and intelligently selects a suitable physical model for motion field simulation. The motion M obtained by the physical model simulation is used to guide the trained two-stream texture feature learning network to generate actual natural fluid animation.

[0012] On the other hand, the present invention also provides a fluid animation generating device combining physics and data drive, comprising:

[0013] A self-supervised optical flow capture and generation module based on consistency constraints of bidirectional sequences, which uses bidirectional video sequence input and performs consistency constraints on bidirectional optical flow output, uses real fluid video data in the original video dataset S for self-supervised training, and uses the trained motion optical flow capture network to generate average optical flow data F and construct training data;

[0014] A self-attention driven multi-scale two-stream texture learning and generation module inputs the training data into a self-attention driven multi-scale two-stream texture feature learning network, enhances the relevance of texture features through a multi-scale deformation structure, and combines the original video data set S with the average optical flow data F for training;

[0015] The intelligent physical model selection and motion field simulation module driven by fluid scene perception is used to perform fluid scene perception on the input image and intelligently select a suitable physical model for motion field simulation. The motion field M obtained by simulating the physical model is used to guide the trained two-stream texture feature learning network to generate actual natural fluid animation.

[0016] The beneficial effects of the present invention are:

[0017] The present invention designs a self-supervised motion optical flow capture network using bidirectional sequences, which provides optical flow data for the training of animation generation models and can estimate the optical flow of fluids more accurately, especially in complex water flow scenes.

[0018] A self-attention driven multi-scale two-stream texture feature learning network is designed, which can better capture long-distance dependencies and improve the understanding and processing capabilities of complex texture features. The network uses the bidirectional optical flow data generated by the self-supervised motion optical flow capture model of the bidirectional sequence for training. During training, it can obtain richer and more comprehensive motion information, improve the ability to capture texture details, and thus enhance its learning and texture association capabilities for image texture details. This method effectively reduces the problems of holes and blur in the generated fluid animation, making the generated animation more delicate and realistic, and greatly improving the quality of animation generation;

[0019] The designed fluid scene perception-driven intelligent physical model selection and motion field simulation module introduces fluid scene perception to perform intelligent selection of physical simulation for complex fluid scenes. This method enables the animation generation model to be applicable to a variety of scenes, such as rivers, waterfalls, mountain springs, etc., and is more in line with the physical characteristics of the actual scene. It provides real physical guidance for the generation model, effectively alleviating the problem of the generation model not conforming to physical laws due to the lack of physical guidance. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A flow chart of a method for generating fluid animation combining physics and data drive according to the present invention;

[0021] Figure 2 It is a schematic diagram of the principle of the self-supervised motion optical flow capture network based on bidirectional sequence of the present invention;

[0022] Figure 3 A schematic diagram of the principle of a self-attention driven multi-scale two-stream texture feature learning network of the present invention;

[0023] Figure 4 This is a schematic diagram of the principles of the intelligent physical model selection and motion field simulation module driven by fluid scene perception of the present invention. DETAILED DESCRIPTION

[0024] The present invention provides a method and device for generating fluid animation by combining physics and data drive, which improves the video quality and physical authenticity of the generated fluid animation.

[0025] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below through the accompanying drawings and embodiments.

[0026] like Figure 1 As shown, a method for generating fluid animation combining physics and data drive provided by an embodiment of the present invention includes the following steps:

[0027] Step S1: In the stage of constructing the transparent fluid motion field, a self-supervised motion optical flow capture network based on a bidirectional sequence is used to impose consistency constraints on the optical flow of the bidirectional video sequence, thereby enhancing the ability to capture complex motion patterns. The real fluid video data in the original video dataset S is used for self-supervised training, and the trained motion optical flow capture network is used to generate average optical flow data F and construct training data.

[0028] Step S2: in the animation generation model training phase, the training data is input into a multi-scale two-stream texture feature learning network driven by self-attention, the relevance of texture features is enhanced through a multi-scale deformation structure, and training is performed by combining the original data set S with the average optical flow data F;

[0029] Step S3: In the fluid animation generation stage, based on the intelligent physical model selection and motion field simulation module driven by fluid scene perception, the input image is subjected to fluid scene perception and a suitable physical model is intelligently selected for motion field simulation; the motion field M obtained by physical perception simulation is used to guide the trained two-stream texture feature learning network to perform actual natural fluid animation generation;

[0030] Optionally, the present invention may further include step S4: a fluid animation generation platform, which integrates a multi-scale dual-stream texture feature learning network driven by self-attention and an intelligent physical model selection and motion field simulation module driven by fluid scene perception. Users can generate fluid animations by inputting a single image and sparse motion cues, and supports users to edit and modify the generated effects.

[0031] like Figure 2 As shown, a schematic diagram of the structure of the self-supervised motion optical flow capture model using bidirectional sequences is shown.

[0032] The present invention designs a self-supervised motion optical flow capture model using bidirectional sequences. The model extracts joint texture features from the extracted multi-frame image features, calculates the correlation between the multiple frames, and finally obtains the bidirectional optical flow by iterating an optical flow calculation module to provide optical flow data for the training of the animation generation model. The model uses bidirectional video sequences as input, and strengthens the pixel-level consistency constraints on the output bidirectional optical flow, strengthens the model's learning of details, and can more accurately estimate the optical flow of the fluid, especially in complex water flow scenes.

[0033] In one embodiment, the above step S1 includes:

[0034] Step S11: For the video sequence in the original video dataset S , construct the forward video sequence and reverse video sequence ,in An image representing the 1st, 2nd, 3rd, ..., nth frame of the video data;

[0035] Step S12: forward video sequence and reverse video sequence At the same time, it is sent to a motion optical flow capture network to estimate the corresponding video optical flow sequence and , and perform consistency constraints;

[0036] (1)

[0037] in, Indicates the calculation of forward optical flow and reverse optical flow The Euclidean distance between represents the total number of pixels, i represents the pixel index, Represents the estimated forward video optical flow sequence The obtained pixel point is and Directional optical flow value, , Represents the estimated inverse video optical flow sequence The obtained pixel point is and Directional optical flow value. This loss function further emphasizes the bidirectional consistency constraint at the pixel level, aiming to enhance the model's ability to accurately estimate complex motions. In addition to strengthening the consistency constraint between bidirectional optical flows, self-supervision is also used for training to address the problem that the original video dataset lacks high-quality optical flow data as supervision.

[0038] Step S13: Use the trained motion optical flow capture network to perform motion optical flow capture on the video sequence in the original video data set S to obtain the video sequence The estimated optical flow sequence ,in Represents the optical flow of two adjacent frames respectively, and estimates the optical flow sequence The video sequence is obtained by averaging the optical flows in The average optical flow data F.

[0039] Step S14: For the video sequence in the original video dataset S Perform preprocessing and randomly extract a starting frame , an intermediate frame and an end frame , construct data pairs As training data for the two-stream texture feature learning network, As network input, Acts as true labels to constrain the network’s generative capabilities;

[0040] like Figure 3 As shown, a schematic diagram of the structure of the two-stream texture feature learning model is shown.

[0041] The present invention designs a dual-stream texture feature learning network. The network is based on a self-attention network structure, and the network can better capture long-distance dependencies, improve the understanding and processing capabilities of complex texture features, and the network uses an independent Z channel to extract scene features from images and perform mixed feature splicing with fluid features. The self-attention network structure extracts fluid features from images and performs multi-scale feature deformation, which strengthens the network's learning of feature details. The network is trained using bidirectional optical flow data generated by a self-supervised motion optical flow capture model of a bidirectional sequence during training. The network can obtain richer and more comprehensive motion information during training, improve the ability to capture texture details, and thus enhance its learning of image texture details and texture association capabilities. This method effectively alleviates problems such as holes and blurs in generating fluid animations, making the generated animations more delicate and realistic, and greatly improving the generation quality of animations.

[0042] In one embodiment, the above step S2 includes:

[0043] Step S21: For data pair , the starting frame , end frame Input into the attention texture feature encoder to extract the starting frame image features respectively and end frame image features ; At the same time, the start frame , end frame Input to the convolutional neural network encoder connected to the attention texture feature encoder to extract the starting frame combination weight parameter and end frame combined weight parameter , where each layer of features of the attention texture feature encoder is sent to the convolutional neural network encoder and concatenated with the convolutional features from the previous layer of the convolutional neural network encoder, and the concatenated features are sent to the next layer of convolution of the convolutional neural network encoder;

[0044] (2)

[0045] in, Indicates the concatenation of two features in the channel dimension. Represents the convolutional neural network encoder The convolutional features of the layer, Represents the attention texture feature encoder The result of the layer, and They represent the convolutional neural network encoder Layer and attention texture feature encoder layer;

[0046] Step S22: For data pair , according to Euler integration, we get the starting frame To middle frame The first optical flow and intermediate frames To end frame The second optical flow ; For the extracted starting frame image features and end frame image features , respectively, according to the first optical flow and the second optical flow Deform to obtain the deformed starting frame image features And the end frame image features after deformation :

[0047] (3)

[0048] (4)

[0049] in, It indicates that the feature displacement deformation process of pixels is carried out under the guidance of optical flow, and the feature of the starting frame image after deformation is And the end frame image after deformation Perform a linear combination:

[0050] (5)

[0051] Where e represents the exponential function, Indicates time The final feature representation of and Respectively represent the start time and end time;

[0052] Step S23: For each layer of the attention texture feature encoder in step S21, feature deformation and linear combination are performed according to step S22 to obtain different scales at time The feature representation ,in, Represents the attention texture feature encoder The layer characteristic deformation combination is formed at time The characteristic of ;

[0053] Step S24: For , except that the result of the last layer is directly input into an attention texture feature decoder, all other layers are input into the attention texture feature decoder in the form of skip connection, as shown in formula (6):

[0054] (6)

[0055] in, Represents the attention texture feature decoder The characteristics of the layer, Represents the attention texture feature decoder layer, and finally obtain the intermediate frame through the attention texture feature decoder The predicted image .

[0056] like Figure 4 ,shown, is a flow chart of the physical perception simulation.

[0057] The present invention designs a physical perception simulation module, and for complex fluid scenes, it introduces fluid scene perception to perform intelligent selection of physical simulation. The method first performs fluid scene perception on the input image, selects a physical model based on the perception result, selects a suitable physical model to estimate the motion field, and simultaneously performs data preprocessing on the input to obtain a scene depth map and a dense motion field to meet the conditions of various physical models. The method introduces two physical models: the particle grid method PIC and the shallow water equation method SWE, to deal with scenes with obvious high drop and shallow water slow flow scenes respectively. This method enables the animation generation model to be applicable to more diverse scenes, such as rivers, waterfalls, mountain springs, etc., and is more in line with the physical characteristics of actual scenes.

[0058] In one embodiment, the above step S3 includes:

[0059] Step S31: The image features of the input image are obtained through step S21. Here, only the image features of the initial frame are used. The image features of the initial frame are input into the classification network to obtain the scene classification result Class. According to the scene classification result Class, the corresponding physical model is selected for simulation;

[0060] Step S32: For the input image, use a monocular depth estimation network to estimate the depth of the input image , and construct the 3D surface of the liquid based on the depth information and the set camera pose:

[0061] (7)

[0062] in, represents the pixel index, represents the extracted particle position, Represents depth information, Represents the pixel coordinates on the image, and the camera parameters are set as , , , , , Respectively represent the scaling factors of the x-axis and y-axis directions on the image plane, , Represent the positions of the principal point in the x-axis and y-axis directions, respectively, corresponding to a perspective camera with a vertical viewing angle of 90 degrees;

[0063] Step S33: In addition to the image input by the user, the user will also input a sparse motion cue to guide the generation of the animation, initialize the input sparse motion cue, and generate a pixel-level dense motion field through the nearest neighbor averaging method. , pixels The velocity value at is the exponential average of the adjacent elements, as shown in formula (8):

[0064] (8)

[0065] in, Represents the velocity of a pixel in the fluid region. Indicates Marking speed. Represents the pixels on the image With The Euclidean distance between the locations of the markers. is a parameter related to the image size;

[0066] Step S34: For pixel-level dense motion field , by converting the set camera pose into the velocity field corresponding to the particles on the 3D liquid surface , as shown in formula (9):

[0067] (9)

[0068] in, is the speed of a two-dimensional pixel, are the camera parameters, , Respectively represent the scaling factors of the x-axis and y-axis directions on the image plane, represents the pixel index, represents the extracted particle position;

[0069] Step S35: For the input image and velocity field According to the scene classification result Class, select the particle grid method or shallow water equation method for physical simulation to obtain the physical simulation motion field , as shown in formula (10):

[0070] (10)

[0071] in, They represent the physical simulation processes of the particle mesh method and the shallow water equation method respectively;

[0072] Step S36: Using a motion smoothing network to simulate the motion field of the physical simulation Smoothing to get a smooth motion field ;

[0073] Step S37: When actually generating the animation, set the data pair for step S37. ,in, is the input image, The smoothed motion field simulated for physical perception will be used as the smoothed motion field arrive The motion field of the input graph can be obtained through Euler integration. The motion field of the frame, i.e. arrive The playground; at this time, the construction generates the The data pair of the frame is Repeat this process to get the data pair sequence ,Will The data in step S2 are sequentially generated to obtain the corresponding predicted frame sequence. ,in, Indicates Frame images, thereby obtaining the generated fluid animation video sequence;

[0074] Step S38: Build a fluid animation generation platform to facilitate users to create high-quality fluid animations. The platform supports users to generate fluid animations by uploading a single image and providing motion prompts. In addition, the platform also provides editing functions for elements such as scenes, backgrounds, and fluid types, and supports users to modify the generated effects. The modified effects and corresponding motion prompt information will be saved for subsequent model updates and training to improve the accuracy and quality of animation generation.

[0075] On the other hand, a fluid animation generation device combining physics and data drive, the device can optionally perform each step of the above method, and the device specifically includes:

[0076] A self-supervised optical flow capture and generation module based on consistency constraints of bidirectional sequences, which uses bidirectional video sequence input and performs consistency constraints on bidirectional optical flow output, uses real fluid video data in the original video dataset S for self-supervised training, and uses the trained motion optical flow capture network to generate average optical flow data F and construct training data;

[0077] A self-attention driven multi-scale two-stream texture learning and generation module inputs the training data into a self-attention driven multi-scale two-stream texture feature learning network, enhances the relevance of texture features through a multi-scale deformation structure, and combines the original video data set S with the average optical flow data F for training;

[0078] The intelligent physical model selection and motion field simulation module driven by fluid scene perception is used to perform fluid scene perception on the input image and intelligently select a suitable physical model for motion field simulation. The motion field M obtained by simulating the physical model is used to guide the trained two-stream texture feature learning network to generate actual natural fluid animation.

[0079] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for generating fluid animation by combining physics and data drive, characterized in that: include: Step S1: In the stage of constructing the transparent fluid motion field, a self-supervised motion optical flow capture network based on the bidirectional sequence is used to impose consistency constraints on the optical flow of the bidirectional video sequence. The real fluid video data in the original video dataset S is used for self-supervised training. The trained motion optical flow capture network is used to generate average optical flow data F and construct training data. Step S2: in the animation generation model training phase, the training data is input into a multi-scale two-stream texture feature learning network driven by self-attention, the relevance of texture features is enhanced through a multi-scale deformation structure, and training is performed by combining the original video data set S with the average optical flow data F; Step S3: In the fluid animation generation stage, the intelligent physical model selection and motion field simulation module driven by fluid scene perception performs fluid scene perception on the input image and intelligently selects a suitable physical model for motion field simulation. The motion field M obtained by the physical model simulation is used to guide the trained two-stream texture feature learning network to perform actual natural fluid animation generation. Wherein, the step S1 comprises: Step S11: For the video sequence in the original video dataset S , construct the forward video sequence and reverse video sequence ,in An image representing the 1st, 2nd, 3rd, ..., nth frame of the video data; Step S12: forward video sequence and reverse video sequence At the same time, it is fed into the motion optical flow capture network with consistency constraints to estimate the corresponding video optical flow sequences respectively. and , and perform consistency constraints: (1) in, The Euclidean distance of each pixel optical flow value used to calculate the bidirectional optical flow represents the total number of pixels, i represents the pixel index, , Represents the estimated forward video optical flow sequence The obtained pixel point is and Directional optical flow value, Represents the estimated inverse video optical flow sequence The obtained pixel point is and Directional optical flow value; use self-supervision to train to address the problem that the original video data does not have high-quality optical flow; Step S13: Use the trained motion optical flow capture network to perform motion optical flow capture on the video sequence in the original video data set S to obtain the video sequence The estimated optical flow sequence ,in Represents the optical flow of two adjacent frames respectively, and estimates the optical flow sequence The video sequence is obtained by averaging the optical flows in The average optical flow data F; Step S14: For the video sequence in the original video dataset S Perform preprocessing and randomly extract a starting frame , an intermediate frame and an end frame , construct data pairs As training data for the two-stream texture feature learning network, As network input, Acts as true labels to constrain the network’s generation capabilities.

2. The method for generating fluid animation by combining physics and data drive according to claim 1, characterized in that: The step S2 comprises: Step S21: For data pair , the starting frame , end frame Input into the attention texture feature encoder to extract the starting frame image features respectively and end frame image features ; Set the start frame , end frame Input to the convolutional neural network encoder connected to the attention texture feature encoder to extract the starting frame combination weight parameter and end frame combined weight parameter , where each layer of features of the attention texture feature encoder is sent to the convolutional neural network encoder and concatenated with the convolutional features from the previous layer of the convolutional neural network encoder. The concatenated features are sent to the next layer of convolution of the convolutional neural network encoder: (2) in, Indicates the concatenation of two features in the channel dimension. Represents the convolutional neural network encoder The convolutional features of the layer, Represents the attention texture feature encoder The result of the layer, and They represent the convolutional neural network encoder Layer and attention texture feature encoder layer; Step S22: For data pair , according to Euler integration, we get the starting frame To middle frame The first optical flow and intermediate frames To end frame The second optical flow ; For the extracted starting frame image features and end frame image features , respectively, according to the first optical flow and the second optical flow Deform to obtain the deformed starting frame image features And the end frame image features after deformation : (3) (4) in, It indicates that the feature displacement deformation process of pixels is carried out under the guidance of optical flow, and the feature of the starting frame image after deformation is And the end frame image features after deformation Perform a linear combination: (5) Where e represents the exponential function, Indicates time The final feature representation is, and Represent the start time and end time respectively. , , respectively represent the starting time To the middle moment The process and end time of feature change To the middle moment The process of feature change, , Respectively indicate the end time To the middle moment Frame difference and start time To the middle moment Frame difference; Step S23: For each layer of the attention texture feature encoder in step S21, feature deformation and linear combination are performed according to step S22 to obtain different scales at time The feature representation ,in, Represents the attention texture feature encoder The layer feature deformation and linear combination form the The characteristic representation of ; Step S24: For , except that the result of the last layer is directly input into the attention texture feature decoder, the results of the remaining layers are input into the attention texture feature decoder in the form of skip connections, as shown in formula (6): (6) in, Represents the attention texture feature decoder The characteristics of the layer, Represents the attention texture feature decoder layer, and finally obtain the intermediate frame through the attention texture feature decoder The predicted image .

3. The method for generating fluid animation by combining physics and data drive according to claim 2, characterized in that: The step S3 comprises: Step S31: In the actual animation generation stage, the user inputs a single image and the corresponding sparse motion prompt, and obtains the image features of the input image through step S21 , where only the image features of the starting frame are used, the image features of the starting frame are input into the classification network to obtain the scene classification result Class, and the corresponding physical model is selected for simulation according to the scene classification result Class; Step S32: For the input image, use a monocular depth estimation network to estimate the depth information of the input image , and based on the depth information The 3D surface of the liquid is constructed with the set camera pose, as shown in formula (7): , (7) , in, represents the pixel index, represents the extracted particle position, Represents depth information, Represents the pixel coordinates on the image, and the camera parameters are set as , , , , , Respectively represent the scaling factors of the x-axis and y-axis directions on the image plane, , Represent the positions of the principal point in the x-axis and y-axis directions, respectively, corresponding to a perspective camera with a vertical viewing angle of 90 degrees; Step S33: In addition to the input image, a sparse motion cue is input to guide the generation of the animation, the input sparse motion cue is initialized, and a pixel-level dense motion field is generated by the nearest neighbor averaging method. , at coordinates ( ) is the exponential average of the adjacent elements, as shown in formula (8): (8) in, represents the velocity of the pixel in the fluid region, Indicates Marking speed, Represents the pixel coordinates on the image With The Euclidean distance between the marker positions, is a parameter related to the image size; Step S34: For pixel-level dense motion field , by converting the set camera pose into the velocity field corresponding to the particles on the 3D liquid surface , as shown in formula (9): (9) in, is the velocity of a two-dimensional pixel; Step S35: For the input image and velocity field According to the scene classification result Class, select the particle grid method or shallow water equation method for physical simulation to obtain the physical simulation motion field , as shown in formula (10): (10) in, They represent the physical simulation processes of the particle mesh method and the shallow water equation method respectively; Step S36: Using a motion smoothing network to simulate the motion field of the physical simulation Smooth the motion field ; Step S37: When actually generating the animation, ,in, For the input image, smooth motion field As To the next frame The motion field of the input image is obtained by Euler integration. The motion field of the frame, i.e. arrive The playground of The data pair of the frame is , repeat the construction The process of the frame data pair obtains the data pair sequence ,Will The data in step S2 are sequentially generated to obtain the corresponding predicted frame sequence. ,in, Indicates The image of the frame is obtained to obtain the generated fluid animation video sequence; Step S38: Build a fluid animation generation platform to support the correction of the generated fluid animation effects. The corrected effects and corresponding motion prompt information will be saved for subsequent model updating and training.

4. A fluid animation generation device combining physics and data drive, used to execute the method according to any one of claims 1 to 3, characterized in that: include: A self-supervised optical flow capture and generation module based on consistency constraints of bidirectional sequences, which uses bidirectional video sequence input and performs consistency constraints on bidirectional optical flow output, uses real fluid video data in the original video dataset S for self-supervised training, and uses the trained motion optical flow capture network to generate average optical flow data F and construct training data; A self-attention driven multi-scale two-stream texture learning and generation module inputs the training data into a self-attention driven multi-scale two-stream texture feature learning network, enhances the relevance of texture features through a multi-scale deformation structure, and combines the original video data set S with the average optical flow data F for training; The intelligent physical model selection and motion field simulation module driven by fluid scene perception is used to perform fluid scene perception on the input image and intelligently select a suitable physical model for motion field simulation. The motion field M obtained by simulating the physical model is used to guide the trained two-stream texture feature learning network to generate actual natural fluid animation.

Citation Information

Patent Citations

  • Interactive video character tracking method and system based on self-supervised optical flow learning

    CN117392180A

  • Multi-direction texture fusion prior guided double-flow image restoration method

    CN118429222A