Training Method and Device for a Hierarchical Neural Network for Fluid Simulation

Through a layered neural network, the background layer and fluid layer are separated and the fluid characteristic information is adjusted, the problem of background layer fluctuations in fluid simulation animation is solved, achieving a more stable and realistic fluid simulation effect.

CN114662397BActive Publication Date: 2025-07-22SENSETIME GRP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210346497.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-07-22
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

In the prior art, non-fluid parts in fluid simulation animations are prone to fluctuations, affecting the display effect.

Method used

The background layer and fluid layer in the sample video frame are separated by a layered neural network, and are represented by background feature information and fluid feature information respectively. By adjusting the fluid feature information, the simulated fluid flow effect is achieved, and the background feature information and the adjusted fluid feature information are then fused. The background layer part in the generated predicted video frame is not easily affected by fluid.

Benefits of technology

It improves the stability and authenticity of fluid simulation animations, reduces the possibility of fluctuations in the background layer, and the generated fluid simulation animation is more realistic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114662397B_ABST
    Figure CN114662397B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for training a hierarchical neural network for fluid simulation, including: obtaining multiple sample video frames in a sample video, and supervised video frames corresponding to the multiple sample video frames in the sample video, and optical flow information between the multiple sample video frames and the supervised video frames respectively; inputting the multiple sample video frames into a hierarchical neural network to be trained, and determining background feature information and fluid feature information respectively corresponding to the multiple sample video frames; adjusting the fluid feature information corresponding to each sample video frame based on the optical flow information corresponding to each sample video frame to obtain adjusted fluid feature information; constructing a predicted video frame based on the background feature information and the adjusted fluid feature information, and training the hierarchical neural network to be trained based on the predicted video frame and the supervised video frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a method and apparatus for training a hierarchical neural network for fluid simulation. Background Art

[0002] Fluid simulation refers to simulating the flow process of a fluid, and a fluid simulation animation refers to an animation showing the simulated fluid flow. Generally, a fluid simulation animation is generated based on an original image. However, in related technologies, the original image is generally processed as a whole, which results in fluctuations in non-fluid parts in the generated fluid simulation animation, affecting the display effect. Summary of the Invention

[0003] Embodiments of the present disclosure at least provide a method and apparatus for training a hierarchical neural network for fluid simulation.

[0004] In a first aspect, an embodiment of the present disclosure provides a method for training a hierarchical neural network for fluid simulation, including:

[0005] Obtaining multiple frames of sample video frames in a sample video, and corresponding supervised video frames in the sample video, and optical flow information between each of the multiple frames of sample video frames and the corresponding supervised video frames;

[0006] Inputting the multiple frames of sample video frames into a hierarchical neural network to be trained, and determining background feature information and fluid feature information corresponding to each of the multiple frames of sample video frames;

[0007] Adjusting the fluid feature information corresponding to each sample video frame based on the optical flow information corresponding to each sample video frame to obtain adjusted fluid feature information;

[0008] Constructing a predicted video frame based on the background feature information and the adjusted fluid feature information, and training the hierarchical neural network to be trained based on the predicted video frame and the supervised video frame.

[0009] In the above method, the background layer and the fluid layer in the sample video frame can be separated by a hierarchical neural network and represented by background feature information and fluid feature information respectively, and then by adjusting the fluid feature information, the effect of simulating fluid flow can be achieved. Then, the background feature information and the adjusted fluid feature information are fused. In the obtained predicted video frame, the background layer part is not easily affected by the fluid, and the possibility and range of fluctuations are small, further improving the stability and authenticity of the generated fluid simulation animation.

[0010] In a possible implementation manner, the background feature information corresponding to the sample video frame includes a background feature map for characterizing the background part in the sample video frame and a first transparency mask of the background feature map; the fluid feature information corresponding to the sample video includes a fluid feature map for characterizing the fluid part in the sample video frame and a second transparency mask of the fluid feature map.

[0011] In a possible implementation manner, the multiple video frames include a first video frame and a second video frame, and the first video frame is located before the second video frame;

[0012] Constructing the predicted video frame based on the background feature information and the adjusted fluid feature information includes:

[0013] Fusing the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame to determine the fused fluid feature information;

[0014] Constructing the predicted video frame based on the background feature information corresponding to the first video frame and the fused fluid feature information.

[0015] By fusing the adjusted fluid feature information of the first video frame and the second video frame, the adjusted fluid position can be corrected, and the generated predicted video frame is more realistic.

[0016] In a possible implementation manner, fusing the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame to determine the fused fluid feature information includes:

[0017] Inputting the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame into a to-be-trained adjustment network to determine the fused fluid feature information; wherein, the adjustment network is trained following the hierarchical neural network.

[0018] Here, when training the hierarchical network, the adjustment network can be trained synchronously, avoiding training the network separately for fusion, and improving the training speed of the model.

[0019] In a possible implementation manner, training the to-be-trained hierarchical neural network based on the predicted video frame and the supervised video frame includes:

[0020] Based on the multiple sample video frames and the intermediate video frames between the multiple sample video frames, constructing a first background supervision image; and, obtaining a pre-annotated second background supervision image, wherein the annotation information in the second background supervision image is used to annotate the background part that does not include the fluid part;

[0021] Train the hierarchical neural network to be trained based on the first background supervision image, the second background supervision image, the supervision video frame, and the predicted video frame.

[0022] In this way, by introducing multiple supervision data, it is possible to supervise the hierarchical network from multiple aspects, thereby improving the network accuracy of the hierarchical network.

[0023] In a possible implementation manner, the training of the hierarchical neural network to be trained based on the first background supervision image, the second background supervision image, the supervision video frame, and the predicted video frame includes:

[0024] Calculate a background prediction loss based on the first background supervision image and the background feature maps in the background feature information corresponding to each sample video frame; and,

[0025] Calculate a fluid prediction loss based on the second background supervision image and the transparency mask of the fluid feature maps in the fluid feature information corresponding to each sample video frame; and,

[0026] Calculate an image prediction loss based on the supervision video frame and the predicted video frame;

[0027] Train the hierarchical neural network to be trained based on the background prediction loss, the fluid prediction loss, and the image prediction loss.

[0028] In a second aspect, an embodiment of the present disclosure further provides a method for generating a fluid simulation animation, including:

[0029] Obtain an initial image and the optical flow information of the initial image;

[0030] Input the initial image and the optical flow information of the initial image into a target hierarchical neural network. The target hierarchical neural network extracts the background feature information and the fluid feature information of the initial image, adjusts the fluid feature information through the optical flow information, and constructs a predicted video frame of the target fluid in the initial image under the optical flow information based on the adjusted fluid feature information and the background feature information;

[0031] Construct a fluid simulation animation corresponding to the target fluid in the initial image based on the initial image and the predicted video frame.

[0032] In the above method, since the target hierarchical network can divide the initial image into a background part and a fluid part when determining the predicted video frame, and obtain the predicted video frame by only changing the fluid part, the background part in the fluid simulation animation obtained by this method will not change, and the generated fluid simulation animation is more realistic.

[0033] In a possible implementation, after inputting the initial image and the optical flow information of the initial image into the target hierarchical neural network, the target hierarchical neural network is used to perform the following processing procedures:

[0034] Determine the background feature information and fluid feature information corresponding to the initial image;

[0035] Adjust the fluid feature information corresponding to the initial image based on the optical flow information to obtain the adjusted fluid feature information corresponding to the initial image;

[0036] Fuse the fluid feature information corresponding to the initial image and the adjusted fluid feature information corresponding to the initial image to obtain the fused fluid feature information corresponding to the initial image;

[0037] Construct a predicted video frame corresponding to the initial image based on the background feature information corresponding to the initial image and the fused fluid feature information corresponding to the initial image.

[0038] Here, fusing the fluid feature information corresponding to the initial image and the adjusted fluid feature information corresponding to the initial image can adjust the adjusted fluid feature information by the fluid feature information to improve the authenticity of the target fluid.

[0039] In a possible implementation, the optical flow information of the initial image includes the flow velocity and flow direction of each pixel point of the initial image;

[0040] The obtaining of the initial image includes:

[0041] Obtain an initial image, the depth information of the initial image, the key-point optical flow marking information corresponding to the initial image, and the fluid mask image corresponding to the initial image;

[0042] The method further includes determining the optical flow information of the initial image according to the following method:

[0043] Construct a first three-dimensional grid map of the target scene corresponding to the initial image based on the depth information of the initial image; and,

[0044] Determine the initial optical flow information of each pixel point of the initial image based on the key-point optical flow marking information corresponding to the initial image;

[0045] Determine the optical flow information of the initial image based on the first three-dimensional grid map and the initial optical flow information.

[0046] By this method, the user only needs to mark the flow velocity information of some key points to determine the flow velocity information of the entire initial image, and then can determine the predicted video frame based on the flow velocity information of the entire initial image.

[0047] In one possible implementation, determining the optical flow information of the initial image based on the first three-dimensional grid map and the initial optical flow information includes:

[0048] Performing a preliminary adjustment on the initial optical flow information based on the first three-dimensional grid map to obtain intermediate optical flow information;

[0049] Inputting the intermediate optical flow information and the initial image into a pre-trained optical flow prediction network to determine the optical flow information of the initial image.

[0050] In one possible implementation, the method further includes:

[0051] Responding to a target virtual object addition instruction, adding a target virtual object at a target position in the initial image;

[0052] Constructing a second three-dimensional grid map based on the depth information of the initial image after adding the target virtual object;

[0053] Based on the second three-dimensional grid map and the initial optical flow information, re-determining the optical flow information of the initial image.

[0054] In one possible implementation, the method further includes:

[0055] Editing the target virtual object according to an editing instruction for the target virtual object;

[0056] Wherein, the editing instruction is used to edit at least one of the following:

[0057] Size, position, angle, object type.

[0058] In a third aspect, an embodiment of the present disclosure further provides a training device for a hierarchical neural network for fluid simulation, including:

[0059] A first acquisition module, configured to acquire multiple frames of sample video frames in a sample video, and supervised video frames corresponding to the multiple frames of sample video frames in the sample video, and optical flow information between each of the multiple frames of sample video frames and the supervised video frames;

[0060] A first determination module, configured to input the multiple frames of sample video frames into a hierarchical neural network to be trained, and determine background feature information and fluid feature information respectively corresponding to the multiple frames of sample video frames;

[0061] An adjustment module, configured to adjust the fluid feature information corresponding to each sample video frame based on the optical flow information corresponding to each sample video frame to obtain adjusted fluid feature information;

[0062] A training module, configured to construct a predicted video frame based on the background feature information and the adjusted fluid feature information, and train the hierarchical neural network to be trained based on the predicted video frame and the supervised video frame.

[0063] In a possible implementation, the background feature information corresponding to the sample video frame includes a background feature map for characterizing the background part in the sample video frame and a first transparency mask of the background feature map; the fluid feature information corresponding to the sample video includes a fluid feature map for characterizing the fluid part in the sample video frame and a second transparency mask of the fluid feature map.

[0064] In a possible implementation, the multi-frame video frames include a first video frame and a second video frame, and the first video frame is before the second video frame;

[0065] When constructing the predicted video frame based on the background feature information and the adjusted fluid feature information, the training module is configured to:

[0066] Fuse the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame to determine the fused fluid feature information;

[0067] Construct the predicted video frame based on the background feature information corresponding to the first video frame and the fused fluid feature information.

[0068] In a possible implementation, when fusing the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame to determine the fused fluid feature information, the training module is configured to:

[0069] Input the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame into the adjustment network to be trained to determine the fused fluid feature information; wherein, the adjustment network is trained following the hierarchical neural network.

[0070] In a possible implementation, when training the hierarchical neural network to be trained based on the predicted video frame and the supervised video frame, the training module is configured to:

[0071] Construct a first background supervision image based on the multi-frame sample video frames and the intermediate video frames between the multi-frame sample video frames; and obtain a pre-annotated second background supervision image, wherein the annotation information in the second background supervision image is used to annotate the background part without the fluid part;

[0072] Train the hierarchical neural network to be trained based on the first background supervision image, the second background supervision image, the supervision video frame, and the predicted video frame.

[0073] In a possible implementation manner, when training the hierarchical neural network to be trained based on the first background supervision image, the second background supervision image, the supervision video frame, and the predicted video frame, the training module is configured to:

[0074] Calculate a background prediction loss based on the first background supervision image and the background feature maps in the background feature information corresponding to each sample video frame; and,

[0075] Calculate a fluid prediction loss based on the second background supervision image and the transparency masks of the fluid feature maps in the fluid feature information corresponding to each sample video frame; and,

[0076] Calculate an image prediction loss based on the supervision video frame and the predicted video frame;

[0077] Train the hierarchical neural network to be trained based on the background prediction loss, the fluid prediction loss, and the image prediction loss.

[0078] In a fourth aspect, an embodiment of the present disclosure further provides a fluid simulation animation generation device, including:

[0079] A second acquisition module, configured to acquire an initial image and the optical flow information of the initial image;

[0080] A second determination module, configured to input the initial image and the optical flow information of the initial image into a target hierarchical neural network, extract the background feature information and fluid feature information of the initial image by the target hierarchical neural network, adjust the fluid feature information through the optical flow information, and construct a predicted video frame of the target fluid in the initial image under the optical flow information based on the adjusted fluid feature information and the background feature information;

[0081] A construction module, configured to construct a fluid simulation animation corresponding to the target fluid in the initial image based on the initial image and the predicted video frame.

[0082] In a possible implementation manner, after inputting the initial image and the optical flow information of the initial image into the target hierarchical neural network, the second determination module uses the target hierarchical neural network to perform the following processing procedure:

[0083] Determine the background feature information and fluid feature information corresponding to the initial image;

[0084] Adjust the fluid feature information corresponding to the initial image based on the optical flow information to obtain the adjusted fluid feature information corresponding to the initial image;

[0085] Fuse the fluid feature information corresponding to the initial image and the adjusted fluid feature information corresponding to the initial image to obtain the fused fluid feature information corresponding to the initial image;

[0086] Construct a predicted video frame corresponding to the initial image based on the background feature information corresponding to the initial image and the fused fluid feature information corresponding to the initial image.

[0087] In a possible implementation, the optical flow information of the initial image includes the flow velocity and flow direction of each pixel point of the initial image;

[0088] The second acquisition module, when acquiring the initial image, is used for:

[0089] Acquire an initial image, the depth information of the initial image, the key-point optical flow marking information corresponding to the initial image, and the fluid mask image corresponding to the initial image;

[0090] The second acquisition module is further used to determine the optical flow information of the initial image according to the following method:

[0091] Construct a first three-dimensional mesh map of the target scene corresponding to the initial image based on the depth information of the initial image; and,

[0092] Determine the initial optical flow information of each pixel point of the initial image based on the key-point optical flow marking information corresponding to the initial image;

[0093] Determine the optical flow information of the initial image based on the first three-dimensional mesh map and the initial optical flow information.

[0094] In a possible implementation, when determining the optical flow information of the initial image based on the first three-dimensional mesh map and the initial optical flow information, the second acquisition module is used for:

[0095] Preliminarily adjust the initial optical flow information based on the first three-dimensional mesh map to obtain intermediate optical flow information;

[0096] Input the intermediate optical flow information and the initial image into a pre-trained optical flow prediction network to determine the optical flow information of the initial image.

[0097] In a possible implementation, the device further includes an interaction module, which is used for:

[0098] Respond to a target virtual object addition instruction and add a target virtual object at the target position of the initial image;

[0099] The second acquisition module is further configured to:

[0100] Construct a second three-dimensional mesh graph based on the depth information of the initial image after adding the target virtual object;

[0101] Redetermine the optical flow information of the initial image based on the second three-dimensional mesh graph and the initial optical flow information.

[0102] In a possible implementation manner, the interaction module is further configured to:

[0103] Edit the target virtual object according to an editing instruction for the target virtual object;

[0104] Wherein, the editing instruction is used to edit at least one of the following:

[0105] Size, position, angle, object type.

[0106] In a fifth aspect, an embodiment of the present disclosure further provides a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps in the first aspect, or any possible implementation manner in the first aspect, or the steps in the second aspect, or any possible implementation manner in the second aspect are executed.

[0107] In a sixth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps in the first aspect, or any possible implementation manner in the first aspect, or the steps in the second aspect, or any possible implementation manner in the second aspect are executed.

[0108] For the effect description of the above training device for a hierarchical neural network for fluid simulation, fluid simulation animation generation device, computer device, and computer-readable storage medium, refer to the description of the above method for training a hierarchical neural network for fluid simulation and method for generating a fluid simulation animation, which will not be elaborated here.

[0109] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specific embodiments are given in conjunction with the accompanying drawings and described in detail as follows. Description of the Drawings

[0110] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the accompanying drawings required for the embodiments. The accompanying drawings here are incorporated into the specification and form a part of this specification. These accompanying drawings show the embodiments that conform to the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following accompanying drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related accompanying drawings can be obtained based on these accompanying drawings without creative efforts.

[0111] Figure 1 The flowchart of a training method for a hierarchical neural network for fluid simulation provided by the embodiments of the present disclosure is shown;

[0112] Figure 2 The schematic flowchart of a training method for a hierarchical neural network for fluid simulation provided by the embodiments of the present disclosure is shown;

[0113] Figure 3 The flowchart of a fluid simulation animation generation method provided by the embodiments of the present disclosure is shown;

[0114] Figure 4 The overall schematic flowchart of a fluid simulation animation generation method provided by the embodiments of the present disclosure is shown;

[0115] Figure 5 The schematic architecture diagram of a training device for a hierarchical neural network for fluid simulation provided by the embodiments of the present disclosure is shown;

[0116] Figure 6 The schematic architecture diagram of a fluid simulation animation generation device provided by the embodiments of the present disclosure is shown;

[0117] Figure 7 The schematic structure diagram of a computer device provided by the embodiments of the present disclosure is shown. Detailed implementation manners

[0118] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. Components of the embodiments of the present disclosure generally described and illustrated in the accompanying drawings herein may be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of the present disclosure.

[0119] It has been found through research that in the related art, when generating a fluid simulation animation, the original image is generally processed as a whole, which results in fluctuations in the non-fluid part in the generated fluid simulation animation, affecting the display effect.

[0120] Based on the above research, the present disclosure provides a training method and device for a hierarchical neural network for fluid simulation. The background layer and the fluid layer in a sample video frame can be separated through the hierarchical neural network and represented by background feature information and fluid feature information respectively. Then, by adjusting the fluid feature information, the effect of simulating fluid flow is achieved. After that, the background feature information and the adjusted fluid feature information are fused. In the obtained predicted video frame, the background layer part is not easily affected by the fluid, and the possibility and range of fluctuations are relatively small, further improving the stability and authenticity of the generated fluid simulation animation.

[0121] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0122] For ease of understanding of this embodiment, first, a training method for a hierarchical neural network for fluid simulation disclosed in the embodiments of the present disclosure will be introduced in detail. The execution subject of the training method for the hierarchical neural network for fluid simulation provided by the embodiments of the present disclosure is generally a computer device with certain computing capabilities. Such a computer device includes, for example: a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the training method for the hierarchical neural network for fluid simulation may be implemented by a processor calling computer-readable instructions stored in a memory.

[0123] See Figure 1 As shown, it is a flowchart of a training method for a hierarchical neural network for fluid simulation provided by the embodiments of the present disclosure. The method includes steps 101 to 104, where:

[0124] Step 101: Obtain multiple frames of sample video frames in a sample video, as well as the supervised video frames corresponding to the multiple frames of sample video frames in the sample video, and the optical flow information between each of the multiple frames of sample video frames and the supervised video frames.

[0125] Step 102: Input the multiple frames of sample video frames into the hierarchical neural network to be trained, and determine the background feature information and fluid feature information corresponding to each of the multiple frames of sample video frames.

[0126] Step 103: Adjust the fluid feature information corresponding to each sample video frame based on the optical flow information corresponding to each sample video frame to obtain adjusted fluid feature information.

[0127] Step 104: Construct a predicted video frame based on the background feature information and the adjusted fluid feature information, and train the hierarchical neural network to be trained based on the predicted video frame and the supervised video frame.

[0128] The following is a detailed introduction to the above steps.

[0129] Regarding step 101,

[0130] The multiple frames of sample video frames generally refer to multiple groups of sample video frames, and each group of sample video frames contains two sample video frames.

[0131] In a possible implementation, when obtaining multiple sample video frames in the sample video, multiple sample video frames in the sample video can be obtained at any time interval within a preset time period.

[0132] Exemplarily, if the preset time period is 20 seconds, multiple sample video frames can be obtained in the same sample video at time intervals of 0 second, 5 seconds, 15 seconds, and 20 seconds respectively. Multiple sample video frames within the same time period form a group of sample video frames.

[0133] In practical applications, there can be multiple sample videos, and there can also be multiple groups of sample video frames obtained in each sample video. Below, taking the sample video frames obtained in each sample video as a group of sample video frames, that is, obtaining two sample video frames as an example, the method provided by the present disclosure will be introduced.

[0134] Here, the supervised video frame generally refers to a certain video frame, for example, it can be a video frame randomly selected from the multiple video frames.

[0135] The optical flow information between any one sample video frame and the supervised video frame can refer to the flow velocity and flow direction between each pixel point in this any one sample video frame and the corresponding pixel point in the supervised video frame. Generally, the background part is the main content part in the video frame, and its optical flow information between multiple video frames is empty or 0, while the optical flow information of the fluid part is not empty.

[0136] In a possible implementation, when determining the optical flow information between the sample video frame and the supervised video frame, it can be predicted through a dense optical flow estimation model exemplarily. Specifically, the sample video frame, the supervised video frame, and the time interval between the sample video frame and the supervised video frame are input into the dense optical flow estimation model, and the dense optical flow estimation model can output the optical flow information between the sample video frame and the supervised video frame.

[0137] Regarding step 102,

[0138] Since the fluid in the sample video frame may be transparent or semi - transparent, and the sample video frame is two - dimensional, the pixel points in the sample video frame may be both background and fluid. Therefore, it is difficult to segment the background part and the fluid part of the sample video frame based on traditional semantic segmentation methods.

[0139] In a possible implementation, the background feature information corresponding to the sample video frame includes a background feature map for characterizing the background part in the sample video frame and a first transparency mask of the background feature map; the fluid feature information corresponding to the sample video includes a fluid feature map for characterizing the fluid part in the sample video frame and a second transparency mask of the fluid feature map.

[0140] Among them, the transparency mask contains the transparency information of each pixel point in the corresponding feature map. The pixel points in the above background feature map and the respective transparencies included in the first transparency mask are in one-to-one correspondence, and the pixel points in the above fluid feature map and the respective transparencies included in the second transparency mask are also in one-to-one correspondence.

[0141] Through this method, the background part and the fluid part in the sample video frame can be distinguished. For the partial pixel points that are both background and fluid, the feature dimensions can be distinguished through transparency.

[0142] Regarding step 103,

[0143] Here, the adjustment of the fluid feature information may include adjusting the fluid feature map and adjusting the transparency mask.

[0144] In a possible implementation, the multi-frame video frames are the first video frame and the second video frame, and the first video frame is before the second video frame; the optical flow information between the first video frame and the supervised video frame is the first optical flow information, and the optical flow information between the second video frame and the supervised video frame is the second optical flow information. When adjusting the fluid feature information corresponding to each sample video frame based on the optical flow information corresponding to each sample video frame, it may be based on the first optical flow information to adjust the fluid feature information corresponding to the first video frame, and based on the second optical flow information to adjust the fluid feature information corresponding to the second video frame.

[0145] Exemplarily, video frame 1 and video frame 10 in the sample video frame can be obtained (the video frame number represents the acquisition time sequence, and the smaller the number, the earlier the acquisition time), then video frame 5 is used as the supervised video frame, and then the first optical flow information between video frame 1 and video frame 5 is calculated, and the second optical flow information between video frame 5 and video frame 10 is calculated. Then, video frame 1 and video frame 10 are input into the hierarchical neural network to be trained to determine the background feature information and fluid feature information corresponding to video frame 1, and the background feature information and fluid feature information corresponding to video frame 10. Then, the fluid feature information corresponding to video frame 1 is adjusted based on the first optical flow information, and the fluid feature information corresponding to video frame 10 is adjusted based on the second optical flow information.

[0146] The flow of a fluid, at the image level, refers to the movement of the pixel points corresponding to the fluid. Therefore, in one possible implementation, when adjusting the fluid feature information, it may refer to adjusting the position information of the pixel points of the fluid part in the fluid feature map, and adjusting the position information of the transparency corresponding to each pixel point in the second transparency mask.

[0147] Specifically, when adjusting the fluid feature map based on the optical flow information, exemplarily, the position information of each pixel point in the fluid feature map after movement can be determined based on the optical flow information, and the pixel point can be moved to the position corresponding to the position information, and the transparency corresponding to the moved pixel point in the second transparency mask can also be moved to the corresponding position.

[0148] Here, since the movement of pixel points and transparency is in units of pixels, and the determined position information may be in the middle of multiple pixel positions, for this case, the pixel position closest to the determined position information can be used as the moved position.

[0149] In the fluid feature map, when multiple pixel points move to the same position, the value at that position after movement can be the average value of the pixel values of the pixel points located at that position; in the second transparency mask, when multiple transparencies move to the same position, the value at that position after movement (i.e., the moved transparency corresponding to that position) can be the average value of the transparencies at that position.

[0150] In practical applications, only the fluid part may move. Therefore, it is also possible to first determine, based on the second transparency mask, the fluid pixel points corresponding to non-zero transparencies (these pixel points are the pixel points that may be the fluid), and then move the fluid pixel points to the corresponding positions based on the optical flow information.

[0151] Regarding step 104,

[0152] Taking the obtained multiple frames of sample video frames as I t0 and I tn as an example, I t0 represents the video frame at time t0, I tn represents the video frame at time tn, t0 is earlier than tn, the supervised video frame is the video frame at the t-th moment, and the purpose of the hierarchical neural network to be trained is to predict the video frame at the t-th moment. Therefore, the optical flow information corresponding to I t0 can predict to which position the optical flow pixel points in I t0 should flow at the t-th moment, and the optical flow information corresponding to I tn can predict from which position the optical flow pixel points in I tn should start to flow at the t-th moment. Theoretically, the adjustment of the fluid feature information for I t0 and I tn is the same.

[0153] After adjusting the fluid feature map based on the optical flow information, the fluid pixel points in the fluid feature map are actually moved. After the movement, the situation of "fluid break" may occur, or the situation that the fluid flows into the non-flow area (for example, the water in the river may flow to the shore) may occur. Therefore, the adjusted fluid feature information corresponding to two frames of sample video frames can be mutually adjusted.

[0154] Therefore, in a possible implementation manner, when constructing a predicted video frame based on the background feature information and the adjusted fluid feature information, the following steps may be included by way of example:

[0155] Step A1: Fuse the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame to determine the fused fluid feature information;

[0156] Specifically, when fusing the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame to determine the fused fluid feature information, the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame can be input into the adjustment network to be trained to determine the fused fluid feature information. Among them, the adjustment network is trained following the hierarchical neural network.

[0157] Here, the fused fluid feature information includes a fused fluid feature map and a fused transparency mask. That is to say, the adjustment network is not only used to adjust the fluid feature map, but also used to adjust the transparency mask.

[0158] By way of example, when determining the fused fluid feature information, it can be calculated by the following formula:

[0159]

[0160] α' f (T i ) = α' f (T0) + α' f (T n ) (2)

[0161] Among them, I' f (T i ) represents the fused fluid feature map at time T i , I' f (T0) represents the adjusted fluid feature map at time T0, α' f (T0) represents the adjusted transparency mask of the adjusted fluid feature map at time T0, and I' f (T n ) represents Tn Adjusted fluid feature map at a moment, α' f (T n ) represents the adjusted transparency mask of the adjusted fluid feature map at moment T n Adjusted transparency mask of the adjusted fluid feature map at moment T, α' f (T i ) represents the transparency mask of the fused fluid feature map at moment T i The video frame at moment T0 is the above-mentioned first video frame, and the video frame at moment T n is the above-mentioned second video frame.

[0162] Step A2, construct the predicted video frame based on the background feature information corresponding to the first video frame and the fused fluid feature information.

[0163] Specifically, when constructing the predicted video frame based on the background feature information corresponding to the first video frame and the fused fluid feature information, exemplarily, the calculation can be performed according to the following formula:

[0164]

[0165] Where, I(T i ) represents the predicted video frame at moment T i , I b represents the background feature map of the first video frame, and α b represents the first transparency mask of the background feature map of the first video frame.

[0166] In a possible implementation manner, when training the to-be-trained hierarchical neural network based on the predicted video frame and the supervised video frame, exemplarily, the loss value can be directly calculated based on the predicted video frame and the supervised video frame, and the network parameters of the to-be-trained hierarchical neural network can be adjusted based on the loss value.

[0167] In another possible implementation manner, in order to improve the network accuracy of the hierarchical neural network, other loss values can be introduced.

[0168] Exemplarily, based on the multi-frame sample video frames and the intermediate video frames between the multi-frame sample video frames, a first background supervision image can be constructed, and a pre-annotated second background supervision image can be obtained, where the annotation information in the second background supervision image is used to annotate the background part that does not include the fluid part; then, based on the first background supervision image, the second background supervision image, the supervised video frame, and the predicted video frame, the to-be-trained hierarchical neural network can be trained.

[0169] Here, the first background supervision image can be exemplarily obtained by adding multiple video frames and the intermediate video frames between the multiple video frames pixel by pixel and then taking the average to obtain the first background supervision image.

[0170] In this way, through adding pixel by pixel and taking the average, the pixel points that completely belong to the background part and do not contain any fluid part will not change, while for the areas where the fluid flows through, the pixel values will change accordingly after addition. The first background supervision image constructed by this method can enhance the features of the background part. Supervising the training process of the hierarchical neural network based on the first background supervision image can improve the ability of the hierarchical neural network to distinguish between the background and the fluid.

[0171] The annotation information of the second background supervision image can refer to the background part that does not contain any fluid part in any video frame of the manually annotated sample video. Supervising the hierarchical neural network through the second background supervision image can also improve the ability of the hierarchical neural network to distinguish between the background and the fluid.

[0172] Specifically, when training the hierarchical neural network to be trained based on the first background supervision image, the second background supervision image, the supervised video frame, and the predicted video frame, three aspects of loss values can be calculated, specifically including:

[0173] 1. Calculate the background prediction loss based on the first background supervision image and the background feature maps in the background feature information corresponding to each sample video frame.

[0174] 2. Calculate the fluid prediction loss based on the second background supervision image and the transparency information of the fluid feature maps in the fluid feature information corresponding to each sample video frame.

[0175] 3. Calculate the image prediction loss based on the supervised video frame and the predicted video frame.

[0176] Specifically, the above loss values can be exemplarily the loss of a Generative Adversarial Network (GAN), L1 loss, perceptual loss, etc.

[0177] It should be noted here that the adjustment network for determining the fused fluid feature information is trained together with the hierarchical neural network. Therefore, when adjusting the network parameters of the to-be-trained hierarchical neural network based on the above loss value, it is also necessary to adjust the network parameters of the adjustment network. During the training process of the hierarchical neural network, the parameters of the adjustment network are iteratively updated synchronously with the parameters of the hierarchical neural network. Alternatively, during the training process of the hierarchical neural network, after the parameters of the hierarchical neural network are iteratively updated several times, the parameters of the adjustment network can be iteratively updated one or more times. In this way, the hierarchical neural network and the adjustment network can achieve a better video frame prediction effect synchronously.

[0178] In addition, the steps of constructing the predicted video frame in the above steps 102 to 104 can all be executed by the hierarchical neural network, that is, the hierarchical neural network is an end-to-end hierarchical neural network, the input of the network is the sample video frame, and the output of the network is the predicted video frame.

[0179] Next, the training method of the above hierarchical neural network for fluid simulation will be introduced in combination with specific schematic diagrams.

[0180] See Figure 2 As shown in the flowchart of a training method of a hierarchical neural network for fluid simulation provided by an embodiment of the present disclosure, specifically:

[0181] The input of the hierarchical neural network is the first sample video frame I(T0) and the second sample video frame I(T n ). After the first sample video frame I(T0) undergoes background extraction and transparency regression, background feature information (including the background feature map I b (T0) and the first transparency mask α b (T0)) of the background feature map can be obtained. After the first sample video frame I(T0) undergoes fluid encoding and transparency regression, fluid feature information (including the fluid feature map I f (T0) and the second transparency mask α f (T0)) of the fluid feature map can be obtained. In addition, after the second sample video frame I(T n ) undergoes background extraction and fluid encoding, the background feature information (including the background feature map I b (T n ) and the third transparency mask α b (T n )) and the fluid feature information (including the fluid feature map I f (T n ) and the fourth transparency mask α f (T n )) of the first sample image can also be obtained.

[0182] Then, adjust the fluid feature information of the first sample video frame and the fluid feature information of the second sample video frame according to the optical flow information at the T i th moment. After adjustment, input them into the adjustment network for fusion to obtain the fused fluid feature information. Multiply the fused fluid feature map and the fused transparency mask in the fused fluid feature information to obtain the fused fluid layer α' f (T i )I' f (T i ).

[0183] Furthermore, fuse the fused fluid layer α' f (T i )I' f (T i ) with the background feature information of the first sample video frame to obtain the predicted video frame I(T i ) at the T i th moment.

[0184] Then, based on the background feature map I b (T0) of the first sample video frame and the background feature map I b (T n ) of the second sample video frame, as well as the first background supervision image, calculate the background prediction loss; based on the second transparency mask α f (T0) of the first sample video frame and the second transparency mask α b (T n ) of the second sample video frame, as well as the second background supervision image, calculate the fluid prediction loss; based on the predicted video frame I(T i ) at the T i th moment and the video frame at the T i th moment of the sample video frame, calculate the image prediction loss, and then adjust the network parameters of the hierarchical neural network based on the three parts of losses.

[0185] In the training method of the hierarchical neural network for fluid simulation provided by the embodiments of the present disclosure, the background layer and the fluid layer in the sample video frame can be separated through the hierarchical neural network and represented by the background feature information and the fluid feature information respectively. Then, by adjusting the fluid feature information, the effect of simulating fluid flow is achieved. After fusing the background feature information and the adjusted fluid feature information, in the obtained predicted video frame, the background layer part is not easily affected by the fluid, and the possibility and range of fluctuations are small, further improving the stability and authenticity of the generated fluid simulation animation.

[0186] Based on the same concept, the embodiments of the present disclosure also provide a method for generating a fluid simulation animation. See Figure 3As shown in the figure, the flowchart of a method for generating a fluid simulation animation provided by an embodiment of the present disclosure includes the following steps:

[0187] Step 301, obtain an initial image and the optical flow information of the initial image.

[0188] Step 302, input the initial image and the optical flow information of the initial image into a target hierarchical neural network. After the target hierarchical neural network extracts the background feature information and fluid feature information of the initial image, adjust the fluid feature information through the optical flow information, and based on the adjusted fluid feature information and the background feature information, construct a predicted video frame of the target fluid in the initial image under the optical flow information.

[0189] Step 303, based on the initial image and the predicted video frame, construct a fluid simulation animation corresponding to the target fluid in the initial image.

[0190] The following is a detailed description of the above steps.

[0191] Here, the target hierarchical neural network can be trained based on the above training method of the hierarchical neural network for fluid simulation.

[0192] In the above embodiment, the input of the hierarchical neural network to be trained during the training process is multiple frames of sample video frames. After the hierarchical neural network is trained to obtain the target hierarchical neural network, the inference process of the target hierarchical neural network is different from the above training process, and its overall process is as Figure 4 shown.

[0193] Specifically, after inputting the initial image and the optical flow information of the initial image into the target hierarchical neural network, the processing process of the target hierarchical neural network is as Figure 4 shown, including the following processing processes:

[0194] Step B1, determine the background feature information and fluid feature information corresponding to the initial image.

[0195] Step B2, adjust the fluid feature information corresponding to the initial image based on the optical flow information to obtain the adjusted fluid feature information corresponding to the initial image.

[0196] Step B3, fuse the fluid feature information corresponding to the initial image and the adjusted fluid feature information corresponding to the initial image to obtain the fused fluid feature information corresponding to the initial image.

[0197] Step B4, based on the background feature information corresponding to the initial image and the fused fluid feature information corresponding to the initial image, construct a predicted video frame corresponding to the initial image.

[0198] The specific implementation process of the above steps B1 to B4 is the same as the training process and will not be repeated here.

[0199] Different from the training process, in the above step B3, during the inference process, the fluid feature information corresponding to the initial image and the adjusted fluid feature information corresponding to the initial image are directly fused, while in the training process, the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame are fused.

[0200] Alternatively, it can be understood that in the inference process of the target hierarchical neural network, the input is two identical initial images, one for stratification and adjustment of the fluid, and the other for correction of the adjusted fluid feature information.

[0201] In the above step B1, the optical flow information of the initial image includes the flow velocity and flow direction of each pixel point of the initial image.

[0202] Specifically, when acquiring the initial image, the initial image, the depth information of the initial image, the key point optical flow marking information corresponding to the initial image, and the fluid mask image corresponding to the initial image can be acquired simultaneously; illustratively, the key point optical flow marking information corresponding to the initial image and the fluid mask image corresponding to the initial image can be as follows: Figure 4 As shown, Figure 4 The medium gray area represents the fluid part, the dark area represents the background part, and the arrows are the key point optical flow marking information, including the fluid flow direction and flow rate.

[0203] After determining the key point optical flow marking information corresponding to the initial image, the optical flow information of the initial image may also be determined by the following method:

[0204] Step C1: constructing a first three-dimensional grid map of a target scene corresponding to the initial image based on the depth information of the initial image.

[0205] Step C2: determining initial optical flow information of each pixel point of the initial image based on the key point optical flow label information corresponding to the initial image.

[0206] Step C3: Determine the optical flow information of the initial image based on the first three-dimensional grid image and the initial optical flow information.

[0207] Here, step C2 refers to estimating the initial optical flow information of other pixel points based on the key-point optical flow labeling information marked by the user. After determining the initial optical flow information, since the determined initial optical flow information is only the optical flow information in the two-dimensional plane, and the optical flow information is affected by factors such as terrain, it is necessary to establish the first three-dimensional grid map of the target scene, and then estimate the three-dimensional optical flow information of the fluid based on the first three-dimensional grid map. That is to say, the optical flow information of the initial image obtained in step 301 is three-dimensional optical flow information.

[0208] Exemplarily, when based on the initial optical flow information of each pixel point of the initial image, it can be determined by the following formula:

[0209]

[0210] where, v i,j represents the optical flow information of the pixel point (i, j), V k represents the optical flow information of the Kth marked key point, d(k, ij) represents the distance between the pixel point (i, j) and the K marked key points (for example, it can be the Euclidean distance), and σ represents a hyperparameter proportional to the image size.

[0211] In a possible implementation manner, the initial optical flow information of each pixel point of the initial image may refer to the initial optical flow information of the pixel points in the fluid part of the initial image. The pixel points in the non-fluid part (i.e., the background part) of the initial image do not move, so the corresponding initial optical flow information is 0.

[0212] After constructing the first three-dimensional grid map, based on the initial optical flow information, the optical flow information at time T i can be inferred, and then video frame prediction is performed based on the optical flow information at time T i .

[0213] In another possible implementation manner, since the accuracy of the three-dimensional grid map established according to a single image is limited, after inferring the optical flow information through the three-dimensional grid map, the inference result can also be adjusted.

[0214] Exemplarily, after initially adjusting the initial optical flow information based on the first three-dimensional grid map, intermediate optical flow information can be obtained, and then the intermediate optical flow information and the initial image can be input into a pre-trained optical flow prediction network to determine the optical flow information of the initial image.

[0215] Specifically, when the optical flow prediction network is trained, it can be through the following steps:

[0216] Step D1, obtain multiple initial sample images and the key-point optical flow labeling information of the initial sample images.

[0217] Step D2: Determine the optical flow information at the S-th moment after the initial sample image according to the sample video where the initial sample image is located.

[0218] Step D3: Establish a three-dimensional grid map corresponding to the initial sample image, and determine the initial optical flow information of each pixel point of the initial sample image based on the key-point optical flow marking information corresponding to the initial image.

[0219] Step D4: Input the initial optical flow information of each pixel point of the initial sample image and the initial sample image into the optical flow prediction network to be trained, and obtain the predicted optical flow information at the S-th moment.

[0220] Step D5: Train the optical flow prediction network to be trained according to the optical flow information at the S-th moment determined in Step D2 and the predicted optical flow information at the S-th moment obtained in Step D4.

[0221] In the above Step 302, there can be multiple predicted video frames determined by the target hierarchical network, and different predicted video frames correspond to different optical flow information. After determining multiple predicted video frames, when constructing the fluid simulation animation corresponding to the target fluid in the initial image, the initial image and the predicted video frames can be arranged in the corresponding time sequence to obtain the fluid simulation animation.

[0222] In a possible implementation manner, the method provided by the present disclosure can also combine the AR display effect. Specifically, before adding a target virtual object to the target position of the initial image, the target virtual object can be edited first according to the editing instruction for the target virtual object;

[0223] wherein, the editing instruction is used to edit at least one of the following:

[0224] Size, position, angle, object type.

[0225] Then, in response to the target virtual object adding instruction, the edited target virtual object can be added to the target position of the initial image.

[0226] Or after responding to the target virtual object adding instruction and adding the target virtual object to the target position of the initial image, then in response to the editing instruction for the target virtual object, the target virtual object can be edited.

[0227] After adding the target virtual object, it may affect the flow direction and flow velocity of the target fluid in the initial image. Therefore, based on the depth information of the initial image after adding the target virtual object, a second three-dimensional grid map can be reconstructed, and then based on the second three-dimensional grid map and the initial optical flow information determined in step C2, the optical flow information of the initial image can be re-determined.

[0228] In this way, when predicting the predicted video frames after adding the target virtual object later, prediction can be performed according to the initial image and the re-determined optical flow information. Thus, interaction between the user and the fluid can be achieved, and the generated fluid simulation animation is more realistic.

[0229] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order that constitutes any limitation to the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0230] Based on the same inventive concept, the embodiments of the present disclosure also provide a training device for a hierarchical neural network for fluid simulation corresponding to the training method of the hierarchical neural network for fluid simulation. Since the principle of solving problems by the device in the embodiments of the present disclosure is similar to the above training method of the hierarchical neural network for fluid simulation in the embodiments of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be elaborated.

[0231] Refer to Figure 5 As shown, it is a schematic architecture diagram of a training device for a hierarchical neural network for fluid simulation provided by an embodiment of the present disclosure. The device includes: a first acquisition module 501, a first determination module 502, an adjustment module 503, and a training module 504; wherein,

[0232] The first acquisition module 501 is configured to acquire multiple sample video frames in a sample video, as well as the supervised video frames corresponding to the multiple sample video frames in the sample video, and the optical flow information between each of the multiple sample video frames and the supervised video frames;

[0233] The first determination module 502 is configured to input the multiple sample video frames into the hierarchical neural network to be trained, and determine the background feature information and fluid feature information respectively corresponding to the multiple sample video frames;

[0234] The adjustment module 503 is configured to adjust the fluid feature information corresponding to each sample video frame based on the optical flow information corresponding to each sample video frame to obtain adjusted fluid feature information;

[0235] A training module 504, configured to construct a predicted video frame based on the background feature information and the adjusted fluid feature information, and train the hierarchical neural network to be trained based on the predicted video frame and the supervised video frame.

[0236] In a possible implementation, the background feature information corresponding to the sample video frame includes a background feature map for characterizing the background part in the sample video frame and a first transparency mask of the background feature map; the fluid feature information corresponding to the sample video includes a fluid feature map for characterizing the fluid part in the sample video frame and a second transparency mask of the fluid feature map.

[0237] In a possible implementation, the multi-frame video frames include a first video frame and a second video frame, and the first video frame is before the second video frame;

[0238] When constructing the predicted video frame based on the background feature information and the adjusted fluid feature information, the training module 504 is configured to:

[0239] Fuse the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame to determine fused fluid feature information;

[0240] Construct the predicted video frame based on the background feature information corresponding to the first video frame and the fused fluid feature information.

[0241] In a possible implementation, when fusing the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame to determine the fused fluid feature information, the training module 504 is configured to:

[0242] Input the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame into an adjustment network to be trained to determine the fused fluid feature information; wherein, the adjustment network is trained following the hierarchical neural network.

[0243] In a possible implementation, when training the hierarchical neural network to be trained based on the predicted video frame and the supervised video frame, the training module 504 is configured to:

[0244] Construct a first background supervision image based on the multi-frame sample video frames and the intermediate video frames between the multi-frame sample video frames; and obtain a pre-annotated second background supervision image, wherein the annotation information in the second background supervision image is used to annotate the background part without the fluid part;

[0245] Train the hierarchical neural network to be trained based on the first background supervision image, the second background supervision image, the supervision video frame, and the predicted video frame.

[0246] In a possible implementation manner, when training the hierarchical neural network to be trained based on the first background supervision image, the second background supervision image, the supervision video frame, and the predicted video frame, the training module 504 is configured to:

[0247] Calculate a background prediction loss based on the first background supervision image and the background feature maps in the background feature information corresponding to each sample video frame; and,

[0248] Calculate a fluid prediction loss based on the second background supervision image and the transparency mask of the fluid feature maps in the fluid feature information corresponding to each sample video frame; and,

[0249] Calculate an image prediction loss based on the supervision video frame and the predicted video frame;

[0250] Train the hierarchical neural network to be trained based on the background prediction loss, the fluid prediction loss, and the image prediction loss.

[0251] Descriptions of the processing flows of the modules in the device and the interaction flows between the modules can refer to the relevant descriptions in the above method embodiments and will not be elaborated here.

[0252] Based on the same concept, an embodiment of the present disclosure further provides a fluid simulation animation generation device. Refer to Figure 6 As shown, it is a schematic architecture diagram of a fluid simulation animation generation device provided by an embodiment of the present disclosure, including a second acquisition module 601, a second determination module 602, a construction module 603, and an interaction module 604. Specifically:

[0253] The second acquisition module 601 is configured to acquire an initial image and the optical flow information of the initial image;

[0254] The second determination module 602 is configured to input the initial image and the optical flow information of the initial image into the target hierarchical neural network, extract the background feature information and fluid feature information of the initial image by the target hierarchical neural network, adjust the fluid feature information through the optical flow information, and construct a predicted video frame of the target fluid in the initial image under the optical flow information based on the adjusted fluid feature information and the background feature information;

[0255] The construction module 603 is configured to construct a fluid simulation animation corresponding to the target fluid in the initial image based on the initial image and the predicted video frame.

[0256] In a possible implementation, after inputting the initial image and the optical flow information of the initial image into the target hierarchical neural network, the second determination module 602 uses the target hierarchical neural network to perform the following processing procedures:

[0257] Determine the background feature information and fluid feature information corresponding to the initial image;

[0258] Adjust the fluid feature information corresponding to the initial image based on the optical flow information to obtain the adjusted fluid feature information corresponding to the initial image;

[0259] Fuse the fluid feature information corresponding to the initial image and the adjusted fluid feature information corresponding to the initial image to obtain the fused fluid feature information corresponding to the initial image;

[0260] Construct a predicted video frame corresponding to the initial image based on the background feature information corresponding to the initial image and the fused fluid feature information corresponding to the initial image.

[0261] In a possible implementation, the optical flow information of the initial image includes the flow velocity and flow direction of each pixel point of the initial image;

[0262] The second acquisition module 601, when acquiring the initial image, is used for:

[0263] Acquire an initial image, the depth information of the initial image, the key point optical flow marking information corresponding to the initial image, and the fluid mask image corresponding to the initial image;

[0264] The second acquisition module 601 is further used to determine the optical flow information of the initial image according to the following method:

[0265] Construct a first three-dimensional grid map of the target scene corresponding to the initial image based on the depth information of the initial image; and,

[0266] Determine the initial optical flow information of each pixel point of the initial image based on the key point optical flow marking information corresponding to the initial image;

[0267] Determine the optical flow information of the initial image based on the first three-dimensional grid map and the initial optical flow information.

[0268] In a possible implementation, when the second acquisition module 601 determines the optical flow information of the initial image based on the first three-dimensional grid map and the initial optical flow information, it is used for:

[0269] Preliminarily adjust the initial optical flow information based on the first three-dimensional grid map to obtain intermediate optical flow information;

[0270] Input the intermediate optical flow information and the initial image into a pre-trained optical flow prediction network to determine the optical flow information of the initial image.

[0271] In a possible implementation, the device further includes an interaction module 604, configured to:

[0272] Respond to a target virtual object addition instruction, and add a target virtual object at a target position of the initial image;

[0273] The second acquisition module 601 is further configured to:

[0274] Construct a second three-dimensional grid map based on the depth information of the initial image after adding the target virtual object;

[0275] Based on the second three-dimensional grid map and the initial optical flow information, re-determine the optical flow information of the initial image.

[0276] In a possible implementation, the interaction module 604 is further configured to:

[0277] Edit the target virtual object according to an editing instruction for the target virtual object;

[0278] Wherein, the editing instruction is used to edit at least one of the following:

[0279] Size, position, angle, object type.

[0280] Based on the same inventive concept, an embodiment of the present disclosure further provides a computer device. Referring to Figure 7 As shown, it is a schematic structural diagram of a computer device 700 provided by an embodiment of the present disclosure, including a processor 701, a memory 702, and a bus 703. Among them, the memory 702 is used to store execution instructions, including an internal memory 7021 and an external memory 7022; the internal memory 7021 here is also called the main memory, and is used to temporarily store the operation data in the processor 701 and the data exchanged with the external memory 7022 such as a hard disk. The processor 701 exchanges data with the external memory 7022 through the internal memory 7021. When the computer device 700 runs, the processor 701 communicates with the memory 702 through the bus 703, so that the processor 701 executes the following instructions:

[0281] Obtain multiple frames of sample video frames in a sample video, and the supervised video frames corresponding to the multiple frames of sample video frames in the sample video, and the optical flow information between each of the multiple frames of sample video frames and the supervised video frames;

[0282] Input the multi-frame sample video frames into the hierarchical neural network to be trained, and determine the background feature information and fluid feature information corresponding to the multi-frame sample video frames respectively;

[0283] Adjust the fluid feature information corresponding to each sample video frame based on the optical flow information corresponding to each sample video frame to obtain the adjusted fluid feature information;

[0284] Construct a predicted video frame based on the background feature information and the adjusted fluid feature information, and train the hierarchical neural network to be trained based on the predicted video frame and the supervised video frame.

[0285] Alternatively, the processor 701 may execute the following instructions:

[0286] Obtain an initial image and the optical flow information of the initial image;

[0287] Input the initial image and the optical flow information of the initial image into the target hierarchical neural network. The target hierarchical neural network extracts the background feature information and fluid feature information of the initial image, adjusts the fluid feature information through the optical flow information, and constructs a predicted video frame of the target fluid in the initial image under the optical flow information based on the adjusted fluid feature information and the background feature information;

[0288] Construct a fluid simulation animation corresponding to the target fluid in the initial image based on the initial image and the predicted video frame.

[0289] An embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the training method of the hierarchical neural network for fluid simulation and the fluid simulation animation generation method described in the above method embodiments. Wherein, the storage medium may be a volatile or non-volatile computer-readable storage medium.

[0290] An embodiment of the present disclosure further provides a computer program product, which carries program codes. The instructions included in the program codes can be used to execute the steps of the training method of the hierarchical neural network for fluid simulation and the fluid simulation animation generation method described in the above method embodiments. For details, please refer to the above method embodiments and will not be elaborated here.

[0291] Among them, the above computer program product can be specifically implemented in a manner of hardware, software or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0292] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some communication interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0293] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0294] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0295] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0296] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting it. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions described in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A training method for a hierarchical neural network for fluid simulation, characterized in that, Including: Obtaining multiple frames of sample video frames in a sample video, as well as corresponding supervised video frames in the sample video, and optical flow information between each of the multiple frames of sample video frames and the corresponding supervised video frames; Inputting the multiple frames of sample video frames into a hierarchical neural network to be trained, and determining background feature information and fluid feature information corresponding to each of the multiple frames of sample video frames; Adjusting the fluid feature information corresponding to each sample video frame based on the optical flow information corresponding to each sample video frame to obtain adjusted fluid feature information; Constructing a predicted video frame based on the background feature information and the adjusted fluid feature information, and training the hierarchical neural network to be trained based on the predicted video frame and the supervised video frame; Wherein, the training of the hierarchical neural network to be trained based on the predicted video frame and the supervised video frame includes: Constructing a first background supervision image based on the multiple frames of sample video frames and intermediate video frames between the multiple frames of sample video frames; and obtaining a pre-annotated second background supervision image, wherein the annotation information in the second background supervision image is used to annotate the background part that does not include the fluid part; Calculating a background prediction loss based on the first background supervision image and the background feature maps in the background feature information corresponding to each sample video frame; and Calculating a fluid prediction loss based on the second background supervision image and the transparency masks of the fluid feature maps in the fluid feature information corresponding to each sample video frame; and Calculating an image prediction loss based on the supervised video frame and the predicted video frame; Training the hierarchical neural network to be trained based on the background prediction loss, the fluid prediction loss, and the image prediction loss.

2. The method according to claim 1, characterized in that, The background feature information corresponding to the sample video frame includes a background feature map for characterizing the background part in the sample video frame and a first transparency mask of the background feature map; the fluid feature information corresponding to the sample video includes a fluid feature map for characterizing the fluid part in the sample video frame and a second transparency mask of the fluid feature map.

3. The method according to claim 1 or 2, characterized in that, The multiple frames of sample video frames include a first video frame and a second video frame, and the first video frame is before the second video frame; The constructing a predicted video frame based on the background feature information and the adjusted fluid feature information includes: Fusing the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame to determine fused fluid feature information; Constructing the predicted video frame based on the background feature information corresponding to the first video frame and the fused fluid feature information.

4. The method according to claim 3, wherein The fusing the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame to determine fused fluid feature information includes: Inputting the adjusted fluid feature information corresponding to the first video frame and the adjusted fluid feature information corresponding to the second video frame into an adjustment network to be trained to determine the fused fluid feature information; wherein, the adjustment network is trained following the hierarchical neural network.

5. A method for generating a fluid simulation animation, characterized in that, Including: Obtain an initial image and the optical flow information of the initial image; Input the initial image and the optical flow information of the initial image into a target hierarchical neural network. The target hierarchical neural network extracts the background feature information and fluid feature information of the initial image, adjusts the fluid feature information through the optical flow information, and constructs a predicted video frame of the target fluid in the initial image under the optical flow information based on the adjusted fluid feature information and the background feature information. Wherein, the target hierarchical neural network is trained by the training method of the hierarchical neural network for fluid simulation according to any one of claims 1 to 4; Construct a fluid simulation animation corresponding to the target fluid in the initial image based on the initial image and the predicted video frame.

6. The method according to claim 5, wherein After inputting the initial image and the optical flow information of the initial image into the target hierarchical neural network, the target hierarchical neural network is used to execute the following processing procedure: Determine the background feature information and fluid feature information corresponding to the initial image; Adjust the fluid feature information corresponding to the initial image based on the optical flow information to obtain the adjusted fluid feature information corresponding to the initial image; Fuse the fluid feature information corresponding to the initial image and the adjusted fluid feature information corresponding to the initial image to obtain the fused fluid feature information corresponding to the initial image; Construct a predicted video frame corresponding to the initial image based on the background feature information corresponding to the initial image and the fused fluid feature information corresponding to the initial image.

7. The method according to claim 5, wherein The optical flow information of the initial image includes the flow velocity and flow direction of each pixel point of the initial image; The obtaining of the initial image includes: Obtain an initial image, the depth information of the initial image, the key point optical flow marking information corresponding to the initial image, and the fluid mask image corresponding to the initial image; The method further includes determining the optical flow information of the initial image according to the following method: Construct a first three-dimensional grid map of the target scene corresponding to the initial image based on the depth information of the initial image; and, Determine the initial optical flow information of each pixel point of the initial image based on the key point optical flow marking information corresponding to the initial image; Determine the optical flow information of the initial image based on the first three-dimensional grid map and the initial optical flow information.

8. The method according to claim 7, wherein The determining of the optical flow information of the initial image based on the first three-dimensional grid map and the initial optical flow information includes: Preliminarily adjust the initial optical flow information based on the first three-dimensional grid map to obtain intermediate optical flow information; Input the intermediate optical flow information and the initial image into a pre-trained optical flow prediction network to determine the optical flow information of the initial image.

9. The method according to claim 7 or 8, characterized in that, The method further includes: Respond to a target virtual object addition instruction, and add a target virtual object at the target position of the initial image; Construct a second three-dimensional grid map based on the depth information of the initial image after adding the target virtual object; Re-determine the optical flow information of the initial image based on the second three-dimensional grid map and the initial optical flow information.

10. The method according to claim 9, characterized in that, The method further includes: Edit the target virtual object according to an edit instruction for the target virtual object; wherein the edit instruction is used to edit at least one of the following: size, position, angle, object type.

11. A training device for a hierarchical neural network for fluid simulation, characterized in that, Including: A first acquisition module, configured to acquire multiple sample video frames in a sample video, as well as supervised video frames corresponding to the multiple sample video frames in the sample video, and optical flow information between the multiple sample video frames and the supervised video frames respectively; A first determination module, configured to input the multiple sample video frames into a hierarchical neural network to be trained, and determine background feature information and fluid feature information corresponding to the multiple sample video frames respectively; An adjustment module, configured to adjust the fluid feature information corresponding to each sample video frame based on the optical flow information corresponding to each sample video frame to obtain adjusted fluid feature information; A training module, configured to construct a predicted video frame based on the background feature information and the adjusted fluid feature information, and train the hierarchical neural network to be trained based on the predicted video frame and the supervised video frame; wherein the training module is specifically configured to: Construct a first background supervision image based on the multiple sample video frames and intermediate video frames between the multiple sample video frames; and obtain a pre-annotated second background supervision image, wherein the annotation information in the second background supervision image is used to annotate the background part that does not include the fluid part; Calculate a background prediction loss based on the first background supervision image and the background feature maps in the background feature information corresponding to each sample video frame; and Calculate a fluid prediction loss based on the second background supervision image and the transparency masks of the fluid feature maps in the fluid feature information corresponding to each sample video frame; and Calculate an image prediction loss based on the supervised video frame and the predicted video frame; Train the hierarchical neural network to be trained based on the background prediction loss, the fluid prediction loss, and the image prediction loss.

12. A fluid simulation animation generation device, characterized in that, Including: A second acquisition module, configured to acquire an initial image and optical flow information of the initial image; A second determination module, configured to input the initial image and the optical flow information of the initial image into a target hierarchical neural network, extract background feature information and fluid feature information of the initial image by the target hierarchical neural network, adjust the fluid feature information through the optical flow information, and construct a predicted video frame of the target fluid in the initial image under the optical flow information based on the adjusted fluid feature information and the background feature information; wherein the target hierarchical neural network is trained by the training method of the hierarchical neural network for fluid simulation according to any one of claims 1 to 4; A construction module, configured to construct a fluid simulation animation corresponding to the target fluid in the initial image based on the initial image and the predicted video frame.

13. A computer device, characterized in that, Including: A processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the training method of the hierarchical neural network for fluid simulation according to any one of claims 1 to 4 are performed, or the steps of the fluid simulation animation generation method according to any one of claims 5 to 10 are performed.

14. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by the processor, the steps of the training method of the hierarchical neural network for fluid simulation according to any one of claims 1 to 4 are performed, or the steps of the fluid simulation animation generation method according to any one of claims 5 to 10 are performed.

Citation Information

Patent Citations

  • Model training method, video frame insertion method and corresponding device

    CN113542651A