A Dynamic Fluid Surface Reconstruction Method and System Based on the SwinLSTM Network
Through the method based on the SwinLSTM network, a simulated fluid data set is constructed and pre-processed. A dynamic fluid prediction network model is constructed in combination with the Swin-Transformer block and the LSTM unit, which solves the problem of space-time consistency in dynamic fluid surface reconstruction and realizes high-precision dynamic fluid surface reconstruction.
Patent Information
- Application Number
- CN202510553585.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Existing deep learning-based methods are difficult to effectively capture the consistency of dynamic fluid surfaces in time and space, resulting in limited accuracy of reconstruction of the reconstruction of the depth and normal field information.
Using a method based on SwinLSTM network, a dynamic fluid prediction network model is constructed by constructing a simulated fluid data set and preprocessing it, combining Swin-Transformer blocks and LSTM units to construct a dynamic fluid prediction network model, and a depth field and normal field prediction are predicted on the reference image and distorted image, and quantitatively evaluated with the real depth field information and normal field information to obtain the trained dynamic fluid prediction network model.
The reconstruction accuracy of the model on simulated and real dynamic fluid data sets is improved, and the three-dimensional morphology of the dynamic fluid surface can be accurately restored, ensuring the spatial and temporal consistency of the reconstruction results.
Smart Images

Figure CN120068739B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of dynamic fluid surface reconstruction, and in particular, to a method and system for dynamic fluid surface reconstruction based on the SwinLSTM network. Background Art
[0002] Modeling and reconstructing dynamic fluid surfaces from images is crucial in many scientific and engineering fields such as hydraulics, hydrodynamics, fluid simulation, and computer graphics. However, these fluid surfaces pose unique challenges. When light passes through the invisible air-fluid interface, it deviates from its original straight propagation path, resulting in severe distortion of the images captured by the camera. In addition, the time-varying fluctuations of the dynamic fluid surface make it more complex to extract reliable and stable image features. All these make it difficult to accurately recover a series of spatially and temporally consistent fluid surface shapes and motions. Currently, deep learning-based methods usually combine convolutional neural networks (CNNs) and recurrent neural networks (RNNs) for dynamic fluid surface reconstruction, but these methods often cannot effectively capture the global spatio-temporal consistency of dynamic fluids, resulting in limited accuracy in reconstructing the depth field and normal field information. Summary of the Invention
[0003] To solve the above technical problems, the purpose of the present invention is to provide a method and system for dynamic fluid surface reconstruction based on the SwinLSTM network, which can effectively capture the consistency information of fluid sequences in time and space, thereby improving the reconstruction accuracy of the model on simulated and real dynamic fluid data sets.
[0004] The first technical solution adopted by the present invention is: A method for dynamic fluid surface reconstruction based on the SwinLSTM network, comprising the following steps:
[0005] Simulate the fluid behavior based on a preset wave equation to construct a simulated fluid data set;
[0006] Preprocess the simulated fluid data set to obtain a reference image, a distorted image, true depth field information, and true normal field information;
[0007] Based on the Swin-Transformer block and the LSTM unit, construct a dynamic fluid prediction network model;
[0008] Predict the depth field and normal field for the reference image and the distorted image based on the dynamic fluid prediction network model, and conduct quantitative evaluation by combining the true depth field information and the true normal field information to obtain a trained dynamic fluid prediction network model;
[0009] Predict the depth field and normal field for the dynamic fluid based on the trained dynamic fluid prediction network model to obtain the dynamic fluid surface reconstruction result.
[0010] Furthermore, the step of constructing a simulated fluid data set based on a preset wave equation specifically includes:
[0011] Select the shallow water wave equation to simulate the fluid wave behavior under shallow water conditions to obtain the first fluid behavior simulation data;
[0012] Select the Grestner wave equation for fast Fourier transform and perform ocean wave simulation in computer graphics to obtain the second fluid behavior simulation data;
[0013] Select the Gaussian equation to simulate the water ripples with damping effect to obtain the third fluid behavior simulation data;
[0014] Perform weighted linear combination on the first fluid behavior simulation data, the second fluid behavior simulation data, and the third fluid behavior simulation data to construct a simulated fluid data set.
[0015] Furthermore, the expression of the shallow water wave equation is specifically as follows:
[0016] ;
[0017] In the above formula, represents the liquid density, represents the velocity of the fluid in the direction, represents the velocity of the fluid in the direction, represents the coordinate position of the Euler grid, represents the dynamic fluid surface height under the shallow water wave equation, represents the time point;
[0018] The expression of the Grestner wave equation is specifically as follows:
[0019] ;
[0020] In the above formula, represents the dynamic fluid surface height under the Gerstner wave equation, represents the Fourier amplitude, represents the imaginary unit, and are integers ranging from and where, and are the dimensions of the grid, represents the coordinate position of the Euler grid;
[0021] The expression of the Gaussian equation is specifically as follows:
[0022] ;
[0023] In the above formula, represents the dynamic fluid surface height under the Gaussian equation, represents the wave amplitude, represents the damping factor, represents the phase factor, represents time, represents the coordinate position of the Euler grid.
[0024] Furthermore, the step of preprocessing the simulated fluid data set to obtain the reference image, the distorted image, the true depth field information and the true normal field information specifically includes:
[0025] Determine the reference image and the true depth field information based on the simulated fluid data set, and obtain the air refractive index and the liquid refractive index;
[0026] Perform a partial derivative calculation on the true depth field information to obtain the true normal field information;
[0027] Combine the true normal field information, the air refractive index and the liquid refractive index, and perform single refraction according to the reference image to obtain the distorted image.
[0028] Furthermore, the dynamic fluid prediction network model specifically includes a graph partitioning layer, a graph embedding layer, a cyclic downsampling layer, a cyclic upsampling layer and a graph splicing layer. The graph partitioning layer, the graph embedding layer, the cyclic downsampling layer, the cyclic upsampling layer and the graph splicing layer are connected in sequence, where:
[0029] The cyclic downsampling layer includes a first SwinLSTM core prediction unit and a tile merging unit. The first SwinLSTM core prediction unit includes a first SwinTransformer block and a first LSTM unit. The first SwinTransformer block includes a first sub-module and a second sub-module. The first sub-module and the second sub-module have the same structure. Among them, the first sub-module includes a first layer normalization layer, a first sliding window multi-head attention mechanism layer, a first multi-layer perceptron layer and a second layer normalization layer;
[0030] The cyclic upsampling layer includes a second SwinLSTM core prediction unit and a tile expansion unit. The second SwinLSTM core prediction unit includes a second SwinTransformer block and a second LSTM unit. The second SwinTransformer block includes a third sub-module and a fourth sub-module. The third sub-module and the fourth sub-module have the same structure. Among them, the third sub-module includes a third layer normalization layer, a second sliding window multi-head attention mechanism layer, a second multi-layer perceptron layer and a fourth layer normalization layer.
[0031] Furthermore, the expression of the SwinLSTM core prediction unit is specifically as follows:
[0032] ;
[0033] In the above formula, represents the fluid frame information at time represents the hidden state information of the fluid data at time represents the hidden state information of the fluid data at time represents the cell state information of the fluid data at time represents the cell state information of the fluid data at time represents the intermediate variable information related to the fluid frame at time represents the linear mapping layer, represents the concatenation of two tensors along the channel, represents the Swin-Transformer block, represents the hyperbolic tangent activation function,
[0034] represents the sigmoid activation function,
[0035] Furthermore, the expression of the SwinTransformer block is specifically as follows:
[0036] ; In the above formula, represents the feature information of the th layer, represents the feature information of the th layer, represents the feature information of the th layer, represents the intermediate feature information of the th layer, represents the intermediate feature information of the th layer, represents the layer normalization layer, represents the sliding window multi-head attention mechanism layer, represents the multi-layer perceptron layer.
[0037] Specifically, the expression of the loss function of the dynamic fluid prediction network model is specifically as follows:
[0038] ;
[0039] In the above formula, represents the loss function of the dynamic fluid prediction network model, represents the mean squared error loss function, represents the structural similarity loss function, represents the scale consistency loss function.
[0040] Specifically, the step of predicting the depth field and the normal field for the reference image and the distorted image based on the dynamic fluid prediction network model, and performing quantitative evaluation by combining the real depth field information and the real normal field information to obtain the trained dynamic fluid prediction network model specifically includes:
[0041] Input the reference image and the distorted image into the dynamic fluid prediction network model for training;
[0042] Based on the graph partition layer of the dynamic fluid prediction network model, perform partition processing on the reference image and the distorted image to obtain non-overlapping linear small patches;
[0043] Based on the graph embedding layer of the dynamic fluid prediction network model, flatten and perform feature mapping on the non-overlapping linear small patches to obtain linear feature small patches;
[0044] Obtain the current hidden state and the current cell state according to the linear feature small patches;
[0045] Based on the recurrent downsampling layer of the dynamic fluid prediction network model, perform recurrent downsampling on the linear feature small patches, the current hidden state and the current cell state to obtain the hidden state at the next moment and the cell state at the next moment;
[0046] Based on the recurrent upsampling layer of the dynamic fluid prediction network model, perform upsampling on the hidden state at the next moment and the cell state at the next moment to obtain the upsampled hidden state and the upsampled cell state;
[0047] Based on the graph stitching layer of the dynamic fluid prediction network model, stitch the upsampled hidden state and the upsampled cell state to obtain the predicted depth field information and the predicted normal field information;
[0048] Through the structural similarity index and the peak signal-to-noise ratio index, combine the real depth field information and the real normal field information to perform quantitative evaluation on the predicted depth field information and the predicted normal field information to obtain a quantitative evaluation result;
[0049] If the quantitative evaluation result does not meet the preset requirements, retrain the dynamic fluid prediction network model until the preset requirements are met, and output the trained dynamic fluid prediction network model.
[0050] The second technical solution adopted by the present invention is: a dynamic fluid surface reconstruction system based on the SwinLSTM network, including:
[0051] The first module is used to simulate the fluid behavior based on a preset wave equation and construct a simulation fluid data set;
[0052] The second module is used to preprocess the simulation fluid data set to obtain a reference image, a distorted image, real depth field information and real normal field information;
[0053] The third module is used to construct a dynamic fluid prediction network model based on the Swin-Transformer block and the LSTM unit;
[0054] The fourth module is used to predict the depth field and the normal field of the reference image and the distorted image based on the dynamic fluid prediction network model, and perform quantitative evaluation by combining the real depth field information and the real normal field information to obtain a trained dynamic fluid prediction network model;
[0055] The fifth module is used to predict the depth field and the normal field of the dynamic fluid based on the trained dynamic fluid prediction network model to obtain the dynamic fluid surface reconstruction result.
[0056] The beneficial effects of the method and system of the present invention are as follows: First, the present invention simulates the fluid behavior based on a preset wave equation, constructs a simulation fluid data set and performs preprocessing to obtain a reference image, a distorted image, real depth field information and real normal field information. A method of combining and superimposing three different forms of wave equations is used to simulate various types of fluids, and a dynamic fluid simulation data set containing rich samples is constructed. Then, based on the Swin-Transformer block and the LSTM unit, a dynamic fluid prediction network model is constructed. By replacing the convolutional layer in the traditional convolutional neural network with a self-attention mechanism and introducing the designed SwinLSTM core prediction module, this network solves the problem that the traditional convolutional neural network ignores the spatio-temporal consistency of the fluid sequence when learning local spatial information, enabling the network to effectively capture the spatio-temporal consistency information of the fluid sequence. Further, based on the dynamic fluid prediction network model, the depth field and the normal field of the reference image and the distorted image are predicted, and quantitative evaluation is performed by combining the real depth field information and the real normal field information. Finally, based on the trained dynamic fluid prediction network model, the depth field and the normal field of the dynamic fluid are predicted, which improves the reconstruction accuracy of the model on the simulation and real dynamic fluid data sets, can make full use of the prediction information of the dynamic fluid feed-forward frames, ensures the spatio-temporal consistency of the dynamic fluid reconstruction result, and can accurately restore its surface three-dimensional morphology. Description of the Drawings
[0057] Figure 1It is the flowchart of the steps of a dynamic fluid surface reconstruction method based on the SwinLSTM network according to the present invention;
[0058] Figure 2 It is the structural block diagram of a dynamic fluid surface reconstruction system based on the SwinLSTM network according to the present invention;
[0059] Figure 3 It is the schematic diagram of the simulation fluid data set provided by the specific embodiment of the present invention;
[0060] Figure 4 It is the structural schematic diagram of the dynamic fluid prediction network model provided by the specific embodiment of the present invention;
[0061] Figure 5 It is the structural schematic diagram of the SwinLSTM core prediction unit provided by the specific embodiment of the present invention;
[0062] Figure 6 It is the structural schematic diagram of the Swin-Transfomer block provided by the specific embodiment of the present invention;
[0063] Figure 7 It is the schematic diagram of the prediction effect on the real fluid data provided by the specific embodiment of the present invention. Detailed implementation manners
[0064] The following further elaborates the present invention in detail with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0065] Refer to Figure 1 , the present invention provides a dynamic fluid surface reconstruction method based on the SwinLSTM network, and the method includes the following steps:
[0066] S100. Simulate the fluid behavior based on a preset wave equation to construct a simulation fluid data set;
[0067] S110. Select the shallow water wave equation to simulate the fluid wave behavior under shallow water conditions to obtain the first fluid behavior simulation data;
[0068] In this embodiment, the shallow water wave equation is a set of partial differential equations derived from the Navier-Stokes equation. This equation describes the incompressible property of the fluid, including mass conservation and linear momentum conservation. The shallow water wave equation is as follows:
[0069] ;
[0070] In the above formula, represents the liquid density, represents the velocity of the fluid in the direction, represents the velocity of the fluid in the direction, represents the coordinate position of the Eulerian grid, represents the dynamic fluid surface height under the shallow water wave equation, represents the time point.
[0071] S120. Select the Gerstner wave equation for fast Fourier transform and perform wave simulation in computer graphics to obtain the second fluid behavior simulation data;
[0072] In this embodiment, the Gerstner wave equation is widely used in wave simulation in computer graphics. The embodiment of the present invention uses it to simulate the movement of fluids with relatively large volumes and calculates the Gerstner equation in the form of fast Fourier transform (FFT). The equation representation based on FFT is:
[0073] ;
[0074] In the above formula, represents the dynamic fluid surface height under the Gerstner wave equation, represents the Fourier amplitude, represents the imaginary unit, and are integers ranging from and where, and are the dimensions of the grid, represents the coordinate position of the Eulerian grid.
[0075] S130. Select the Gaussian equation to simulate the water ripple with damping effect to obtain the third fluid behavior simulation data;
[0076] In this embodiment, the Gaussian equation is used to simulate the water ripple with damping effect, represents the dynamic fluid surface height, and the Gaussian equation representation is:
[0077] ;
[0078] In the above formula, represents the dynamic fluid surface height under the Gaussian equation, represents the wave amplitude, represents the damping factor, represents the phase factor, represents the time, Represents the coordinate position of the Eulerian grid.
[0079] S140. Perform a weighted linear combination of the first fluid behavior simulation data, the second fluid behavior simulation data, and the third fluid behavior simulation data to construct a simulated fluid data set.
[0080] In this embodiment, a weighted linear combination of the first fluid behavior simulation data, the second fluid behavior simulation data, and the third fluid behavior simulation data is performed, and its expression is specifically as follows:
[0081] ;
[0082] In the above formula, represents the dynamic fluid surface height under the combined equation, represents the dynamic fluid surface height under the shallow water wave equation, represents the dynamic fluid surface height under the Gerstner wave equation, represents the dynamic fluid surface height under the Gaussian equation, and represent the weight coefficients.
[0083] Furthermore, it should be noted that by superimposing different weight coefficient ratios, various physically simulated fluid effects can be obtained to construct multiple sets of fluid sequence data sets, where the Eulerian grid size of all simulated fluids is set to , and the simulated fluid data set constructed by the combined equation is as shown in Figure 3 (a) therein, including a 3-channel undistorted reference variable image, a 3-channel distorted image, a 1-channel depth field image, and a 3-channel normal field image.
[0084] S200. Preprocess the simulated fluid data set to obtain a reference image, a distorted image, real depth field information, and real normal field information;
[0085] Specifically, based on the simulated fluid data set, determine the reference image and the real depth field information, and obtain the air refractive index and the liquid refractive index; perform a partial derivative calculation on the real depth field information to obtain the real normal field information; combine the real normal field information, the air refractive index, and the liquid refractive index, and perform a single refraction according to the reference image to obtain the distorted image.
[0086] In this embodiment, the schematic diagram of the data set construction principle is as shown in Figure 3 (b) therein. First, the height field information of different dynamic fluid surfaces can be simulated through the above combined equation. By taking the partial derivative of the height field information, the fluid surface normal field information can be further solved. By combining the normal field with the refractive index of the first medium and the refractive index of the second medium , the distorted image sequence after single refraction can be calculated from the undistorted reference image below the liquid surface.
[0087] Furthermore, it should be noted that 12,000 pairs of data (network input data and network true label data) are constructed from the above combined equations for training and testing. Among them, the total size of the network input data is . Among them, 12,000 represents the total number of samples, 10 represents the number of adjacent frames in the fluid sequence, 256 represents the length and width dimensions of the input tensor, and 6 represents the concatenation of the 3-channel distorted image and the 3-channel undistorted reference image in the channel dimension. Among them, the total size of the network true label data is . Among them, 12,000 represents the total number of samples, 10 represents the number of adjacent frames in the fluid sequence, 256 represents the length and width dimensions of the output tensor, and 4 represents the concatenation of the 1-channel predicted depth field and the 3-channel predicted normal field in the channel dimension.
[0088] Then, the data set is divided into a training set and a test set, where the training set accounts for 80% and the test set accounts for 20%. Among them, the total network input size of the training set is , and the total network label size of the training set is , the total network input size of the test set is , and the total network label size of the test set is .
[0089] S300. Build a dynamic fluid prediction network model based on the Swin-Transformer block and the LSTM unit;
[0090] Specifically, use the Pytorch deep learning framework to build an end-to-end network model based on Swin-Transformer and LSTM for dynamic fluid depth field and normal field prediction. As Figure 4 shown, the dynamic fluid prediction network model specifically includes a graph partitioning layer, a graph embedding layer, a recurrent downsampling layer, a recurrent upsampling layer, and a graph concatenation layer. The graph partitioning layer, the graph embedding layer, the recurrent downsampling layer, the recurrent upsampling layer, and the graph concatenation layer are connected in sequence, where:
[0091] Both the recurrent downsampling layer and the recurrent upsampling layer have N layers.
[0092] The cyclic downsampling layer includes a first SwinLSTM core prediction unit and a tile merging unit. The first SwinLSTM core prediction unit includes a first SwinTransformer block and a first LSTM unit. The first SwinTransformer block includes a first sub-module and a second sub-module, and the first sub-module and the second sub-module have the same structure. Among them, the first sub-module includes a first layer normalization layer, a first sliding window multi-head attention mechanism layer, a first multi-layer perceptron layer, and a second layer normalization layer.
[0093] The cyclic upsampling layer includes a second SwinLSTM core prediction unit and a tile expansion unit. The second SwinLSTM core prediction unit includes a second SwinTransformer block and a second LSTM unit. The second SwinTransformer block includes a third sub-module and a fourth sub-module, and the third sub-module and the fourth sub-module have the same structure. Among them, the third sub-module includes a third layer normalization layer, a second sliding window multi-head attention mechanism layer, a second multi-layer perceptron layer, and a fourth layer normalization layer.
[0094] S400. Perform depth field and normal field prediction on the reference image and the distorted image based on the dynamic fluid prediction network model, and conduct quantitative evaluation by combining the real depth field information and the real normal field information to obtain the trained dynamic fluid prediction network model;
[0095] Specifically, first, the input distorted images and reference images of 10 groups of fluid sequence frames are divided into non-overlapping linear small tiles through the graph partitioning layer, and after flattening each small tile, feature embedding is performed to map it to any dimension C (C represents the number of channels). Then it is sent into the SwinLSTM core prediction unit in the cyclic downsampling module. This core prediction unit accepts 3 input elements: the feature tile mapped to any dimension C, the hidden state and the cell state . And it outputs two elements: the hidden state and the cell state . Among them is copied into two copies, one for the subsequent cyclic upsampling module, and the other is one of the inputs for the next upsampling loop module. The cell state is then passed into the tile expansion unit, and the obtained output and are used as the inputs for the next round of the cyclic upsampling module. After performing N rounds of downsampling loops, the obtained output and the copied hidden state vector group are used as the inputs for the cyclic upsampling module. Similarly, after performing N rounds of upsampling loops, the obtained output is used as the input for the final graph stitching layer, and thus the depth field and normal field information of the current 10-frame continuous fluid sequence frames can be obtained.
[0096] S410. Input the reference image and the distorted image into the dynamic fluid prediction network model for training;
[0097] S420. Based on the graph partitioning layer of the dynamic fluid prediction network model, partition the reference image and the distorted image to obtain non-overlapping linear small patches;
[0098] S430. Based on the graph embedding layer of the dynamic fluid prediction network model, flatten and perform feature mapping on the non-overlapping linear small patches to obtain linear feature small patches;
[0099] S440. Obtain the current hidden state and the current cell state according to the linear feature small patches;
[0100] S450. Based on the recurrent downsampling layer of the dynamic fluid prediction network model, perform recurrent downsampling on the linear feature small patches, the current hidden state and the current cell state to obtain the hidden state at the next moment and the cell state at the next moment;
[0101] S460. Based on the recurrent upsampling layer of the dynamic fluid prediction network model, perform upsampling on the hidden state at the next moment and the cell state at the next moment to obtain the upsampled hidden state and the upsampled cell state;
[0102] First of all, it should be noted that the structures and functions of the SwinLSTM core prediction units and the Swin-Transformer blocks in the recurrent downsampling layer and the recurrent upsampling layer are the same, and the embodiments of the present invention are described and explained uniformly.
[0103] In this embodiment, the SwinLSTM core prediction unit is as Figure 5 shown. In the SwinLSTM unit, the unit state and the hidden state are updated in the horizontal direction to capture the temporal dependencies of dynamic fluid data in the long and short stages. At the same time, the Swin-Transformer block learns the global spatial dependencies of dynamic fluid data in the vertical direction. The formula description of the SwinLSTM core prediction unit is as follows:
[0104] ;
[0105] In the above formula, represents the fluid frame information at time , represents the hidden state information of the fluid data at time , represents the hidden state information of the fluid data at time , represents the cell state information of the fluid data at time . Represents The cell state information of the fluid data at a moment And Intermediate variable information related to the fluid frame at a moment Represents a linear mapping layer Represents the concatenation of two tensors along the channel Represents a Swin-Transformer block Represents the hyperbolic tangent activation function Represents the sigmoid activation function Represents the Hadamard product, which is an operation of multiplying the corresponding elements of two matrices
[0106] Furthermore, it should be noted that the Swin-Transformer block is as Figure 6 shown. In the Swin-Transformer block, it consists of two sub-modules. The first sub-module consists of two layer normalization layers, a sliding window multi-head attention mechanism layer, and a multi-layer perceptron layer. A layer normalization operation and a residual connection operation are respectively applied before and after each sliding window multi-head attention mechanism layer and multi-layer perceptron layer; the second sub-module is exactly the same as the first sub-module. The formula description of the Swin-Transformer block is as follows:
[0107] ;
[0108] In the above formula, Represents the feature information of the th layer Represents the feature information of the th layer Represents the feature information of the th layer Represents the intermediate feature information of the th layer Represents the intermediate feature information of the th layer Represents a layer normalization layer Represents a sliding window multi-head attention mechanism layer Represents a multi-layer perceptron layer
[0109] S470. The graph concatenation layer based on the dynamic fluid prediction network model concatenates the upsampled hidden state and the upsampled cell state to obtain the predicted depth field information and the predicted normal field information;
[0110] S480. By using the structural similarity index and the peak signal-to-noise ratio index, combining the real depth field information and the real normal field information, quantitatively evaluate the predicted depth field information and the predicted normal field information to obtain the quantitative evaluation result;
[0111] Specifically, the performance of the model is quantitatively evaluated using the constructed simulation test set data. The SSIM (Structural Similarity Index) and PSNR (Peak Signal-to-Noise Ratio) are selected as the quantitative evaluation indicators to judge the quality of the depth field and normal field recovery of the simulated fluid data by the trained model. The specific formulas are as follows:
[0112] ;
[0113] In the above formula, and respectively represent the average values of images and , and respectively represent the variances of images and , represents the covariance of images and , represents the minimum constant introduced for stability, represents the possible maximum pixel value in the image, and the maximum value is 1 for the normalized data, represents the mean square error between the predicted image and the label image.
[0114] S490. If the quantitative evaluation result does not meet the preset requirements, the dynamic fluid prediction network model is retrained until the preset requirements are met, and the trained dynamic fluid prediction network model is output.
[0115] Specifically, the constructed Swin-Transformer and LSTM are instantiated, and the model is trained using the training set constructed in step S100. After training, the network model parameters are saved.
[0116] It should also be noted that according to the particularity of the dynamic fluid data set, a corresponding loss function is designed, and its expression is specifically as follows:
[0117] ;
[0118] In the above formula, represents the loss function of the dynamic fluid prediction network model, represents the mean square error loss function, which is used to ensure the basic convergence of the network, represents the structural similarity loss function, which is used to ensure the structural similarity of the depth field and normal field information, thus accelerating the convergence of the network, represents the scale consistency loss function, which is used to ensure the scale consistency of the depth field and normal field of adjacent frames in the fluid sequence prediction.
[0119] S500. Predict the depth field and normal field of the dynamic fluid based on the trained dynamic fluid prediction network model to obtain the dynamic fluid surface reconstruction result.
[0120] Furthermore, as shown in Table 1, the performance comparison of the method proposed in the embodiments of the present invention with common CNN and RNN structure networks under the above quantization metrics is as follows;
[0121] Table 1 Comparison data table
[0122]
[0123] According to Table 1, it can be seen that the Swin-Transformer and LSTM network structures proposed by the present invention perform significantly better than the network structures designed based on CNN and RNN in the above various quantization metrics. It shows that the performance of this method is good in the dynamic fluid three-dimensional reconstruction task.
[0124] Furthermore, capture the dynamic fluid dataset in the real environment, and use the real dataset to qualitatively verify the model performance. The qualitative evaluation results are as Figure 7 shown, Figure 7 The first column is the sequence of distorted images of the dynamic fluid captured in the real environment, Figure 7 The second column is the undistorted reference image placed at the bottom of the dynamic fluid, Figure 7 The third column is the depth map prediction result of the model proposed by the present invention on the real fluid data, Figure 7 The fourth column is the normal map prediction result of the model proposed by the present invention on the real fluid data, Figure 7 The fifth column is the three-dimensional visualization display of the depth map prediction result.
[0125] In summary, in the embodiments of the present invention, a simulation fluid training dataset is first constructed using a combined wave equation and divided into a training set and a test set. To solve the problem that traditional convolutional neural networks ignore the spatio-temporal consistency of fluid sequences when learning local spatial information, the present invention designs an end-to-end neural network structure integrating Swin-Transformer and LSTM. Specifically, the input of the network consists of multiple frames of reference images and distorted images of the fluid sequence, and the output is the depth field and normal field information of the dynamic fluid surface. The present invention replaces the convolutional layer in the traditional convolutional neural network with a self-attention mechanism and introduces the designed SwinLSTM core prediction module, enabling the network to effectively capture the consistency information of the fluid sequence in time and space. Experimental results show that the method proposed in the present invention significantly improves the reconstruction accuracy of the model on simulation and real dynamic fluid datasets. First, the network is trained on the constructed simulation fluid training set, then quantitatively evaluated on the test dataset, and finally qualitatively evaluated on the constructed real dynamic fluid dataset. The trained model is used to predict the real fluid distorted image sequence and the reference image sequence to achieve the three-dimensional reconstruction of the real dynamic fluid surface, which has broad application potential in related fields such as the three-dimensional reconstruction of dynamic fluid surfaces.
[0126] Referring to Figure 2 , a dynamic fluid surface reconstruction system based on a SwinLSTM network, comprising:
[0127] A first module 201 for simulating fluid behavior based on a preset wave equation to construct a simulation fluid dataset;
[0128] A second module 202 for preprocessing the simulation fluid dataset to obtain reference images, distorted images, real depth field information, and real normal field information;
[0129] A third module 203 for constructing a dynamic fluid prediction network model based on Swin-Transformer blocks and LSTM cells;
[0130] A fourth module 204 for predicting the depth field and normal field of the reference images and distorted images based on the dynamic fluid prediction network model, and performing quantitative evaluation by combining the real depth field information and the real normal field information to obtain a trained dynamic fluid prediction network model;
[0131] A fifth module 205 for predicting the depth field and normal field of the dynamic fluid based on the trained dynamic fluid prediction network model to obtain a dynamic fluid surface reconstruction result.
[0132] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0133] The above is a specific description of the preferred embodiments of the present invention. However, the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention. These equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A dynamic fluid surface reconstruction method based on the SwinLSTM network, characterized in that, It includes the following steps: Select the shallow water wave equation to simulate the fluid wave behavior under shallow water conditions, and obtain the first fluid behavior simulation data; Select the Gerstner wave equation for fast Fourier transform and perform ocean wave simulation in computer graphics to obtain the second fluid behavior simulation data; Select the Gaussian equation to simulate the water ripple with damping effect, and obtain the third fluid behavior simulation data; Perform weighted linear combination on the first fluid behavior simulation data, the second fluid behavior simulation data and the third fluid behavior simulation data to construct a simulated fluid data set; Determine the reference image and the true depth field information based on the simulated fluid data set, and obtain the air refractive index and the liquid refractive index; Perform partial derivative calculation on the true depth field information to obtain the true normal field information; Combine the true normal field information, the air refractive index and the liquid refractive index, and perform single refraction according to the reference image to obtain a distorted image; Based on the Swin-Transformer block and the LSTM unit, construct a dynamic fluid prediction network model; Based on the dynamic fluid prediction network model, predict the depth field and the normal field of the reference image and the distorted image, and perform quantitative evaluation by combining the true depth field information and the true normal field information to obtain the trained dynamic fluid prediction network model; Based on the trained dynamic fluid prediction network model, predict the depth field and the normal field of the dynamic fluid to obtain the dynamic fluid surface reconstruction result.
2. The dynamic fluid surface reconstruction method based on the SwinLSTM network according to claim 1, wherein, The expression of the shallow water wave equation is specifically as follows: ; In the above formula, represents the liquid density, represents the velocity of the fluid in the direction, represents the velocity of the fluid in the direction, and represent the coordinate positions of the Eulerian grid, represents the dynamic fluid surface height under the shallow water wave equation, represents time; The expression of the Gerstner wave equation is specifically as follows: ; In the above formula, represents the dynamic fluid surface height under the Gerstner wave equation, represents the Fourier amplitude, represents the imaginary unit, and are integers ranging from to and are the dimensions of the grid, 、 represent the coordinate positions of the Eulerian grid; The expression of the Gaussian equation is specifically as follows: ; In the above formula, represents the dynamic fluid surface height under the Gaussian equation, represents the wave amplitude, represents the damping factor, represents the phase factor, represents time, and represent the coordinate positions of the Eulerian grid.
3. The dynamic fluid surface reconstruction method based on the SwinLSTM network according to claim 2, wherein The dynamic fluid prediction network model specifically includes a graph partitioning layer, a graph embedding layer, a cyclic downsampling layer, a cyclic upsampling layer and a graph splicing layer. The graph partitioning layer, the graph embedding layer, the cyclic downsampling layer, the cyclic upsampling layer and the graph splicing layer are connected in sequence, where: The cyclic downsampling layer includes a first SwinLSTM core prediction unit and a graph block merging unit. The first SwinLSTM core prediction unit includes a first SwinTransformer block and a first LSTM unit. The first SwinTransformer block includes a first sub-module and a second sub-module. The first sub-module and the second sub-module have the same structure. Among them, the first sub-module includes a first layer normalization layer, a first sliding window multi-head attention mechanism layer, a first multi-layer perceptron layer and a second layer normalization layer; The cyclic upsampling layer includes a second SwinLSTM core prediction unit and a graph block expansion unit. The second SwinLSTM core prediction unit includes a second SwinTransformer block and a second LSTM unit. The second SwinTransformer block includes a third sub-module and a fourth sub-module. The third sub-module and the fourth sub-module have the same structure. Among them, the third sub-module includes a third layer normalization layer, a second sliding window multi-head attention mechanism layer, a second multi-layer perceptron layer and a fourth layer normalization layer.
4. The dynamic fluid surface reconstruction method based on the SwinLSTM network according to claim 3, wherein, The expression of the SwinLSTM core prediction unit is specifically as follows: ; In the above formula, represents the fluid frame information at the moment, represents the fluid data hiding state information at the moment, represents the fluid data hiding state information at the moment, represents the fluid data cell state information at the moment, represents the fluid data cell state information at the moment, represents and the intermediate variable information related to the fluid frame at the moment, represents the linear mapping layer, represents the concatenation of two tensors along the channel, represents the Swin-Transformer block, represents the hyperbolic tangent activation function, represents the sigmoid activation function, represents the Hadamard product, which represents the operation of multiplying the corresponding elements of two matrices.
5. The dynamic fluid surface reconstruction method based on the SwinLSTM network according to claim 4, characterized in that, The expression of the SwinTransformer block is specifically as follows: ; In the above formula, represents the feature information of the th layer, represents the feature information of the th layer, represents the feature information of the th layer, represents the intermediate feature information of the th layer, represents the intermediate feature information of the th layer, represents the layer normalization layer, represents the sliding window multi-head attention mechanism layer, represents the multi-layer perceptron layer.
6. The dynamic fluid surface reconstruction method based on the SwinLSTM network according to claim 5, wherein The expression of the loss function of the dynamic fluid prediction network model is specifically as follows: ; In the above formula, represents the loss function of the dynamic fluid prediction network model, represents the mean squared error loss function, represents the structural similarity loss function, represents the scale consistency loss function.
7. The dynamic fluid surface reconstruction method based on the SwinLSTM network according to claim 6, wherein The step of using the dynamic fluid prediction network model to predict the depth field and normal field of the reference image and the distorted image, and combining the real depth field information and the real normal field information for quantitative evaluation to obtain the trained dynamic fluid prediction network model specifically includes: Input the reference image and the distorted image into the dynamic fluid prediction network model for training; Based on the graph partition layer of the dynamic fluid prediction network model, perform partition processing on the reference image and the distorted image to obtain non-overlapping linear small patches; Based on the graph embedding layer of the dynamic fluid prediction network model, flatten and feature map the non-overlapping linear small patches to obtain linear feature small patches; Obtain the current hidden state and the current cell state according to the linear feature small patches; Based on the recurrent downsampling layer of the dynamic fluid prediction network model, perform recurrent downsampling on the linear feature small patches, the current hidden state and the current cell state to obtain the hidden state at the next moment and the cell state at the next moment; Based on the recurrent upsampling layer of the dynamic fluid prediction network model, perform upsampling on the hidden state at the next moment and the cell state at the next moment to obtain the upsampled hidden state and the upsampled cell state; Based on the graph concatenation layer of the dynamic fluid prediction network model, concatenate the upsampled hidden state and the upsampled cell state to obtain the predicted depth field information and the predicted normal field information; Through the structural similarity index and the peak signal-to-noise ratio index, combine the real depth field information and the real normal field information to quantitatively evaluate the predicted depth field information and the predicted normal field information to obtain a quantitative evaluation result; If the quantitative evaluation result does not meet the preset requirements, retrain the dynamic fluid prediction network model until the preset requirements are met, and output the trained dynamic fluid prediction network model.
8. A dynamic fluid surface reconstruction system based on the SwinLSTM network, characterized in that, It includes the following modules: The first module is used to select the shallow water wave equation to simulate the fluid wave behavior under shallow water conditions to obtain the first fluid behavior simulation data; Select the Gerstner wave equation for fast Fourier transform and perform ocean wave simulation in computer graphics to obtain the second fluid behavior simulation data; Select the Gaussian equation to simulate the water ripple with damping effect to obtain the third fluid behavior simulation data; Perform weighted linear combination on the first fluid behavior simulation data, the second fluid behavior simulation data and the third fluid behavior simulation data to construct a simulation fluid data set; The second module is used to determine the reference image and the real depth field information based on the simulation fluid data set, and obtain the air refractive index and the liquid refractive index; Perform partial derivative calculation on the real depth field information to obtain the real normal field information; Combine the real normal field information, the air refractive index and the liquid refractive index, and perform single refraction according to the reference image to obtain the distorted image; The third module is used to build a dynamic fluid prediction network model based on Swin-Transformer blocks and LSTM cells; The fourth module is used to perform depth field and normal field predictions on the reference image and the distorted image based on the dynamic fluid prediction network model, and conduct quantitative evaluation by combining the real depth field information and the real normal field information to obtain the trained dynamic fluid prediction network model; The fifth module is used to perform depth field and normal field predictions on the dynamic fluid based on the trained dynamic fluid prediction network model to obtain the dynamic fluid surface reconstruction result.
Citation Information
Patent Citations
Liquid surface three-dimensional reconstruction method and system based on data driving
CN117993302A
Hydromechanics partial differential equation solving device based on autoregressive neural network
CN118586311A