Dynamic fluid surface reconstruction method and system based on SwinLSTM network

By adopting the SwinLSTM network in dynamic fluid surface reconstruction, combining the Swin-Transformer block and the LSTM unit, the space-time consistency information of dynamic fluid is captured, and the problem of limited reconstruction accuracy in the prior art is solved, and higher reconstruction accuracy and space-time consistency are achieved.

CN120068739AActive Publication Date: 2025-05-30FOSHAN UNIVERSITY

Patent Information

Application Number
CN202510553585.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the global spatiotemporal consistency of dynamic fluids, resulting in limited accuracy of reconstruction of depth field and normal field information of dynamic fluid surface reconstruction.

Method used

Using a dynamic fluid surface reconstruction method based on SwinLSTM network, a dynamic fluid prediction network model is constructed, and a Swin-Transformer block and LSTM unit is used to capture the consistency information of the fluid sequence in time and space.

Benefits of technology

The reconstruction accuracy of the model on simulated and real dynamic fluid data sets is improved, and the three-dimensional morphology of the dynamic fluid surface can be accurately restored and the space-time consistency is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068739A_ABST
    Figure CN120068739A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic fluid surface reconstruction method and system based on a SwinLSTM network, and the method comprises the steps: carrying out the fluid behavior simulation based on a preset wave equation, constructing a simulation fluid data set, and obtaining a reference image, a distortion image, real depth field information and real normal field information; the method comprises the following steps: constructing a dynamic fluid prediction network model on the basis of a Swin-Transform block and an LSTM (Long Short Term Memory) unit; performing depth field and normal field prediction on the reference image and the distorted image, and performing quantitative evaluation in combination with real depth field information and real normal field information to obtain a trained dynamic fluid prediction network model; and performing depth field and normal field prediction to obtain a dynamic fluid surface reconstruction result. According to the method, the consistency information of the fluid sequence in time and space can be effectively captured, so that the reconstruction precision of the model in simulation and real dynamic fluid data sets is improved. The dynamic fluid surface reconstruction method and system based on the SwinLSTM network can be widely applied to the technical field of dynamic fluid surface reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dynamic fluid surface reconstruction, and in particular, to a method and system for dynamic fluid surface reconstruction based on a SwinLSTM network. Background Art

[0002] Modeling and reconstructing dynamic fluid surfaces from images is crucial in many scientific and engineering fields such as hydraulics, hydrodynamics, fluid simulation, and computer graphics. However, these fluid surfaces pose unique challenges. When light passes through the invisible air-fluid interface, it deviates from its original straight-line propagation path, resulting in severe distortion of the images captured by the camera. In addition, the time-varying fluctuations of the dynamic fluid surface make it more complex to extract reliable and stable image features. All these make it difficult to accurately recover a series of spatially and temporally consistent fluid surface shapes and motions. Currently, deep learning-based methods usually combine convolutional neural networks (CNNs) and recurrent neural networks (RNNs) for dynamic fluid surface reconstruction, but these methods often cannot effectively capture the global spatio-temporal consistency of dynamic fluids, resulting in limited accuracy in reconstructing the depth field and normal field information. Summary of the Invention

[0003] In order to solve the above technical problems, the purpose of the present invention is to provide a method and system for dynamic fluid surface reconstruction based on a SwinLSTM network, which can effectively capture the consistency information of fluid sequences in time and space, thereby improving the reconstruction accuracy of the model on simulation and real dynamic fluid data sets.

[0004] The first technical solution adopted by the present invention is: A method for dynamic fluid surface reconstruction based on a SwinLSTM network, comprising the following steps: Simulate fluid behavior based on a preset wave equation to construct a simulated fluid data set; Preprocess the simulated fluid data set to obtain reference images, distorted images, true depth field information, and true normal field information; Construct a dynamic fluid prediction network model based on Swin-Transformer blocks and LSTM units; Predict the depth field and normal field for the reference image and the distorted image based on the dynamic fluid prediction network model, and perform quantitative evaluation by combining the true depth field information and the true normal field information to obtain a trained dynamic fluid prediction network model; Predict the depth field and normal field for the dynamic fluid based on the trained dynamic fluid prediction network model to obtain the dynamic fluid surface reconstruction result.

[0005] Further, the step of constructing the simulated fluid data set based on the preset wave equation specifically includes: Select the shallow water wave equation to simulate the fluid wave behavior under shallow water conditions, and obtain the first fluid behavior simulation data; Select the Grestner wave equation to perform a fast Fourier transform and conduct a sea wave simulation in computer graphics to obtain the second fluid behavior simulation data; Select the Gaussian equation to simulate the water ripple with a damping effect to obtain the third fluid behavior simulation data; Perform a weighted linear combination of the first fluid behavior simulation data, the second fluid behavior simulation data, and the third fluid behavior simulation data to construct a simulated fluid data set.

[0006] Furthermore, the expression of the shallow water wave equation is specifically as follows: ; In the above formula, represents the liquid density, represents the velocity of the fluid in the direction, represents the velocity of the fluid in the direction, represents the coordinate position of the Euler grid, represents the dynamic fluid surface height under the shallow water wave equation, represents the time point; The expression of the Grestner wave equation is specifically as follows: ; In the above formula, represents the dynamic fluid surface height under the Gerstner wave equation, represents the Fourier amplitude, represents the imaginary unit, and are integers ranging from and where, and are the dimensions of the grid, represents the coordinate position of the Euler grid; The expression of the Gaussian equation is specifically as follows: ; In the above formula, represents the dynamic fluid surface height under the Gaussian equation, represents the wave amplitude, represents the damping factor, represents the phase factor, represents the time, represents the coordinate position of the Euler grid.

[0007] Further, the step of preprocessing the simulated fluid data set to obtain a reference image, a distorted image, real depth field information, and real normal field information specifically includes: Determine the reference image and real depth field information based on the simulated fluid data set, and obtain the air refractive index and the liquid refractive index; Perform a partial derivative calculation on the real depth field information to obtain the real normal field information; Combine the real normal field information, the air refractive index, and the liquid refractive index, and perform single refraction based on the reference image to obtain the distorted image.

[0008] Further, the dynamic fluid prediction network model specifically includes a graph partitioning layer, a graph embedding layer, a recurrent downsampling layer, a recurrent upsampling layer, and a graph splicing layer. The graph partitioning layer, the graph embedding layer, the recurrent downsampling layer, the recurrent upsampling layer, and the graph splicing layer are connected in sequence, where: The recurrent downsampling layer includes a first SwinLSTM core prediction unit and a tile merging unit. The first SwinLSTM core prediction unit includes a first SwinTransformer block and a first LSTM unit. The first SwinTransformer block includes a first sub-module and a second sub-module. The first sub-module and the second sub-module have the same structure. Among them, the first sub-module includes a first layer normalization layer, a first sliding window multi-head attention mechanism layer, a first multi-layer perceptron layer, and a second layer normalization layer; The recurrent upsampling layer includes a second SwinLSTM core prediction unit and a tile expansion unit. The second SwinLSTM core prediction unit includes a second SwinTransformer block and a second LSTM unit. The second SwinTransformer block includes a third sub-module and a fourth sub-module. The third sub-module and the fourth sub-module have the same structure. Among them, the third sub-module includes a third layer normalization layer, a second sliding window multi-head attention mechanism layer, a second multi-layer perceptron layer, and a fourth layer normalization layer.

[0009] Further, the expression of the SwinLSTM core prediction unit is specifically as follows: ; In the above formula, represents the fluid frame information at time represents the hidden state information of the fluid data at time represents the hidden state information of the fluid data at time represents the cell state information of the fluid data at time represents Fluid data cell state information at a moment, Indicates and Intermediate variable information related to the fluid frame at a moment, Indicates a linear mapping layer, Indicates the concatenation of two tensors along the channel, Indicates a Swin-Transformer block, Indicates a hyperbolic tangent activation function, Indicates a sigmoid activation function, Indicates the Hadamard product, which represents the operation of multiplying the corresponding elements of two matrices.

[0010] Furthermore, the expression of the Swin Transformer block is specifically as follows: ; In the above formula, Indicates the feature information of the th layer, Indicates the feature information of the th layer, Indicates the feature information of the th layer, Indicates the intermediate feature information of the th layer, Indicates the intermediate feature information of the th layer, Indicates a layer normalization layer, Indicates a sliding window multi-head attention mechanism layer, Indicates a multi-layer perceptron layer.

[0011] Specifically, the expression of the loss function of the dynamic fluid prediction network model is specifically as follows: ; In the above formula, Indicates the loss function of the dynamic fluid prediction network model, Indicates the mean squared error loss function, Indicates the structural similarity loss function, Indicates the scale consistency loss function.

[0012] Specifically, the step of using the dynamic fluid prediction network model to predict the depth field and normal field of the reference image and the distorted image, and combining the real depth field information and real normal field information for quantitative evaluation to obtain the trained dynamic fluid prediction network model specifically includes: Input the reference image and the distorted image into the dynamic fluid prediction network model for training; Based on the graph partitioning layer of the dynamic fluid prediction network model, the reference image and the distorted image are partitioned to obtain non-overlapping linear small tiles; Based on the graph embedding layer of the dynamic fluid prediction network model, the non-overlapping linear small tiles are flattened and feature-mapped to obtain linear feature small tiles; The current hidden state and the current cell state are obtained according to the linear feature small tiles; Based on the recurrent downsampling layer of the dynamic fluid prediction network model, the linear feature small tiles, the current hidden state and the current cell state are recurrently downsampled to obtain the hidden state at the next moment and the cell state at the next moment; Based on the recurrent upsampling layer of the dynamic fluid prediction network model, the hidden state at the next moment and the cell state at the next moment are upsampled to obtain the upsampled hidden state and the upsampled cell state; Based on the graph concatenation layer of the dynamic fluid prediction network model, the upsampled hidden state and the upsampled cell state are concatenated to obtain the predicted depth field information and the predicted normal field information; Through the structural similarity index and the peak signal-to-noise ratio index, combined with the real depth field information and the real normal field information, the predicted depth field information and the predicted normal field information are quantitatively evaluated to obtain the quantitative evaluation result; If the quantitative evaluation result does not meet the preset requirements, the dynamic fluid prediction network model is retrained until the preset requirements are met, and the trained dynamic fluid prediction network model is output.

[0013] The second technical solution adopted by the present invention is: a dynamic fluid surface reconstruction system based on the SwinLSTM network, including: The first module is used to simulate the fluid behavior based on the preset wave equation and construct a simulation fluid data set; The second module is used to preprocess the simulation fluid data set to obtain the reference image, the distorted image, the real depth field information and the real normal field information; The third module is used to construct a dynamic fluid prediction network model based on the Swin-Transformer block and the LSTM unit; The fourth module is used to predict the depth field and the normal field of the reference image and the distorted image based on the dynamic fluid prediction network model, and perform quantitative evaluation in combination with the real depth field information and the real normal field information to obtain the trained dynamic fluid prediction network model; The fifth module is used to predict the depth field and the normal field of the dynamic fluid based on the trained dynamic fluid prediction network model to obtain the dynamic fluid surface reconstruction result.

[0014] The beneficial effects of the method and system of the present invention are as follows: First, the present invention simulates fluid behavior based on a preset wave equation, constructs a simulation fluid data set and performs preprocessing to obtain a reference image, a distorted image, real depth field information and real normal field information. A method of combining and superimposing three different forms of wave equations is used to simulate various types of fluids, constructing a dynamic fluid simulation data set containing rich samples. Then, based on the Swin-Transformer block and the LSTM unit, a dynamic fluid prediction network model is constructed. By replacing the convolutional layer in the traditional convolutional neural network with a self-attention mechanism and introducing the designed SwinLSTM core prediction module, this network solves the problem that the traditional convolutional neural network ignores the spatio-temporal consistency of the fluid sequence when learning local spatial information, enabling the network to effectively capture the spatio-temporal consistency information of the fluid sequence. Further, based on the dynamic fluid prediction network model, depth field and normal field predictions are made on the reference image and the distorted image, and quantitative evaluation is carried out by combining the real depth field information and the real normal field information. Finally, based on the trained dynamic fluid prediction network model, depth field and normal field predictions are made on the dynamic fluid, improving the reconstruction accuracy of the model on the simulation and real dynamic fluid data sets, being able to make full use of the prediction information of the dynamic fluid feed-forward frames, ensuring the spatio-temporal consistency of the dynamic fluid reconstruction result, and accurately restoring its surface three-dimensional morphology. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a flowchart of the steps of a method for dynamic fluid surface reconstruction based on a SwinLSTM network according to the present invention; Figure 2 is a structural block diagram of a system for dynamic fluid surface reconstruction based on a SwinLSTM network according to the present invention; Figure 3 is a schematic diagram of a simulation fluid data set provided by a specific embodiment of the present invention; Figure 4 is a schematic diagram of the structure of a dynamic fluid prediction network model provided by a specific embodiment of the present invention; Figure 5 is a schematic diagram of the structure of a SwinLSTM core prediction unit provided by a specific embodiment of the present invention; Figure 6 is a schematic diagram of the structure of a Swin-Transfomer block provided by a specific embodiment of the present invention; Figure 7 is a schematic diagram of the prediction effect on real fluid data provided by a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0017] Referring to Figure 1 , the present invention provides a dynamic fluid surface reconstruction method based on the SwinLSTM network, and the method includes the following steps: S100. Simulate the fluid behavior based on a preset wave equation to construct a simulation fluid data set; S110. Select the shallow water wave equation to simulate the fluid wave behavior under shallow water conditions to obtain the first fluid behavior simulation data; In this embodiment, the shallow water wave equation is a set of partial differential equations derived from the Navier-Stokes equation. This equation describes the incompressible nature of the fluid, including mass conservation and linear momentum conservation. The shallow water wave equation is as follows: ; In the above formula, represents the liquid density, represents the velocity of the fluid in the direction, represents the velocity of the fluid in the direction, represents the coordinate position of the Euler grid, represents the dynamic fluid surface height under the shallow water wave equation, represents the time point.

[0018] S120. Select the Gerstner wave equation to perform a fast Fourier transform and perform wave simulation in computer graphics to obtain the second fluid behavior simulation data; In this embodiment, the Gerstner wave equation is widely used in wave simulation in computer graphics. The embodiment of the present invention uses it to simulate the fluid motion with a relatively large volume and adopts the fast Fourier transform (FFT) form to calculate the Gerstner equation. The equation representation based on FFT is: ;

[0019] In the above formula, represents the dynamic fluid surface height under the Gerstner wave equation, represents the Fourier amplitude, represents the imaginary unit, and are in the range of and an integer between, where, and are the dimensions of the grid, represents the coordinate position of the Eulerian grid.

[0020] S130. Select the Gaussian equation to simulate the water ripple with damping effect, and obtain the third fluid behavior simulation data; In this embodiment, the Gaussian equation is used to simulate the water ripple with damping effect, represents the dynamic fluid surface height, then the expression of the Gaussian equation is: ; In the above formula, represents the dynamic fluid surface height under the Gaussian equation, represents the wave amplitude, represents the damping factor, represents the phase factor, represents the time, represents the coordinate position of the Eulerian grid.

[0021] S140. Perform weighted linear combination on the first fluid behavior simulation data, the second fluid behavior simulation data and the third fluid behavior simulation data to construct a simulation fluid data set.

[0022] In this embodiment, the weighted linear combination of the first fluid behavior simulation data, the second fluid behavior simulation data and the third fluid behavior simulation data is as follows: ; In the above formula, represents the dynamic fluid surface height under the combined equation, represents the dynamic fluid surface height under the shallow water wave equation, represents the dynamic fluid surface height under the Gerstner wave equation, represents the dynamic fluid surface height under the Gaussian equation, and represent the weight coefficients.

[0023] Furthermore, it should be noted that by the superposition of different weight coefficient ratios, various physically simulated fluid effects can be obtained to construct multiple sets of fluid sequence data sets, where the Eulerian grid size of all simulated fluids is set to , and the simulation fluid data set constructed by the combined equation is as shown in Figure 3 (a) in, including a 3-channel undistorted reference variable image, a 3-channel distorted image, a 1-channel depth field image and a 3-channel normal field image.

[0024] S200. Preprocess the simulated fluid dataset to obtain a reference image, a distorted image, true depth field information, and true normal field information; Specifically, determine the reference image and true depth field information based on the simulated fluid dataset, and obtain the air refractive index and the liquid refractive index; perform a partial derivative calculation on the true depth field information to obtain the true normal field information; combine the true normal field information, the air refractive index, and the liquid refractive index, and perform single refraction based on the reference image to obtain the distorted image.

[0025] In this embodiment, the schematic diagram of the dataset construction principle is as shown in Figure 3 Figure (b). First, the height field information of different dynamic fluid surfaces can be simulated through the above combined equation. The normal field information of the fluid surface can be further obtained by taking the partial derivative of the height field information. By combining the normal field with the refractive index of the first medium and the refractive index of the second medium , a sequence of distorted images after single refraction can be calculated from the undistorted reference image below the liquid surface.

[0026] Furthermore, it should be noted that 12,000 pairs of data (network input data and network true label data) are constructed by the above combined equation for training and testing. Among them, the total size of the network input data is . Among them, 12,000 represents the total number of samples, 10 represents the number of adjacent frames in the fluid sequence, 256 represents the length and width dimensions of the input tensor, and 6 represents the concatenation of the 3-channel distorted image and the 3-channel undistorted reference image in the channel dimension. Among them, the total size of the network true label data is . Among them, 12,000 represents the total number of samples, 10 represents the number of adjacent frames in the fluid sequence, 256 represents the length and width dimensions of the output tensor, and 4 represents the concatenation of the 1-channel predicted depth field and the 3-channel predicted normal field in the channel dimension.

[0027] Then, the dataset is divided into a training set and a test set, where the training set accounts for 80% and the test set accounts for 20%. Among them, the total network input size of the training set is , the total network label size of the training set is , the total network input size of the test set is , and the total network label size of the test set is .

[0028] S300. Build a dynamic fluid prediction network model based on Swin-Transformer blocks and LSTM units; Specifically, use the Pytorch deep learning framework to build an end-to-end network model based on Swin-Transformer and LSTM for dynamic fluid depth field and normal field prediction, as shown inFigure 4 As shown, the dynamic fluid prediction network model specifically includes a graph partitioning layer, a graph embedding layer, a cyclic downsampling layer, a cyclic upsampling layer, and a graph splicing layer. The graph partitioning layer, the graph embedding layer, the cyclic downsampling layer, the cyclic upsampling layer, and the graph splicing layer are connected in sequence, where: Both the cyclic downsampling layer and the cyclic upsampling layer have N layers.

[0029] The cyclic downsampling layer includes a first SwinLSTM core prediction unit and a tile merging unit. The first SwinLSTM core prediction unit includes a first SwinTransformer block and a first LSTM unit. The first SwinTransformer block includes a first sub-module and a second sub-module. The first sub-module and the second sub-module have the same structure. Among them, the first sub-module includes a first layer normalization layer, a first sliding window multi-head attention mechanism layer, a first multi-layer perceptron layer, and a second layer normalization layer.

[0030] The cyclic upsampling layer includes a second SwinLSTM core prediction unit and a tile expansion unit. The second SwinLSTM core prediction unit includes a second SwinTransformer block and a second LSTM unit. The second SwinTransformer block includes a third sub-module and a fourth sub-module. The third sub-module and the fourth sub-module have the same structure. Among them, the third sub-module includes a third layer normalization layer, a second sliding window multi-head attention mechanism layer, a second multi-layer perceptron layer, and a fourth layer normalization layer.

[0031] S400. Perform depth field and normal field prediction on the reference image and the distorted image based on the dynamic fluid prediction network model, and conduct quantitative evaluation by combining the real depth field information and the real normal field information to obtain the trained dynamic fluid prediction network model; Specifically, first, the distorted images and reference images of 10 groups of fluid sequence frames input are partitioned into non-overlapping linear small tiles by the graph partitioning layer, and after flattening each small tile, feature embedding is performed to map it to any dimension C (C represents the number of channels). Then it is sent into the SwinLSTM core prediction unit in the cyclic downsampling module. This core prediction unit accepts 3 input elements: namely, the feature tile mapped to any dimension C, the hidden state and the cell state . And it outputs two elements: the hidden state and the cell state . Among them is copied into two copies, one for the subsequent cyclic upsampling module, and the other is one of the inputs for the next upsampling cyclic module. The cell state is then passed into the tile expansion unit, and the resulting output is combined with As the input of the next-round cyclic upsampling module, the output obtained after performing N rounds of downsampling cycles and the copied hidden state vector group are used as the input of the cyclic upsampling module. Similarly, the output obtained after performing N rounds of upsampling cycles is used as the input of the final graph stitching layer, and the depth field and normal field information of the current 10-frame continuous fluid sequence frames can be obtained.

[0032] S410. Input the reference image and the distorted image into the dynamic fluid prediction network model for training; S420. Based on the graph partitioning layer of the dynamic fluid prediction network model, perform partitioning processing on the reference image and the distorted image to obtain non-overlapping linear small tiles; S430. Based on the graph embedding layer of the dynamic fluid prediction network model, flatten and perform feature mapping processing on the non-overlapping linear small tiles to obtain linear feature small tiles; S440. Obtain the current hidden state and the current cell state according to the linear feature small tiles; S450. Based on the cyclic downsampling layer of the dynamic fluid prediction network model, perform cyclic downsampling on the linear feature small tiles, the current hidden state, and the current cell state to obtain the hidden state at the next moment and the cell state at the next moment; S460. Based on the cyclic upsampling layer of the dynamic fluid prediction network model, perform upsampling on the hidden state at the next moment and the cell state at the next moment to obtain the upsampled hidden state and the upsampled cell state; First of all, it should be noted that the structures and functions of the SwinLSTM core prediction units and the Swin-Transformer blocks in the cyclic downsampling layer and the cyclic upsampling layer are the same, and the embodiments of the present invention will be described uniformly.

[0033] In this embodiment, the SwinLSTM core prediction unit is as Figure 5 shown. In the SwinLSTM unit, the unit state and the hidden state are updated in the horizontal direction to capture the time dependence of the dynamic fluid data in the long stage and the short stage. At the same time, the Swin-Transformer block learns the global spatial dependence of the dynamic fluid data in the vertical direction. The formula description of the SwinLSTM core prediction unit is as follows: ; In the above formula, represents the fluid frame information at the moment, represents the hidden state information of the fluid data at the moment, represents Fluid data hiding state information at a moment, denote Fluid data cell state information at a moment, denote Fluid data cell state information at a moment, denote and Intermediate variable information related to the fluid frame at a moment, denote the linear mapping layer, denote the concatenation of two tensors along the channel, denote the Swin-Transformer block, denote the hyperbolic tangent activation function, denote the sigmoid activation function, denote the Hadamard product, which represents the operation of multiplying the corresponding elements of two matrices.

[0034] Furthermore, it should be noted that the Swin-Transformer block, as Figure 6 shown, consists of two sub-modules. The first sub-module consists of two layer normalization layers, a sliding window multi-head attention mechanism layer, and a multi-layer perceptron layer. A layer normalization operation and a residual connection operation are respectively applied before and after each sliding window multi-head attention mechanism layer and multi-layer perceptron layer; the second sub-module is exactly the same as the first sub-module. The formula description of the Swin-Transformer block is as follows: ; In the above formula, denote the feature information of the th layer, denote the feature information of the th layer, denote the feature information of the th layer, denote the intermediate feature information of the th layer, denote the layer normalization layer, denote the sliding window multi-head attention mechanism layer, denote the multi-layer perceptron layer. denote the multi-layer perceptron layer.

[0035] S470. Based on the graph concatenation layer of the dynamic fluid prediction network model, concatenate the upsampled hidden state and the upsampled cell state to obtain the predicted depth field information and the predicted normal field information; ​S480. Quantitatively evaluate the predicted depth field information and the predicted normal field information by combining the structural similarity index and the peak signal-to-noise ratio index with the true depth field information and the true normal field information to obtain a quantitative evaluation result; Specifically, use the data of the constructed simulation test set to quantitatively evaluate the model performance. Select SSIM (structural similarity index) and PSNR (peak signal-to-noise ratio) as the quantitative evaluation indicators to judge the quality of the depth field and normal field recovery of the model obtained by training for the simulation fluid data. The specific formulas are as follows: ; In the above formula, and respectively represent the average values of the images and , and respectively represent the variances of the images and , represents the covariance of the images and , represents the minimum constant introduced for stability, represents the maximum possible pixel value in the image. For the normalized data, the maximum value is 1, represents the mean squared error between the predicted image and the label image.

[0036] S490. If the quantitative evaluation result does not meet the preset requirements, retrain the dynamic fluid prediction network model until the preset requirements are met, and output the trained dynamic fluid prediction network model.

[0037] Specifically, instantiate the constructed Swin-Transformer and LSTM, and use the training set constructed in step S100 to train the model. After the training is completed, save the network model parameters.

[0038] It should also be noted that according to the particularity of the dynamic fluid data set, a corresponding loss function is designed, and its expression is specifically as follows: ; In the above formula, represents the loss function of the dynamic fluid prediction network model, represents the mean squared error loss function, which is used to ensure the basic convergence of the network, represents the structural similarity loss function, which is used to ensure the structural similarity of the depth field and normal field information, thereby accelerating the convergence of the network, represents the scale consistency loss function, which is used to ensure the scale consistency of the depth field and normal field of adjacent frames in the fluid sequence prediction.

[0039] S500. Predict the depth field and normal field of the dynamic fluid based on the trained dynamic fluid prediction network model to obtain the dynamic fluid surface reconstruction result.

[0040] Furthermore, as shown in Table 1, it is a performance comparison of the method proposed in the embodiments of the present invention with common CNN and RNN structure networks under the above quantitative indicators; Table 1 Comparison data table

[0041] According to Table 1, it can be seen that the Swin-Transformer and LSTM network structures proposed by the present invention perform significantly better than the network structures designed based on CNN and RNN in the above various quantitative indicators. It shows that the performance of this method is good in the dynamic fluid three-dimensional reconstruction task.

[0042] Furthermore, capture the dynamic fluid data set in the real environment and use the real data set to qualitatively verify the model performance. The qualitative evaluation results are as Figure 7 shown, Figure 7 The first column is the sequence of distorted images of the dynamic fluid captured in the real environment, Figure 7 The second column is the undistorted reference image placed at the bottom of the dynamic fluid, Figure 7 The third column is the depth map prediction result of the model proposed by the present invention on the real fluid data, Figure 7 The fourth column is the normal map prediction result of the model proposed by the present invention on the real fluid data, Figure 7 The fifth column is the three-dimensional visualization display of the depth map prediction result.

[0043] In summary, in the embodiments of the present invention, a simulation fluid training dataset is first constructed using a combined wave equation and divided into a training set and a test set. To address the problem that traditional convolutional neural networks ignore the spatio-temporal consistency of fluid sequences when learning local spatial information, the present invention designs an end-to-end neural network structure that integrates Swin-Transformer and LSTM. Specifically, the input of the network consists of multiple frames of reference images and distorted images of the fluid sequence, and the output is the depth field and normal field information of the dynamic fluid surface. The present invention replaces the convolutional layer in the traditional convolutional neural network with a self-attention mechanism and introduces the designed SwinLSTM core prediction module, enabling the network to effectively capture the consistency information of the fluid sequence in time and space. Experimental results show that the method proposed in the present invention significantly improves the reconstruction accuracy of the model on simulation and real dynamic fluid datasets. First, the network is trained on the constructed simulation fluid training set, then quantitatively evaluated on the test dataset, and finally qualitatively evaluated on the constructed real dynamic fluid dataset. The trained model is used to predict the real fluid distorted image sequence and the reference image sequence to achieve the three-dimensional reconstruction of the real dynamic fluid surface, which has broad application potential in related fields such as three-dimensional reconstruction of dynamic fluid surfaces.

[0044] Referring to Figure 2 , a dynamic fluid surface reconstruction system based on a SwinLSTM network, comprising: A first module 201 for simulating fluid behavior based on a preset wave equation to construct a simulation fluid dataset; A second module 202 for preprocessing the simulation fluid dataset to obtain reference images, distorted images, real depth field information, and real normal field information; A third module 203 for constructing a dynamic fluid prediction network model based on Swin-Transformer blocks and LSTM cells; A fourth module 204 for predicting the depth field and normal field of the reference image and the distorted image based on the dynamic fluid prediction network model, and performing quantitative evaluation by combining the real depth field information and the real normal field information to obtain a trained dynamic fluid prediction network model; A fifth module 205 for predicting the depth field and normal field of the dynamic fluid based on the trained dynamic fluid prediction network model to obtain the dynamic fluid surface reconstruction result.

[0045] The content in the above method embodiments is applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0046] The above is a specific description of the preferred embodiment of the present invention. However, the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A dynamic fluid surface reconstruction method based on SwinLSTM network, characterized in that: The following steps are involved: Simulate fluid behavior based on preset wave equations and build a simulation fluid data set; Preprocess the simulated fluid data set to obtain reference images, distorted images, true depth field information, and true normal field information; Based on Swin-Transformer blocks and LSTM units, a dynamic fluid prediction network model is constructed; Based on the dynamic fluid prediction network model, the depth field and normal field of the reference image and the distorted image are predicted, and quantitative evaluation is performed by combining the real depth field information and the real normal field information to obtain the trained dynamic fluid prediction network model; Based on the trained dynamic fluid prediction network model, the depth field and normal field of the dynamic fluid are predicted to obtain the dynamic fluid surface reconstruction result.

2. A dynamic fluid surface reconstruction method based on SwinLSTM network according to claim 1, characterized in that: The step of constructing a simulation fluid data set based on a preset wave equation specifically includes: Selecting a shallow water wave equation to simulate the fluid wave behavior under shallow water conditions, and obtaining first fluid behavior simulation data; The Grestner wave equation is selected for fast Fourier transform, and the ocean wave simulation in computer graphics is performed to obtain the simulation data of the second fluid behavior; Selecting Gaussian equation to simulate water ripples with damping effect, and obtaining third fluid behavior simulation data; A weighted linear combination is performed on the first fluid behavior simulation data, the second fluid behavior simulation data, and the third fluid behavior simulation data to construct a simulated fluid data set.

3. A dynamic fluid surface reconstruction method based on SwinLSTM network according to claim 2, characterized in that: The expression of the shallow water wave equation is specifically as follows: ; In the above formula, represents the density of the liquid, Indicates that the fluid The speed of the direction, Indicates that the fluid The speed of the direction, represents the coordinate position of the Euler grid, represents the dynamic fluid surface height under the shallow water wave equation, Indicates a point in time; The expression of the Grestner wave equation is specifically as follows: ; In the above formula, represents the dynamic fluid surface height under the Gerstner wave equation, represents the Fourier amplitude, represents the imaginary unit, and The range is and Integers between , where and is the dimension of the grid, Represents the coordinate position of the Euler grid; The Gaussian equation is specifically expressed as follows: ; In the above formula, represents the dynamic fluid surface height under Gaussian equation, Indicates the amplitude, represents the damping factor, represents the phase factor, Indicates time, Represents the coordinate position of the Euler grid.

4. A dynamic fluid surface reconstruction method based on SwinLSTM network according to claim 3, characterized in that: The step of preprocessing the simulated fluid data set to obtain a reference image, a distorted image, real depth field information and real normal field information specifically includes: Determine the reference image and the real depth field information based on the simulated fluid data set to obtain the air refractive index and the liquid refractive index; Perform partial derivative calculation on the real depth field information to obtain the real normal field information; Combining the true normal field information, the refractive index of air and the refractive index of liquid, a single refraction is performed according to the reference image to obtain a distorted image.

5. A dynamic fluid surface reconstruction method based on SwinLSTM network according to claim 4, characterized in that: The dynamic fluid prediction network model specifically includes a graph partitioning layer, a graph embedding layer, a cyclic downsampling layer, a cyclic upsampling layer and a graph splicing layer, wherein the graph partitioning layer, the graph embedding layer, the cyclic downsampling layer, the cyclic upsampling layer and the graph splicing layer are connected in sequence, wherein: The cyclic downsampling layer includes a first SwinLSTM core prediction unit and a tile merging unit, the first SwinLSTM core prediction unit includes a first SwinTransformer block and a first LSTM unit, the first SwinTransformer block includes a first submodule and a second submodule, the first submodule and the second submodule have the same structure, wherein the first submodule includes a first normalization layer, a first sliding window multi-head attention mechanism layer, a first multi-layer perceptron layer and a second normalization layer; The cyclic upsampling layer includes a second SwinLSTM core prediction unit and a tile expansion unit, the second SwinLSTM core prediction unit includes a second SwinTransformer block and a second LSTM unit, the second SwinTransformer block includes a third submodule and a fourth submodule, the third submodule has the same structure as the fourth submodule, wherein the third submodule includes a third normalization layer, a second sliding window multi-head attention mechanism layer, a second multi-layer perceptron layer and a fourth normalization layer.

6. A dynamic fluid surface reconstruction method based on SwinLSTM network according to claim 5, characterized in that: The expression of the SwinLSTM core prediction unit is specifically as follows: ; In the above formula, express Moment fluid frame information, express The fluid data hides the status information at all times. express The fluid data hides the status information at all times. express Fluid data cell status information at all times, express Fluid data cell status information at all times, Representation and Intermediate variable information related to the fluid frame at the moment, represents a linear mapping layer, Indicates the concatenation of two tensors along the channel. represents a Swin-Transformer block, represents the hyperbolic tangent activation function, represents the sigmoid activation function, The Hadamard product represents the operation of multiplying corresponding elements of two matrices.

7. A dynamic fluid surface reconstruction method based on SwinLSTM network according to claim 6, characterized in that: The expression of the SwinTransformer block is as follows: ; In the above formula, Indicates The characteristic information of the layer, Indicates The characteristic information of the layer, Indicates The characteristic information of the layer, Indicates The intermediate feature information of the layer, Indicates The intermediate feature information of the layer, Representation layer normalization layer, represents the sliding window multi-head attention mechanism layer, Represents a multi-layer perceptron layer.

8. A dynamic fluid surface reconstruction method based on SwinLSTM network according to claim 7, characterized in that: The loss function of the dynamic fluid prediction network model is specifically expressed as follows: ; In the above formula, represents the loss function of the dynamic fluid prediction network model, represents the mean square error loss function, represents the structural similarity loss function, represents the scale consistency loss function.

9. A dynamic fluid surface reconstruction method based on SwinLSTM network according to claim 8, characterized in that: The step of predicting the depth field and the normal field of the reference image and the distorted image based on the dynamic fluid prediction network model, and performing quantitative evaluation in combination with the real depth field information and the real normal field information to obtain the trained dynamic fluid prediction network model specifically includes: Input the reference image and the distorted image into the dynamic fluid prediction network model for training; Based on the graph partitioning layer of the dynamic fluid prediction network model, the reference image and the distorted image are partitioned to obtain non-overlapping linear small blocks; Based on the graph embedding layer of the dynamic fluid prediction network model, non-overlapping linear tiles are flattened and feature mapped to obtain linear feature tiles; Get the current hidden state and current cell state based on the linear feature tile; Based on the cyclic downsampling layer of the dynamic fluid prediction network model, the linear feature tiles, the current hidden state and the current cell state are cyclically downsampled to obtain the hidden state at the next moment and the cell state at the next moment; Based on the cyclic upsampling layer of the dynamic fluid prediction network model, the hidden state at the next moment and the cell state at the next moment are upsampled to obtain the upsampled hidden state and the upsampled cell state; Based on the graph concatenation layer of the dynamic fluid prediction network model, the upsampled hidden state and the upsampled cell state are concatenated to obtain the predicted depth field information and the predicted normal field information; Through the structural similarity index and the peak signal-to-noise ratio index, the predicted depth field information and the predicted normal field information are quantitatively evaluated in combination with the real depth field information and the real normal field information to obtain the quantitative evaluation results; If the quantitative evaluation result does not meet the preset requirements, the dynamic fluid prediction network model is retrained until it meets the preset requirements, and the trained dynamic fluid prediction network model is output.

10. A dynamic fluid surface reconstruction system based on SwinLSTM network, characterized in that: Includes the following modules: The first module is used to simulate fluid behavior based on a preset wave equation and construct a simulation fluid data set; The second module is used to preprocess the simulated fluid data set to obtain a reference image, a distorted image, true depth field information and true normal field information; The third module is used to build a dynamic fluid prediction network model based on Swin-Transformer blocks and LSTM units; The fourth module is used to predict the depth field and normal field of the reference image and the distorted image based on the dynamic fluid prediction network model, and to perform quantitative evaluation based on the real depth field information and the real normal field information to obtain the trained dynamic fluid prediction network model; The fifth module is used to predict the depth field and normal field of the dynamic fluid based on the trained dynamic fluid prediction network model to obtain the dynamic fluid surface reconstruction result.

Citation Information

Patent Citations

  • Liquid surface three-dimensional reconstruction method and system based on data driving

    CN117993302A

  • Hydromechanics partial differential equation solving device based on autoregressive neural network

    CN118586311A

  • Image water level identification method and device based on scene simulation

    CN118762314A

Cited By

  • Fluid motion data prediction method and system based on sensor

    CN121431005A