A multi-modal fusion-based water flow velocity intelligent evaluation method

By combining the ST-GNN network with multimodal fusion and the improved PINN model, along with Fourier neural operators and Navier-Stokes equations, the problems of single data utilization and lack of physical laws in existing water flow velocity assessments are solved, achieving high-precision and reliable velocity assessment and uncertainty quantification.

CN122113628APending Publication Date: 2026-05-29QINGDAO YIHE XINSHUI TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO YIHE XINSHUI TECHNOLOGY CO LTD
Filing Date
2026-02-25
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing methods for assessing water flow velocity rely on single-mode data, making it difficult to fully utilize multi-source heterogeneous information and lacking physical constraints. This results in insufficient prediction accuracy and a lack of uncertainty quantification, affecting the reliability and usability of the assessment results.

Method used

A multimodal fusion approach is adopted, using an ST-GNN network to construct a dynamic graph structure for spatiotemporal feature extraction. An improved PINN model is used to introduce Fourier neural operators and Navier-Stokes equations as strong constraints, and a comprehensive loss function is combined to perform flow field prediction and uncertainty quantification.

Benefits of technology

It achieves high-precision and high-reliability water flow velocity assessment, and can provide physical consistency and uncertainty quantification in complex aquatic environments, significantly improving the accuracy and robustness of velocity assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113628A_ABST
    Figure CN122113628A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-modal fusion's water flow velocity intelligent evaluation method, comprising: S1, acquisition multi-modal data, output original multi-modal data;S2, original multi-modal data is preprocessed and feature extraction, generate standardization space-time data set;S3, utilize pre-trained ST-GNN network, output uniform feature vector;S4, uniform feature vector is input into improved PINN model, introduce Fourier neural operator, output three-dimensional flow field data body and global uncertainty quantification diagram;S5, according to uncertainty quantification diagram, three-dimensional flow field data body is confidence weighted and physically corrected, and generates high-precision evaluation model;S6, evaluation result is rendered as visual image, and superimposed display uncertainty quantification diagram.The application realizes the high-precision, physical consistency prediction of complex water area flow field, and improves the reliability and practicality of evaluation result by uncertainty quantification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent water conservancy and deep learning, and in particular to an intelligent method for evaluating water flow velocity based on multimodal fusion. Background Technology

[0002] PINN and ST-GNN networks, due to their ability to integrate prior physical knowledge and model complex spatiotemporal dependencies respectively, have been widely used in fluid simulation, weather forecasting, and other fields in recent years, becoming an important development direction for intelligent flow field assessment. However, in practical applications, water flow velocity assessment scenarios face many challenges, such as diverse data sources, complex physical mechanisms, and high prediction accuracy requirements, and the deployment effectiveness of existing methods is still constrained by many factors.

[0003] Most current evaluation methods rely on single-modal data or purely data-driven models, making it difficult to fully utilize multi-source heterogeneous information such as surface vision, underwater acoustics, and moving references, or completely ignoring the basic physical laws of fluid dynamics. This results in a lack of comprehensiveness and physical consistency in the modeling of dynamic flow fields. Although some systems use ST-GNN, their feature aggregation lacks a decoding mechanism that can explicitly map the learned features to specific physical fields and be constrained by physical equations, thus limiting the accuracy of the model in physical reconstruction tasks.

[0004] Furthermore, existing models typically focus only on data fitting errors or only partially incorporate physical losses during training, failing to effectively simulate global high-frequency interactions in fluid dynamics. This results in insufficient simulation accuracy for complex phenomena such as eddies and turbulence. Simultaneously, the model prediction process generally lacks a quantitative assessment mechanism for its own uncertainties, failing to provide users with confidence information on the prediction results, severely impacting the usability and reliability of the evaluation results in critical decision-making.

[0005] Therefore, how to provide a smart method for evaluating water flow velocity based on multimodal fusion is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose an intelligent water flow velocity assessment method based on multimodal fusion. This invention fully integrates key steps such as synchronous multimodal data acquisition, spatiotemporal feature extraction, ST-GNN network encoding, improved PINN model physical constraint prediction, and confidence-weighted correction. It constructs an intelligent water flow assessment process with data spatiotemporal alignment, dynamic graph structure evolution, physical law embedding, and uncertainty quantification output, achieving high-precision and high-reliability prediction of three-dimensional flow fields in complex aquatic environments. This invention effectively captures the spatiotemporal evolution characteristics of the flow field through the ST-GNN network and uses the improved PINN model with the Navier-Stokes equations as strong constraints, ensuring the physical consistency of the prediction results. This invention possesses advantages such as multi-dimensional feature extraction fusion, physical consistency of flow field prediction, uncertainty quantification assessment, and high result reliability, significantly improving the accuracy and robustness of global water flow velocity assessment, thereby effectively solving problems such as single data utilization, missing physical laws, and insufficient prediction reliability in existing methods.

[0007] A method for intelligent assessment of water flow velocity based on multimodal fusion according to an embodiment of the present invention includes the following steps: S1. Synchronously acquire multimodal data and output raw multimodal data; S2. Preprocess the original multimodal data, and perform image enhancement, acoustic computation and multimodal feature extraction to generate a standardized spatiotemporal dataset; S3. Input the standardized spatiotemporal dataset into the pre-trained ST-GNN network, discretize the monitored water area into a dynamic graph structure, simulate the physical evolution of the flow field by performing spatiotemporal message passing on the dynamic graph structure, and output a unified feature vector. S4. Input the unified feature vector into the improved PINN model, introduce the Fourier neural operator, construct a comprehensive loss function, use the benchmark flow velocity data for dynamic calibration, and output a three-dimensional flow field data volume and a global uncertainty quantification map. S5. Based on the global uncertainty quantification diagram, the three-dimensional flow field data volume is weighted with confidence and physically corrected to generate a high-precision global water flow velocity assessment model. S6. On the user end, the evaluation results of the high-precision global water flow velocity assessment model are rendered into a visual image, providing interactive functions such as real-time display and data query, and an uncertainty quantification chart is overlaid.

[0008] Optionally, S1 specifically includes: S11. By deploying a synchronization time block on the unmanned surface vessel, a synchronization control signal with a unified timestamp is generated, which simultaneously triggers the visual video acquisition device, the underwater acoustic acquisition device, and the reference current velocity acquisition device to collect the visual video stream, the underwater acoustic data stream, and the reference current velocity data of the monitored water surface, respectively. S12. The acquired raw visual video stream, underwater acoustic data stream and reference current velocity data are encapsulated in chronological order and output as synchronously acquired raw multimodal data.

[0009] Optionally, S2 specifically includes: S21. Read the timestamps of each frame of visual video stream and each reference current velocity data in the original multimodal data. Using the timestamps of the reference current velocity data as a reference, perform time interpolation on the visual video stream. Read the timestamps of the underwater acoustic data stream and each reference current velocity data in the original multimodal data. Using the timestamps of the reference current velocity data as a reference, perform time interpolation on the underwater acoustic data stream. S22. Perform histogram equalization on the time-aligned visual video stream to redistribute pixel grayscale values. Perform Doppler frequency shift calculation on the time-aligned underwater acoustic data stream to convert the frequency offset of the echo signal into water flow velocity value. Linearly scale the visual video stream with image enhancement completed, the underwater acoustic data stream with acoustic calculation completed, and the reference flow velocity data according to a preset uniform value range. S23. Input each frame of the visual video stream after standardization and preprocessing into a pre-trained convolutional neural network, and output a visual feature vector of a set dimension. Input the underwater acoustic data stream after standardization and preprocessing into a pre-trained long short-term memory network, and output an acoustic feature vector of a set dimension. Directly construct a reference data vector of a set dimension from the reference current velocity data after standardization and preprocessing. S24. The visual feature vector, acoustic feature vector and the baseline data vector under the same timestamp are concatenated by dimension to form a multimodal feature vector. The multimodal feature vectors under all timestamps are integrated in chronological order to generate a standardized spatiotemporal dataset.

[0010] Optionally, S3 specifically includes: S31. Input the standardized spatiotemporal dataset into the pre-trained ST-GNN network, read the geographic boundary coordinates of the monitored water area, generate a two-dimensional Cartesian coordinate grid, define the geometric center point of each grid cell as a graph node, assign a unique index to each node, and construct the initial graph structure. S32. Assign an independent multilayer perceptron to each graph node. The input layer receives the visual feature vector, acoustic feature vector and reference data vector corresponding to the graph node index in the standardized spatiotemporal dataset. The two hidden layers perform nonlinear transformations with preset weights and biases and concatenate them into an initial feature vector, which is then assigned to the corresponding graph node. S33. Assign a Gaussian kernel function to each edge in the graph structure, read the two-dimensional coordinates of any two nodes and calculate the Euclidean distance, substitute the distance value into the Gaussian kernel function for calculation, use the calculation result as the weight of the edge connecting the two nodes, traverse all node pairs, and construct the spatial adjacency matrix. S34. Combine the initial feature vectors of all nodes into a feature matrix, and input it together with the spatial adjacency matrix into the graph convolutional network layer. By performing matrix multiplication, multiply the spatial adjacency matrix with the feature matrix, perform weighted aggregation on the feature vectors of each node, apply a linear transformation layer and ReLU activation function to generate an updated spatial fusion feature matrix. S35. Input the updated spatial fusion feature matrix into the gated recurrent unit network layer in chronological order. Receive the spatial fusion feature matrix of the current time step and the hidden state of the previous time step. Calculate the hidden state of the current time step through the reset gate and update gate. S36. The hidden state is used as a spatiotemporal fusion feature matrix containing time dependencies, and is then input into the graph convolutional network layer and the gated recurrent unit network layer for a second time. Multiple alternating iterations are performed. In each iteration, the graph convolutional network layer performs spatial feature aggregation, and the gated recurrent unit network layer performs temporal feature update. This process is repeated until the preset number of iterations is reached, and the final spatiotemporal fusion feature matrix is ​​output. S37. Input the final spatiotemporal fusion feature matrix into the global average pooling layer, calculate the arithmetic mean for each dimension along the node dimension, and compress the features of the graph nodes into a one-dimensional vector by arranging them in order and output them as a unified feature vector.

[0011] Optionally, S4 specifically includes: S41. Input the unified feature vector into the improved PINN model, introduce the Fourier neural operator, and construct the comprehensive loss function; S42. Perform backpropagation, starting with the comprehensive loss function, and use the chain rule to calculate the gradient of each learnable parameter in reverse. Use the Adam algorithm to perform an update operation on each parameter based on the calculated gradient, and repeat the iterative steps of forward propagation, backpropagation and parameter update. S43. When the number of iterations reaches the upper limit, stop training and save the improved PINN model after training. Predict the unified feature vector of the new input and output the three-dimensional flow field data volume and the global uncertainty quantization map.

[0012] Optionally, S41 specifically includes: The unified feature vector is input into the improved PINN model. A fully connected network with three hidden layers is used, and a decoder with a ReLU activation function is connected after each hidden layer to linearly map the unified feature vector into a two-dimensional tensor. The horizontal dimension of the two-dimensional tensor is the total number of graph nodes, and the column vector is the three-dimensional velocity component of each node. The two-dimensional tensor is the three-dimensional velocity field of all graph nodes. The two-dimensional tensor and the initial pressure field tensor with the same dimension set to zero are concatenated by the last dimension to form the initial three-dimensional flow field state tensor, i.e. the three-dimensional flow field data volume. Then, the Fourier neural operator is introduced to perform a two-dimensional fast Fourier transform on the initial three-dimensional flow field state tensor to transform the flow field state from the spatial domain to the frequency domain, resulting in a frequency domain tensor in complex form. The frequency domain tensor is input into a learnable Fourier layer containing a weight matrix. By performing complex multiplication of the frequency domain tensor and the weight matrix, the different frequency components in the frequency domain are weighted to simulate the global interaction in fluid dynamics. A two-dimensional fast Fourier inverse transform is performed on the weighted frequency domain tensor to convert it back from the frequency domain to the spatial domain, resulting in an updated spatial domain tensor. The updated spatial domain tensor is input into a convolutional layer with a 1×1 kernel. The six features of each node are locally linearly mixed, and the intermediate state tensor is output. It is then added element-wise to the initial three-dimensional flow field state tensor introduced by the Fourier neural operator to form a residual connection. The predicted three-dimensional flow field state tensor for the next time step is output. The predicted three-dimensional flow field state tensor for the next time step is then introduced into the Fourier neural operator. The prediction operation is repeated to form a flow field prediction sequence. In the flow field prediction sequence, the predicted three-dimensional flow field state tensor at a specific time corresponding to the timestamp of the baseline flow velocity data is selected, and the predicted three-dimensional velocity value at the graph node index corresponding to the spatial position of the baseline flow velocity collector is extracted. The predicted three-dimensional velocity value is compared with the measured three-dimensional velocity value of the baseline flow velocity data point by point, the square of the difference between the two is calculated, and the average of the squared differences of all reference points at all corresponding times is taken as the data loss. For each predicted three-dimensional flow field state tensor in the flow field prediction sequence, the velocity and pressure components, along with the spatial coordinates of the graph nodes, are substituted into the discrete form of the Navier-Stokes equations. The difference between the left-hand and right-hand terms of the equations is calculated as the residual, and the square of the L2 norm of the residual values ​​at all spatiotemporal points is used as the physical loss. When the frequency domain tensor is input into the Fourier layer, the dropout operation is applied to the weight matrix in the Fourier layer. A portion of the elements in the weight matrix are randomly set to zero with a preset probability p. The process is repeated a preset number of times to obtain the three-dimensional flow field state tensor predicted at the next time step for a preset number of different steps. Calculate the variance of all predicted 3D flow field state tensors at each node and combine them into a variance map, which serves as a global uncertainty quantification map. Calculate the baseline data loss based on the deviation between the variance map and the baseline flow velocity data. The data loss, physical loss, and baseline data loss are weighted and summed using a pre-defined hyperparameter weight vector to construct a comprehensive loss function.

[0013] Optionally, the Navier-Stokes equations specifically include: Subtract the velocity from the velocity at the previous moment from the velocity at the current moment, and divide by the time interval to obtain an acceleration vector. Multiply the velocity of the node in each of the three dimensions by the rate of change of velocity in that direction, and add the three products together to obtain a vector. Add the vectors obtained in the first two steps one by one to obtain the left-hand side of the Navier-Stokes equations. The pressure difference between the node and its neighbors in three dimensions is calculated and divided by the distance to obtain the pressure difference vector. Then, the vector is divided by the density and multiplied by negative one to obtain the first vector. The velocity in three dimensions is calculated by second difference calculation. The result is divided by the square of the distance and multiplied by the viscosity coefficient to obtain the second vector. The known gravitational acceleration vector is directly used as the third vector. The first, second and third vectors are added digit by digit to obtain the right-hand side of the Navier-Stokes equation.

[0014] Optionally, S5 specifically includes: S51. Read the global uncertainty quantization map. Read the value of each pixel in the global uncertainty quantization map, that is, the velocity prediction variance of the graph node at the same position in the three-dimensional flow field data volume. Perform the inverse transformation on the variance value of each graph node in the global uncertainty quantization map to generate a confidence weight map with the same dimension as the three-dimensional flow field data volume. S52. Multiply the three-dimensional velocity components of each graph node in the three-dimensional flow field data volume by the weight value of the corresponding node in the confidence weight graph point by point to generate a preliminary corrected flow field data volume after confidence weighting. S53. Substitute the velocity and pressure components of each node in the preliminary corrected flow field data volume, along with the spatial coordinates of the nodes, into the Navier-Stokes equations to calculate the residual of each node. Multiply the residual of each node by a preset correction coefficient and subtract the product from the velocity components of the corresponding node in the preliminary corrected flow field data volume to generate a high-precision global flow velocity assessment model.

[0015] Optionally, S6 specifically includes: rendering the evaluation results of the high-precision global flow velocity assessment model into a dynamic three-dimensional flow field visualization image on the user end, providing interactive functions for real-time display and data query at any point, and displaying the global uncertainty quantification map in the form of a semi-transparent heat map to intuitively show the confidence distribution of the flow velocity prediction.

[0016] The beneficial effects of this invention are: First, by simultaneously acquiring visual video streams, underwater acoustic data streams, and baseline flow velocity data, a multimodal fusion data foundation was constructed, effectively overcoming the limitations of a single data source in representing complex flow fields and providing comprehensive and multidimensional information input for high-precision assessment.

[0017] Secondly, based on the pre-trained ST-GNN network, a standardized spatiotemporal dataset is encoded. By constructing a dynamic graph structure and performing spatiotemporal message passing, the water area is discretized into nodes, and its physical evolution process is simulated. This generates a unified feature vector encoding the physical information of the global flow field, significantly enhancing the modeling ability for the spatiotemporal dependencies and complex dynamic characteristics of the flow field. Building upon this, an improved PINN model is introduced, incorporating Fourier neural operators and a comprehensive loss function combining data, physical, and benchmark data losses. This not only uses the Navier-Stokes equations as strong constraints to ensure the physical consistency of the prediction results but also achieves global uncertainty quantification through dropout technology, significantly improving the accuracy, robustness, and reliability of the model's predictions.

[0018] Furthermore, on the user side, the high-precision evaluation results and uncertainty quantification graphs are visualized and rendered, providing real-time display and data query functions, enabling users to intuitively grasp the flow velocity distribution and its confidence level, providing a scientific, intuitive and reliable basis for engineering decisions.

[0019] In summary, this invention achieves high accuracy, high reliability, and interpretability in water flow velocity assessment by integrating the spatiotemporal modeling of ST-GNN with the physical constraint prediction of improved PINN, significantly enhancing its application value in the fields of smart water conservancy and environmental monitoring. Attached Figure Description

[0020] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a smart water flow velocity assessment method based on multimodal fusion proposed in this invention; Figure 2 This is a flowchart of the multimodal spatiotemporal data fusion and unified feature vector generation based on the ST-GNN network proposed in this invention; Figure 3 This is a flowchart of the three-dimensional flow field prediction and global uncertainty quantification based on the improved PINN model proposed in this invention. Detailed Implementation

[0021] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0022] refer to Figures 1-3 A method for intelligent assessment of water flow velocity based on multimodal fusion includes the following steps: S1. Synchronously acquire multimodal data, including visual video stream monitoring the water surface, underwater acoustic data stream, and reference flow velocity data as a moving reference anchor point, and output raw multimodal data; S2. Perform spatiotemporal alignment and standardization preprocessing on the original multimodal data, and perform image enhancement, acoustic computation and multimodal feature extraction to generate a standardized spatiotemporal dataset composed of multimodal feature vectors; S3. Input the standardized spatiotemporal dataset into the pre-trained ST-GNN network, discretize the monitored water area into a dynamic graph structure, simulate the physical evolution of the flow field by performing spatiotemporal message passing on the dynamic graph structure, and output a unified feature vector that encodes the physical information of the global flow field. S4. Input the unified feature vector into the improved PINN model, construct a comprehensive loss function that includes data loss, physical loss and benchmark data loss, perform dynamic calibration using benchmark flow velocity data, and output a three-dimensional flow field data volume and a global uncertainty quantification map. S5. Based on the global uncertainty quantification diagram, the three-dimensional flow field data volume is weighted with confidence and physically corrected to generate a high-precision global water flow velocity assessment model. S6. On the user end, the evaluation results of the high-precision global water flow velocity assessment model are rendered into a visual image, providing interactive functions such as real-time display and data query, and an uncertainty quantification chart is overlaid.

[0023] This implementation significantly improves the accuracy and reliability of flow field assessment in complex waterways. By fusing multimodal data, including visual, acoustic, and benchmark data, and utilizing the ST-GNN network to construct a dynamic graph structure, it achieves efficient spatiotemporal modeling of the physical evolution of water flow, effectively overcoming the limitations of traditional single-point measurements. Introducing an improved PINN model, through embedded physical equation constraints and dynamic calibration with benchmark data, ensures that the prediction results are not only mathematically accurate but also physically reasonable. The output global uncertainty quantification graph intuitively reveals the confidence distribution of the prediction results, providing crucial risk assessment basis for decision-making. The resulting high-precision assessment model maintains stable performance even under complex conditions such as gate scheduling and flood season, providing strong technical support for refined scheduling of water conservancy projects, flood warning, and navigation safety.

[0024] In this embodiment, S1 specifically includes: S11. By deploying a synchronization time block on the unmanned surface vessel, a synchronization control signal with a unified timestamp is generated, which simultaneously triggers the visual video acquisition device, the underwater acoustic acquisition device, and the reference current velocity acquisition device to collect the visual video stream, the underwater acoustic data stream, and the reference current velocity data of the monitored water surface, respectively. S12. The acquired raw visual video stream, underwater acoustic data stream and reference current velocity data are encapsulated in chronological order and output as synchronously acquired raw multimodal data.

[0025] In this embodiment, S2 specifically includes: S21. Read the timestamps of each frame of visual video stream and each reference current velocity data in the original multimodal data. Using the timestamps of the reference current velocity data as a reference, perform time interpolation on the visual video stream. Read the timestamps of the underwater acoustic data stream and each reference current velocity data in the original multimodal data. Using the timestamps of the reference current velocity data as a reference, perform time interpolation on the underwater acoustic data stream. S22. Perform histogram equalization on the time-aligned visual video stream to redistribute pixel grayscale values. Perform Doppler frequency shift calculation on the time-aligned underwater acoustic data stream to convert the frequency offset of the echo signal into water flow velocity value. Linearly scale the visual video stream with image enhancement completed, the underwater acoustic data stream with acoustic calculation completed, and the reference flow velocity data according to a preset uniform value range. S23. Input each frame of the visual video stream after standardization and preprocessing into a pre-trained convolutional neural network, and output a visual feature vector of a set dimension. Input the underwater acoustic data stream after standardization and preprocessing into a pre-trained long short-term memory network, and output an acoustic feature vector of a set dimension. Directly construct a reference data vector of a set dimension from the reference current velocity data after standardization and preprocessing. S24. The visual feature vector, acoustic feature vector and the baseline data vector under the same timestamp are concatenated by dimension to form a multimodal feature vector. The multimodal feature vectors under all timestamps are integrated in chronological order to generate a standardized spatiotemporal dataset.

[0026] In this embodiment, S3 specifically includes: S31. Input the standardized spatiotemporal dataset into the pre-trained ST-GNN network, read the geographic boundary coordinates of the monitored water area, generate a two-dimensional Cartesian coordinate grid, define the geometric center point of each grid cell as a graph node, assign a unique index to each node, and construct the initial graph structure. S32. Assign an independent multilayer perceptron to each graph node. The input layer receives the visual feature vector, acoustic feature vector and reference data vector corresponding to the graph node index in the standardized spatiotemporal dataset. The two hidden layers perform nonlinear transformations with preset weights and biases and concatenate them into an initial feature vector, which is then assigned to the corresponding graph node. S33. Assign a Gaussian kernel function to each edge in the graph structure, read the two-dimensional coordinates of any two nodes and calculate the Euclidean distance, substitute the distance value into the Gaussian kernel function for calculation, use the calculation result as the weight of the edge connecting the two nodes, traverse all node pairs, and construct the spatial adjacency matrix. S34. Combine the initial feature vectors of all nodes into a feature matrix, and input it together with the spatial adjacency matrix into the graph convolutional network layer. By performing matrix multiplication, multiply the spatial adjacency matrix with the feature matrix, perform weighted aggregation on the feature vectors of each node, apply a linear transformation layer and ReLU activation function to generate an updated spatial fusion feature matrix. S35. Input the updated spatial fusion feature matrix into the gated recurrent unit network layer in chronological order. Receive the spatial fusion feature matrix of the current time step and the hidden state of the previous time step. Calculate the hidden state of the current time step through the reset gate and update gate. S36. The hidden state is used as a spatiotemporal fusion feature matrix containing time dependencies, and is then input into the graph convolutional network layer and the gated recurrent unit network layer for a second time. Multiple alternating iterations are performed. In each iteration, the graph convolutional network layer performs spatial feature aggregation, and the gated recurrent unit network layer performs temporal feature update. This process is repeated until the preset number of iterations is reached, and the final spatiotemporal fusion feature matrix is ​​output. S37. Input the final spatiotemporal fusion feature matrix into the global average pooling layer, calculate the arithmetic mean for each dimension along the node dimension, and compress the features of the graph nodes into a one-dimensional vector by arranging them in order and output them as a unified feature vector.

[0027] This implementation method achieves deep modeling and feature extraction of the spatiotemporal evolution of water flow fields by constructing a dynamic graph structure and utilizing an ST-GNN network. The monitored water area is discretized into graph nodes, and multimodal features are fused to effectively capture the spatial continuity and correlation of water flow. Through alternating iterations of graph convolutional networks and gated recurrent units, the model can simultaneously aggregate spatial neighborhood information and learn temporal dependencies, thereby encoding the dynamic physical processes of the global flow field. The output unified feature vector highly condenses the spatiotemporal correlation information in the original data, not only laying a solid foundation for accurate prediction under subsequent physical constraints but also significantly enhancing the model's ability to represent unsteady flow characteristics under complex conditions, effectively solving the prediction bias problem caused by the decoupling of spatiotemporal features in traditional methods.

[0028] In this embodiment, S4 specifically includes: S41. Input the unified feature vector into the improved PINN model, introduce the Fourier neural operator, and construct the comprehensive loss function; S42. Perform backpropagation, starting with the comprehensive loss function, and use the chain rule to calculate the gradient of each learnable parameter in reverse. Use the Adam algorithm to perform an update operation on each parameter based on the calculated gradient, and repeat the iterative steps of forward propagation, backpropagation and parameter update. S43. When the number of iterations reaches the upper limit, stop training and save the improved PINN model after training. Predict the unified feature vector of the new input and output the three-dimensional flow field data volume and the global uncertainty quantization map.

[0029] In this embodiment, S41 specifically includes: The unified feature vector is input into the improved PINN model. A fully connected network with three hidden layers is used, and a decoder with a ReLU activation function is connected after each hidden layer to linearly map the unified feature vector into a two-dimensional tensor. The horizontal dimension of the two-dimensional tensor is the total number of graph nodes, and the column vector is the three-dimensional velocity component of each node. The two-dimensional tensor is the three-dimensional velocity field of all graph nodes. The two-dimensional tensor and the initial pressure field tensor with the same dimension set to zero are concatenated by the last dimension to form the initial three-dimensional flow field state tensor, i.e. the three-dimensional flow field data volume. Then, the Fourier neural operator is introduced to perform a two-dimensional fast Fourier transform on the initial three-dimensional flow field state tensor to transform the flow field state from the spatial domain to the frequency domain, resulting in a frequency domain tensor in complex form. The frequency domain tensor is input into a learnable Fourier layer containing a weight matrix. By performing complex multiplication of the frequency domain tensor and the weight matrix, the different frequency components in the frequency domain are weighted to simulate the global interaction in fluid dynamics. A two-dimensional fast Fourier inverse transform is performed on the weighted frequency domain tensor to convert it back from the frequency domain to the spatial domain, resulting in an updated spatial domain tensor. The updated spatial domain tensor is input into a convolutional layer with a 1×1 kernel. The six features of each node are locally linearly mixed, and the intermediate state tensor is output. It is then added element-wise to the initial three-dimensional flow field state tensor introduced by the Fourier neural operator to form a residual connection. The predicted three-dimensional flow field state tensor for the next time step is output. The predicted three-dimensional flow field state tensor for the next time step is then introduced into the Fourier neural operator. The prediction operation is repeated to form a flow field prediction sequence. In the flow field prediction sequence, the predicted three-dimensional flow field state tensor at a specific time corresponding to the timestamp of the baseline flow velocity data is selected, and the predicted three-dimensional velocity value at the graph node index corresponding to the spatial position of the baseline flow velocity collector is extracted. The predicted three-dimensional velocity value is compared with the measured three-dimensional velocity value of the baseline flow velocity data point by point, the square of the difference between the two is calculated, and the average of the squared differences of all reference points at all corresponding times is taken as the data loss. For each predicted three-dimensional flow field state tensor in the flow field prediction sequence, the velocity and pressure components, along with the spatial coordinates of the graph nodes, are substituted into the discrete form of the Navier-Stokes equations. The difference between the left-hand and right-hand terms of the equations is calculated as the residual, and the square of the L2 norm of the residual values ​​at all spatiotemporal points is used as the physical loss. When the frequency domain tensor is input into the Fourier layer, the dropout operation is applied to the weight matrix in the Fourier layer. A portion of the elements in the weight matrix are randomly set to zero with a preset probability p. The process is repeated a preset number of times to obtain the three-dimensional flow field state tensor predicted at the next time step for a preset number of different steps. Calculate the variance of all predicted 3D flow field state tensors at each node and combine them into a variance map, which serves as a global uncertainty quantification map. Calculate the baseline data loss based on the deviation between the variance map and the baseline flow velocity data. The data loss, physical loss, and baseline data loss are weighted and summed using a pre-defined hyperparameter weight vector to construct a comprehensive loss function.

[0030] In this embodiment, the Navier-Stokes equations specifically include: Subtract the velocity from the velocity at the previous moment from the velocity at the current moment, and divide by the time interval to obtain an acceleration vector. Multiply the velocity of the node in each of the three dimensions by the rate of change of velocity in that direction, and add the three products together to obtain a vector. Add the vectors obtained in the first two steps one by one to obtain the left-hand side of the Navier-Stokes equations. The pressure difference between the node and its neighbors in three dimensions is calculated and divided by the distance to obtain the pressure difference vector. Then, the vector is divided by the density and multiplied by negative one to obtain the first vector. The velocity in three dimensions is calculated by second difference calculation. The result is divided by the square of the distance and multiplied by the viscosity coefficient to obtain the second vector. The known gravitational acceleration vector is directly used as the third vector. The first, second and third vectors are added digit by digit to obtain the right-hand side of the Navier-Stokes equation.

[0031] This implementation achieves high-precision flow field prediction and uncertainty quantification under physical constraints by introducing an improved PINN model and combining it with Fourier neural operators. The unified eigenvector is decoded into an initial three-dimensional flow field state, and the Fourier neural operator is used to capture global interactions in the frequency domain, effectively simulating the complex evolution process of fluid dynamics. By constructing a comprehensive loss function that includes data loss, physical loss, and baseline data loss, and using the Navier-Stokes equations as strong physical constraints, the prediction results are ensured to not only match the measured data but also strictly adhere to the fundamental laws of fluid mechanics. The introduced dropout mechanism further generates a global uncertainty quantification map, intuitively revealing the confidence level of the prediction results. This approach can significantly improve the physical consistency and robustness of predictions under complex conditions such as sparse data or unsteady flow, providing reliable data support for the refined scheduling and safety assessment of water conservancy projects.

[0032] In this embodiment, S5 specifically includes: S51. Read the global uncertainty quantization map. Read the value of each pixel in the global uncertainty quantization map, that is, the velocity prediction variance of the graph node at the same position in the three-dimensional flow field data volume. Perform the inverse transformation on the variance value of each graph node in the global uncertainty quantization map to generate a confidence weight map with the same dimension as the three-dimensional flow field data volume. S52. Multiply the three-dimensional velocity components of each graph node in the three-dimensional flow field data volume by the weight value of the corresponding node in the confidence weight graph point by point to generate a preliminary corrected flow field data volume after confidence weighting. S53. Substitute the velocity and pressure components of each node in the preliminary corrected flow field data volume, along with the spatial coordinates of the nodes, into the Navier-Stokes equations to calculate the residual of each node. Multiply the residual of each node by a preset correction coefficient and subtract the product from the velocity components of the corresponding node in the preliminary corrected flow field data volume to generate a high-precision global flow velocity assessment model.

[0033] In this embodiment, S6 specifically includes: rendering the evaluation results of the high-precision global flow velocity assessment model into a dynamic three-dimensional flow field visualization image on the user end, providing interactive functions for real-time display and data query at any point, and displaying the global uncertainty quantification map in the form of a semi-transparent heat map to intuitively show the confidence distribution of the flow velocity prediction.

[0034] Example 1: To verify the feasibility of this invention in complex aquatic environments, it was deployed in the intelligent hydrological monitoring platform of a key water conservancy project in a certain province. This project spans the main channel of the Qingjiang River, undertaking multiple tasks including flood control, power generation, navigation, and water supply to downstream cities. The water flow conditions in its dam area and downstream river section are complex and variable, influenced by multiple factors such as upstream water inflow, gate scheduling, power generation load, and meteorological conditions. Accurate real-time monitoring of water flow velocity is the core foundation for ensuring dam safety, optimizing power generation efficiency, and providing early warning of flood disasters.

[0035] The platform's monitoring range covers a key river section from 5 kilometers upstream to 15 kilometers downstream of the dam, deploying a total of 47 monitoring points, including shore-based high-definition cameras, unmanned surface monitoring vessels, and ADCPs. It generates over 800GB of video, acoustic, and sensor data daily. Traditional flow assessment methods primarily rely on single-point or cross-sectional measurements from ADCPs, combined with simplified hydrodynamic models for interpolation and extrapolation. This approach suffers from insufficient spatial coverage, delayed data updates, and poor physical consistency of the model. Especially under conditions of frequent gate operations or sudden changes in floodwater during the flood season, traditional methods often fail to capture the dynamic changes of the entire flow field in a timely manner, leading to significantly increased prediction errors and making it difficult to meet the needs of refined scheduling and emergency response.

[0036] In practical deployment, the method of this invention rigorously aligns and standardizes the aforementioned multi-source heterogeneous data in the spatiotemporal dimension, constructing a unified data input format. Through a pre-trained ST-GNN network, the entire monitored water area is discretized into dynamic graph nodes, effectively fusing multimodal information such as visual texture features, acoustic Doppler frequency shift, and measured flow velocity at benchmark points, generating a unified feature vector encoding the spatiotemporal evolution of the global flow field. This vector is then input into an improved PINN model, which enhances its ability to capture high-frequency vortex structures through Fourier neural operators and uses the embedded Navier-Stokes equations as strong physical constraints to ensure the physical rationality of the prediction results. Simultaneously, utilizing Dropout technology, the model provides high-precision three-dimensional flow field predictions while also outputting a global uncertainty quantification map, intuitively displaying the confidence distribution of the prediction results. Table 1 below shows the performance comparison data between the method of this invention and traditional interpolation methods under different typical working conditions during a three-month continuous operation period: Table 1. Performance Comparison Data Between the Invention and Traditional Interpolation Methods

[0037] Based on the comparative data shown in Table 1, it can be seen that the water flow velocity assessment method based on ST-GNN and improved PINN proposed in this invention shows significant performance advantages over traditional interpolation methods in predicting flow fields in complex water areas, especially in terms of prediction accuracy, physical consistency and uncertainty quantification.

[0038] In terms of prediction accuracy, this invention controls the average absolute error to within 0.2 m / s in all three typical operating conditions, far lower than traditional methods. For example, during the most complex flood season, traditional interpolation methods suffer from errors as high as 0.65 m / s due to a lack of characterization of the physical mechanisms of flood waves. In contrast, this invention reduces the error to 0.19 m / s by integrating multi-source data and physical constraints, providing reliable data support for flood control decision-making.

[0039] The advantages of this invention are particularly prominent in terms of physical consistency. Traditional interpolation methods are essentially mathematical fitting, and their prediction results often violate the basic laws of fluid mechanics, with physical consistency scores even falling below 0.5 under complex working conditions. In contrast, this invention uses the Navier-Stokes equations as strong constraints, ensuring that all prediction results strictly follow physical laws, with scores consistently above 0.88, thus guaranteeing the scientific validity and reliability of the prediction results.

[0040] More importantly, this invention is the only method that can provide uncertainty quantification information with a coverage of nearly 90%. It can accurately identify areas with low prediction confidence and remind users to use it with caution. This is a capability that traditional methods do not have at all, which greatly improves the security and practicality of the system in critical decision-making.

[0041] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for intelligent assessment of water flow velocity based on multimodal fusion, characterized in that, Includes the following steps: S1. Synchronously acquire multimodal data and output raw multimodal data; S2. Preprocess the original multimodal data, and perform image enhancement, acoustic computation and multimodal feature extraction to generate a standardized spatiotemporal dataset; S3. Input the standardized spatiotemporal dataset into the pre-trained ST-GNN network, discretize the monitored water area into a dynamic graph structure, simulate the physical evolution of the flow field by performing spatiotemporal message passing on the dynamic graph structure, and output a unified feature vector. S4. Input the unified feature vector into the improved PINN model, introduce the Fourier neural operator, construct a comprehensive loss function, use the benchmark flow velocity data for dynamic calibration, and output a three-dimensional flow field data volume and a global uncertainty quantification map. S5. Based on the global uncertainty quantification diagram, the three-dimensional flow field data volume is weighted with confidence and physically corrected to generate a high-precision global water flow velocity assessment model. S6. On the user end, the evaluation results of the high-precision global water flow velocity assessment model are rendered into a visual image, providing interactive functions such as real-time display and data query, and an uncertainty quantification chart is overlaid.

2. The intelligent assessment method for water flow velocity based on multimodal fusion according to claim 1, characterized in that, S1 specifically includes: S11. By deploying a synchronization time block on the unmanned surface vessel, a synchronization control signal with a unified timestamp is generated, which simultaneously triggers the visual video acquisition device, the underwater acoustic acquisition device, and the reference current velocity acquisition device to collect the visual video stream, the underwater acoustic data stream, and the reference current velocity data of the monitored water surface, respectively. S12. The acquired raw visual video stream, underwater acoustic data stream and reference current velocity data are encapsulated in chronological order and output as synchronously acquired raw multimodal data.

3. The intelligent assessment method for water flow velocity based on multimodal fusion according to claim 1, characterized in that, S2 specifically includes: S21. Read the timestamps of each frame of visual video stream and each reference current velocity data in the original multimodal data. Using the timestamps of the reference current velocity data as a reference, perform time interpolation on the visual video stream. Read the timestamps of the underwater acoustic data stream and each reference current velocity data in the original multimodal data. Using the timestamps of the reference current velocity data as a reference, perform time interpolation on the underwater acoustic data stream. S22. Perform histogram equalization on the time-aligned visual video stream to redistribute pixel grayscale values. Perform Doppler frequency shift calculation on the time-aligned underwater acoustic data stream to convert the frequency offset of the echo signal into water flow velocity value. Linearly scale the visual video stream with image enhancement completed, the underwater acoustic data stream with acoustic calculation completed, and the reference flow velocity data according to a preset uniform value range. S23. Input each frame of the visual video stream after standardization and preprocessing into a pre-trained convolutional neural network, and output a visual feature vector of a set dimension. Input the underwater acoustic data stream after standardization and preprocessing into a pre-trained long short-term memory network, and output an acoustic feature vector of a set dimension. Directly construct a reference data vector of a set dimension from the reference current velocity data after standardization and preprocessing. S24. The visual feature vector, acoustic feature vector and the baseline data vector under the same timestamp are concatenated by dimension to form a multimodal feature vector. The multimodal feature vectors under all timestamps are integrated in chronological order to generate a standardized spatiotemporal dataset.

4. The intelligent assessment method for water flow velocity based on multimodal fusion according to claim 1, characterized in that, S3 specifically includes: S31. Input the standardized spatiotemporal dataset into the pre-trained ST-GNN network, read the geographic boundary coordinates of the monitored water area, generate a two-dimensional Cartesian coordinate grid, define the geometric center point of each grid cell as a graph node, assign a unique index to each node, and construct the initial graph structure. S32. Assign an independent multilayer perceptron to each graph node. The input layer receives the visual feature vector, acoustic feature vector and reference data vector corresponding to the graph node index in the standardized spatiotemporal dataset. The two hidden layers perform nonlinear transformations with preset weights and biases and concatenate them into an initial feature vector, which is then assigned to the corresponding graph node. S33. Assign a Gaussian kernel function to each edge in the graph structure, read the two-dimensional coordinates of any two nodes and calculate the Euclidean distance, substitute the distance value into the Gaussian kernel function for calculation, use the calculation result as the weight of the edge connecting the two nodes, traverse all node pairs, and construct the spatial adjacency matrix. S34. Combine the initial feature vectors of all nodes into a feature matrix, and input it together with the spatial adjacency matrix into the graph convolutional network layer. By performing matrix multiplication, multiply the spatial adjacency matrix with the feature matrix, perform weighted aggregation on the feature vectors of each node, apply a linear transformation layer and ReLU activation function to generate an updated spatial fusion feature matrix. S35. Input the updated spatial fusion feature matrix into the gated recurrent unit network layer in chronological order. Receive the spatial fusion feature matrix of the current time step and the hidden state of the previous time step. Calculate the hidden state of the current time step through the reset gate and update gate. S36. The hidden state is used as a spatiotemporal fusion feature matrix containing time dependencies, and is then input into the graph convolutional network layer and the gated recurrent unit network layer for a second time. Multiple alternating iterations are performed. In each iteration, the graph convolutional network layer performs spatial feature aggregation, and the gated recurrent unit network layer performs temporal feature update. This process is repeated until the preset number of iterations is reached, and the final spatiotemporal fusion feature matrix is ​​output. S37. Input the final spatiotemporal fusion feature matrix into the global average pooling layer, calculate the arithmetic mean for each dimension along the node dimension, and compress the features of the graph nodes into a one-dimensional vector by arranging them in order and output them as a unified feature vector.

5. The intelligent assessment method for water flow velocity based on multimodal fusion according to claim 1, characterized in that, S4 specifically includes: S41. Input the unified feature vector into the improved PINN model, introduce the Fourier neural operator, and construct the comprehensive loss function; S42. Perform backpropagation, starting with the comprehensive loss function, and use the chain rule to calculate the gradient of each learnable parameter in reverse. Use the Adam algorithm to perform an update operation on each parameter based on the calculated gradient, and repeat the iterative steps of forward propagation, backpropagation and parameter update. S43. When the number of iterations reaches the upper limit, stop training and save the improved PINN model after training. Predict the unified feature vector of the new input and output the three-dimensional flow field data volume and the global uncertainty quantization map.

6. The intelligent assessment method for water flow velocity based on multimodal fusion according to claim 5, characterized in that, S41 specifically includes: The unified feature vector is input into the improved PINN model. A fully connected network with three hidden layers is used, and a decoder with a ReLU activation function is connected after each hidden layer to linearly map the unified feature vector into a two-dimensional tensor. The horizontal dimension of the two-dimensional tensor is the total number of graph nodes, and the column vector is the three-dimensional velocity component of each node. The two-dimensional tensor is the three-dimensional velocity field of all graph nodes. The two-dimensional tensor and the initial pressure field tensor with the same dimension set to zero are concatenated by the last dimension to form the initial three-dimensional flow field state tensor, i.e. the three-dimensional flow field data volume. Then, the Fourier neural operator is introduced to perform a two-dimensional fast Fourier transform on the initial three-dimensional flow field state tensor to transform the flow field state from the spatial domain to the frequency domain, resulting in a frequency domain tensor in complex form. The frequency domain tensor is input into a learnable Fourier layer containing a weight matrix. By performing complex multiplication of the frequency domain tensor and the weight matrix, the different frequency components in the frequency domain are weighted to simulate the global interaction in fluid dynamics. A two-dimensional fast Fourier inverse transform is performed on the weighted frequency domain tensor to convert it back from the frequency domain to the spatial domain, resulting in an updated spatial domain tensor. The updated spatial domain tensor is input into a convolutional layer with a 1×1 kernel. The six features of each node are locally linearly mixed, and the intermediate state tensor is output. It is then added element-wise to the initial three-dimensional flow field state tensor introduced by the Fourier neural operator to form a residual connection. The predicted three-dimensional flow field state tensor for the next time step is output. The predicted three-dimensional flow field state tensor for the next time step is then introduced into the Fourier neural operator. The prediction operation is repeated to form a flow field prediction sequence. In the flow field prediction sequence, the predicted three-dimensional flow field state tensor at a specific time corresponding to the timestamp of the baseline flow velocity data is selected, and the predicted three-dimensional velocity value at the graph node index corresponding to the spatial position of the baseline flow velocity collector is extracted. The predicted three-dimensional velocity value is compared with the measured three-dimensional velocity value of the baseline flow velocity data point by point, the square of the difference between the two is calculated, and the average of the squared differences of all reference points at all corresponding times is taken as the data loss. For each predicted three-dimensional flow field state tensor in the flow field prediction sequence, the velocity and pressure components, along with the spatial coordinates of the graph nodes, are substituted into the discrete form of the Navier-Stokes equations. The difference between the left-hand and right-hand terms of the equations is calculated as the residual, and the square of the L2 norm of the residual values ​​at all spatiotemporal points is used as the physical loss. When the frequency domain tensor is input into the Fourier layer, the dropout operation is applied to the weight matrix in the Fourier layer. A portion of the elements in the weight matrix are randomly set to zero with a preset probability p. The process is repeated a preset number of times to obtain the three-dimensional flow field state tensor predicted at the next time step for a preset number of different steps. Calculate the variance of all predicted 3D flow field state tensors at each node and combine them into a variance map, which serves as a global uncertainty quantification map. Calculate the baseline data loss based on the deviation between the variance map and the baseline flow velocity data. The data loss, physical loss, and baseline data loss are weighted and summed using a pre-defined hyperparameter weight vector to construct a comprehensive loss function.

7. The intelligent assessment method for water flow velocity based on multimodal fusion according to claim 6, characterized in that, The Navier-Stokes equations specifically include: Subtract the velocity from the velocity at the previous moment from the velocity at the current moment, and divide by the time interval to obtain an acceleration vector. Multiply the velocity of the node in each of the three dimensions by the rate of change of velocity in that direction, and add the three products together to obtain a vector. Add the vectors obtained in the first two steps one by one to obtain the left-hand side of the Navier-Stokes equations. The pressure difference between the node and its neighbors in three dimensions is calculated and divided by the distance to obtain the pressure difference vector. Then, the vector is divided by the density and multiplied by negative one to obtain the first vector. The velocity in three dimensions is calculated by second difference calculation. The result is divided by the square of the distance and multiplied by the viscosity coefficient to obtain the second vector. The known gravitational acceleration vector is directly used as the third vector. The first, second and third vectors are added digit by digit to obtain the right-hand side of the Navier-Stokes equation.

8. The intelligent assessment method for water flow velocity based on multimodal fusion according to claim 1, characterized in that, S5 specifically includes: S51. Read the global uncertainty quantization map. Read the value of each pixel in the global uncertainty quantization map, that is, the velocity prediction variance of the graph node at the same position in the three-dimensional flow field data volume. Perform the inverse transformation on the variance value of each graph node in the global uncertainty quantization map to generate a confidence weight map with the same dimension as the three-dimensional flow field data volume. S52. Multiply the three-dimensional velocity components of each graph node in the three-dimensional flow field data volume by the weight value of the corresponding node in the confidence weight graph point by point to generate a preliminary corrected flow field data volume after confidence weighting. S53. Substitute the velocity and pressure components of each node in the preliminary corrected flow field data volume, along with the spatial coordinates of the nodes, into the Navier-Stokes equations to calculate the residual of each node. Multiply the residual of each node by a preset correction coefficient and subtract the product from the velocity components of the corresponding node in the preliminary corrected flow field data volume to generate a high-precision global flow velocity assessment model.

9. The intelligent assessment method for water flow velocity based on multimodal fusion according to claim 1, characterized in that, S6 specifically includes: rendering the evaluation results of the high-precision global water flow velocity assessment model into a dynamic three-dimensional flow field visualization image on the user end, providing interactive functions for real-time display and data query at any point, and displaying the global uncertainty quantification map in the form of a semi-transparent heat map to intuitively show the confidence distribution of the flow velocity prediction.