A multi-scale res2raft particle image velocimetry method based on global attention mechanism optimization
By introducing the Bottle2neck module of Res2Net and the Res2RAFT model with the global attention mechanism of GAM, the problem of insufficient prediction of fine-grained displacement fields in complex flow fields is solved, and high-precision flow field observation is achieved, which is suitable for ecological monitoring and flow field research in underwater environments.
Patent Information
- Application Number
- CN202411816519.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2044-12-11
AI Technical Summary
Existing particle image velocimetry methods struggle to capture subtle flow structures in complex flow fields, especially underwater or turbulent environments. Traditional methods suffer from insufficient measurement accuracy and adaptability, while deep learning models have limitations in multi-scale feature extraction and fine-grained prediction.
The Bottle2neck module of Res2Net is used to enhance multi-scale feature extraction, and combined with the GAM global attention mechanism, a Res2RAFT model is constructed. Through multi-scale feature extraction and interaction with global information, the accuracy and adaptability of flow field prediction are improved.
It achieves high-precision displacement field prediction for complex flow fields, especially in underwater environments where it can more accurately capture the motion state of particles, improving the resolution and detail preservation of flow field prediction, and is suitable for underwater ecological monitoring and flow field research.
Smart Images

Figure CN119716136B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of particle image velocimetry (PIV), and particularly relates to a multi-scale Res2RAFT particle image velocimetry method based on global attention mechanism optimization. BACKGROUND
[0002] Particle image velocimetry (PIV) is a commonly used flow field measurement method in experimental fluid mechanics, and is widely used to reveal the details of fluid motion. However, the traditional PIV method has limitations in terms of measurement accuracy and adaptability in complex flow fields, especially in underwater or turbulent environments, due to insufficient lighting and irregular flow fields, making it difficult for traditional cross-correlation methods and variational optical flow methods to capture subtle flow structures.
[0003] Cross-correlation method is one of the most commonly used methods in PIV, which estimates the displacement field of fluid by finding the maximum correlation of local windows in consecutive image pairs. This method works well in simple flow fields, but due to the fixed window size, it is difficult to adapt to different scales of flow structures, resulting in subtle flow structures in complex flow fields often being unable to be effectively captured. At the same time, the velocity field generated by this method is relatively sparse, making it difficult to achieve fine-grained tracking of fluid. Variational optical flow method solves the flow field displacement through partial differential equations, which can generate dense velocity field. However, the calculation process of this method is complex, with high computational cost, and errors are prone to occur in areas with sharp gradient changes in the flow field. Variational optical flow method usually relies on physical assumptions and prior knowledge, which is difficult to meet the dynamic characteristics of complex flow fields in practical applications, especially in high gradient and nonlinear flow fields, which is prone to lose accuracy.
[0004] In recent years, with the development of deep learning, optical flow networks have been gradually introduced into PIV tasks, but existing networks still face many challenges in fine-grained prediction in complex flow fields.
[0005] With the development of computer vision and deep learning, convolutional neural networks (CNN) are widely used in PIV tasks, processing particle images in an end-to-end manner. These deep learning models, such as FlowNet and LiteFlowNet, have improved optical flow prediction to a certain level, but there are still some shortcomings: FlowNet series networks introduce end-to-end convolutional network structure in optical flow estimation, with high dense optical flow prediction ability. However, the structure of FlowNet2 is complex, consumes a lot of computing resources, and takes a long time to train the model, while its ability to capture multi-scale information is limited, resulting in insufficient prediction accuracy and generalization ability when dealing with complex turbulent flow fields. LiteFlowNet adopts lightweight design to reduce computational cost, but due to the reduction of parameter quantity, the model performs poorly in high-resolution fine-grained prediction, and it is difficult to balance prediction accuracy and model efficiency, especially in complex flow fields, the ability to capture subtle structures is weak. RAFT is a recursive optical flow estimation network that calculates correlation volumes through all-to-all feature matching, achieving efficient and accurate flow field estimation. Although RAFT performs well in dense optical flow estimation, its multi-scale feature extraction capability is limited, and there is still room for improvement in predicting complex and fine-grained flow structures, especially in flow fields with multi-scale vortex structures, the model's local feature expression is insufficient.
[0006] Res2Net network is a convolutional neural network with multi-scale feature extraction capability. Its Bottle2neck module achieves multi-scale feature extraction and fusion through fine-grained feature division and parallel convolution operation. The structural innovation of Bottle2neck lies in the following two points: first, multi-scale parallel convolution: Bottle2neck module divides the input features into several scales, performs parallel convolution operation on each scale, and fuses features of different scales. This approach improves the model's ability to capture fine-grained flow information, even when the image resolution is low, it can still retain rich multi-scale features. Second, feature layer combination: the module gradually combines feature maps of different scales in convolution operation, so that the network maximizes the use of multi-scale information while maintaining computational efficiency. This feature combination method can help the model maintain sensitivity to local and global information when processing multi-scale motion information in particle images, and adapt to the prediction needs of complex flow fields in PIV tasks.
[0007] GAM is a channel and space mixed global attention mechanism that can enhance the sensitivity of the model to feature selection. Its core lies in the selective enhancement of input features through channel attention and spatial attention. The channel attention mechanism strengthens the channel dependency between features by weighting the channel dimension of the feature map. In the PIV task, channel attention can effectively distinguish the flow field features in different particle images and suppress unimportant feature information. Spatial attention weights the spatial dimension of the feature map, focusing on the significant feature area in the flow field to further optimize the model's attention to flow characteristics. The channel and spatial attention of GAM work together to effectively focus global information on important features, significantly improving the ability to capture details in irregular areas such as vortexes and turbulence in complex flow fields.
[0008] The existing optical flow network models, such as FlowNet, LiteFlowNet and RAFT, have shown certain optical flow estimation performance, but there is still room for improvement in precision and multi-scale feature extraction capability. Therefore, it is of great significance to propose a multi-scale optical flow estimation network based on deep learning to better handle complex flow fields. SUMMARY
[0009] The purpose of the present application is to provide a high-precision optical flow estimation method for particle image velocimetry (PIV), namely a Res2RAFT model based on Res2Net multi-scale feature extraction and GAM global attention mechanism. This method aims to solve the problem of insufficient prediction of fine-grained displacement field in complex flow fields by existing optical flow networks, especially in multi-scale vortex and turbulent flow fields. By designing an optical flow estimation structure with multi-scale feature extraction and global information interaction, Res2RAFT can more accurately capture the subtle flow characteristics in complex flow fields, improving the flow field prediction accuracy and adaptability in PIV tasks.
[0010] Particle image velocimetry (PIV) is an important means of studying flow field distribution in experimental fluid mechanics, providing key technical support for revealing fluid dynamics and velocity field distribution. In particular, in the complex underwater flow field environment of deep-sea plume, PIV method helps to observe particle flow and diffusion behavior, which has important significance for studying the state of underwater particle flow field and its impact on deep-sea ecological environment. The particle motion characteristics in deep-sea plume are complex and the flow structure is subtle, and the traditional PIV method has limitations in resolution and accuracy when observing such fine-grained dynamics.
[0011] The Res2RAFT model of the application enhances the multi-scale feature extraction capability by introducing a Bottle2neck module in the Res2Net network, and combines a GAM global attention mechanism to accurately capture the complex flow field underwater. The method can effectively extract the flow characteristics in the particle image at different scales, thereby more accurately predicting the displacement field distribution in the flow field. Res2RAFT will provide an efficient and reliable particle image velocimetry solution for complex underwater flow fields including deep-sea plumes, meeting the detailed observation needs of the underwater particle flow field state, and providing technical support for marine environment monitoring and protection.
[0012] The application provides a multi-scale Res2RAFT particle image velocimetry method based on global attention mechanism optimization, comprising the following steps:
[0013] S1, obtaining a PIV data set and data preprocessing:
[0014] Each data sample in the PIV data set comprises two consecutive frames of particle images and the particle true speed corresponding to the particle images;
[0015] S2, constructing a particle image velocimetry model, wherein the particle image velocimetry model is based on a RAFT optical flow network,
[0016] The RAFT optical flow network comprises a feature extraction network, a context network, a full correlation layer, and a GRU recursive update module;
[0017] The feature extraction network and the context network of the RAFT optical flow network are replaced by a Res2encoder network,
[0018] to obtain a particle image velocimetry model;
[0019] The Res2Encoder network comprises, in sequence, a convolutional layer one, a RA-Bottle2neck module, a RB-
[0020] Bottle2neck module, RA-Bottle2neck module, RB-Bottle2neck module, RA-Bottle2neck module, RB-Bottle2neck module, and a convolutional layer two;
[0021] The convolutional layer one uses a 7x7 convolutional layer to generate an initial feature map of 64 channels;
[0022] The convolutional layer two uses a 1x1 convolution to generate an output feature map of 128 channels;
[0023] The RA-Bottle2neck module is a bottleneck structure based on the Bottle2neck structure with residual learning, used for down-sampling to maintain multi-scale information, and a multi-scale feature processing method is introduced in each residual block, and the input feature mapping of the Res2Encoder module is divided into s subsets, as follows, and each subset is independently subjected to convolution operation to extract features at several scales to increase the diversity of the feature space:
[0024]
[0025] For the first subset x i , no convolution operation is performed, and the original input is directly retained as the output y i ; the second subset x2 is subjected to a standard 3x3 convolution operation K2(x2) to obtain the output y2; starting from the third subset x3, the convolution operation not only depends on the input x3 of the group, but also combines the output y2 of the previous group, that is, K3(x3+y2), thereby layer-by-layer fusing information from different scales; the layer-by-layer cumulative and recursive feature extraction method continues until the processing of the last subset x s is completed.
[0026] The RB-Bottle2neck module is a bottleneck structure based on the Bottle2neck structure with a global attention mechanism, used for maintaining resolution; the RB-Bottle2neck module adds a global attention mechanism GAM based on the RA-Bottle2neck module, and the global attention mechanism GAM is added after the last layer of the RA-Bottle2neck module; and the global attention mechanism GAM in the RB-Bottle2neck module includes a channel attention and a spatial attention module, which utilizes the interaction of channel and spatial dimensions to make the particle image velocimetry model focus on the particle flow information in the two consecutive frames;
[0027] S3 constructs a loss function, and uses the PIV data set in step S1 to train the particle image velocimetry model in a supervised learning manner.
[0028] Preferably, in step S1, the data preprocessing specifically includes: uniformly processing the particle image size in the PIV data set, setting the length and width size to 256x256, and then performing regularization processing on the image to ensure that the input of the Res2Encoder is adapted.
[0029] Preferably, in step S2, s=4, and the stride and baseWidth parameters of the RA-Bottle2neck module are set to 2 and 26, respectively.
[0030] Preferably, in step S2, the stride and baseWidth parameters of the RB-BottleNeck module are set to 1 and 26 respectively.
[0031] Preferably, in step S3, the loss function is as follows:
[0032]
[0033] where n is the number of samples in the PIV dataset, F i is the optical flow prediction value obtained by inputting the particle images of two consecutive frames of samples in the PIV dataset into the particle image velocimetry model, F gt is the true optical flow value reflecting the true velocity of the particles in the samples in the PIV dataset.
[0034] Preferably, λ i = 0.8 n-i is the weight of each prediction, giving greater weight to later predictions to improve the accuracy of the final result.
[0035] Preferably, the full correlation layer is used to perform full correlation calculation on the features extracted by the feature extraction network to obtain a four-layer correlation pyramid.
[0036] The full correlation layer establishes the correlation between pixels by calculating the similarity between each pixel in two frames. A 4D correlation volume is constructed to perform similarity calculation according to the following formula, which saves the inner product of the feature vectors between all pixel pairs:
[0037] C(g θ (I1),g θ (I2))∈R H×W×H×W ,C ijkl =∑hg θ (I1)ijh·g θ (I2)klh
[0038] In the formula, C is the correlation volume, g θ (I1) and g θ (I2) are the feature maps extracted by the feature encoder from the two frames of images.
[0039] C ijkl represents the weight coefficient of the i, j spatial dimension (such as the width and height of the convolution kernel) and the k, l channel dimension in the convolution kernel.
[0040] h is the input of the convolution calculation, which is the value of the local region of the input image or the feature map of the previous layer. It is a part of the convolution operation, through which the convolution kernel extracts features.
[0041] The formula indicates that each pixel pair (i.e., (i,j) and (k,l)) in the two images obtains a similarity score by calculating an inner product;
[0042] The four-layer correlation pyramid is constructed by performing a pooling operation on the correlation volume.
[0043] Preferably, the GRU recursive update module uses a convolutional GRU and a separable convolutional GRU to recursively update the features, and gradually adjusts the flow field prediction.
[0044] The current optical flow estimation and the hidden state are retained by the update gate and the reset gate, and the optical flow increment is calculated by combining the current features, and the optical flow is iteratively updated;
[0045] The separable convolutional GRU captures more fine-grained feature information through horizontal and vertical separable convolution, so that the optical flow estimation is more accurate.
[0046] The present application has the beneficial effects:
[0047] The present application provides a high-precision optical flow estimation method for particle image velocimetry (PIV), which enhances the model's perception ability of complex flow fields through multi-scale feature extraction and global attention mechanism. The method includes data preprocessing, Res2Encoder feature extraction, full correlation calculation, GRU recursive optimization, upsampling, and high-resolution displacement field output processes. It is particularly suitable for particle displacement field prediction in underwater complex environments, and can cope with common challenges such as insufficient underwater lighting, impurity interference, and irregular flow. It provides a reliable solution for high-precision observation of complex flow fields. Through this method, the model can more accurately capture the motion state of particles when dealing with complex underwater flow fields (such as deep-sea plumes), effectively improving the resolution and detail retention ability of flow field prediction. The present application is suitable for application in the fields of ecological monitoring and flow field research in underwater environments, providing higher observation accuracy and reliability.
[0048] The present application adopts the Bottle2neck module of Res2Net to construct Res2Encoder, and uses multi-scale feature extraction to process the flow information in the particle image in layers. By dividing the input features into different scales and performing parallel convolution processing, this module can effectively extract fine-grained features in the flow field at different scales. This multi-scale structure can ensure that the model can capture multi-level information from large-scale overall flow trends to small-scale vortexes when dealing with complex flow. The fusion of multi-scale features makes the model have higher adaptability and accuracy when dealing with complex flow, effectively improving the ability to capture detailed information.
[0049] The GAM global attention mechanism is introduced into the model to improve the selectivity of important features through the combination of channel and spatial attention. In the PIV task, the particle image contains a large amount of redundant information, and the GAM mechanism can focus on the key flow area and suppress irrelevant features to avoid information dispersion. The channel attention mechanism weights the features in the channel dimension, enhancing the dependency relationship between different channel features; while the spatial attention focuses on the significant area in the image by weighting the features in the spatial dimension. The introduction of GAM enables the Res2RAFT model to better adapt to the flow pattern changes in high-complexity flow fields, especially in dynamic and dramatic scenes such as turbulence and vortex.
[0050] The Res2RAFT model of the present application significantly improves the practicability of the PIV task in underwater complex environments through the synergistic effect of multi-scale feature extraction and global attention mechanism. The model is suitable for various application scenarios in fluid mechanics, such as simulated flow field testing in laboratory environments, deep-sea particle flow observation in marine ecological monitoring, and complex flow field evaluation in industrial applications. Especially in scenes with complex flow characteristics such as deep-sea plume and near-sea plankton particle flow, the present application can provide accurate flow field observation data, which helps to reveal the details of fluid motion and provides technical support for scientific research and environmental protection. In summary, the Res2RAFT particle image velocimetry method proposed in the present application overcomes the shortcomings of traditional PIV methods in handling underwater complex flow fields, realizes high-resolution displacement field prediction of flow fields through fine multi-scale feature extraction and global attention mechanism, and has wide application prospects. BRIEF DESCRIPTION OF DRAWINGS
[0051] Figure 1 : Overall structure diagram of Res2RAFT.
[0052] Figure 2 : Hierarchical structure of Res2Encoder.
[0053] Figure 3 : Multi-scale convolution structure diagram of RA-Bottle2neck module.
[0054] Figure 4 : Multi-scale convolution structure diagram of RB-Bottle2neck module.
[0055] Figure 5 : Structure diagram of GAM global attention mechanism in RB-Bottle2neck module.
[0056] Figure 6 : Prediction result diagram of Reynolds number 200.
[0057] Figure 7: Comparison of experimental results of Res2RAFT with traditional methods.
[0058] Figure 8 : Comparison of experimental results of Res2RAFT with deep learning methods. DETAILED DESCRIPTION
[0059] A multi-scale Res2RAFT particle image velocimetry method based on global attention mechanism optimization, the steps of the present application: preparing a data set and preprocessing; introducing the Bottle2neck module in the Res2Net network to enhance the multi-scale feature extraction capability, and introducing the GAM (global attention mechanism) to improve the global dimension interaction of the features and reduce information diffusion. Subsequently, through the correlation calculation layer, the GRU recursive update module, the up-sampling module and other steps, the improved model Res2RAFT optical flow prediction network can effectively recover the high-resolution optical flow field from the low-resolution features, and improve the accuracy and adaptability of the flow field measurement. Including the following steps:
[0060] Step 1, data set selection and preprocessing:
[0061] In this study, a data set specially used to evaluate the performance of particle image velocimetry, namely PIV Dataset artificially synthesized by Cai et al. using computational fluid dynamics technology, is selected. As shown in Table 1, a total of five categories of 11650 pairs of PIV particle images and standard flow field results are included. Each pair of data in the set includes two consecutive frames of particle images and the true velocity results corresponding to the particle image pair. The data set generates images of various flow processes according to computational fluid dynamics. The detailed information of the PIV Dataset data set is shown in Table 1.
[0062] Table 1: Composition of data set
[0063]
[0064] By obtaining the PIV data set containing two consecutive frames of particle images, the image size of the data set is uniformly processed, and the length and width size are all set to 256x256. The images are normalized to ensure that they are suitable for input into the Res2Encoder. The cross-validation method is used in the experiment, and the data set is divided into ten parts, nine of which are training sets and one of which is a test set. Various classification indicators are calculated to analyze the performance of the algorithm.
[0065] The average end point error is used to measure the accuracy of the model flow field prediction. The calculation formula of the average end point error (AEE) is as follows:
[0066]
[0067] In the formula, N represents the total number of pixels. fgt represents the true optical flow of each pixel (ground truth optical flow), fes,i represents the optical flow estimated in the i-th iteration. ||. || represents the L1 norm, that is, the absolute error of the optical flow. The AEE value is too small to be directly taken, and it is not easy to compare, so the error value of every 100 pixels is taken for comparison in this paper, which is more clear and explicit.
[0068] Regarding loss calculation, through supervised learning, the loss between the optical flow prediction result and the true displacement field is calculated, and the model parameters are updated through back propagation. The sequence loss function is used for loss calculation. Specifically, the difference between the optical flow prediction sequence {F1, F2, …, Fn} and the true optical flow Fgt is measured by weighted L1 loss
[0069]
[0070] In the formula, λ i = 0.8 n-i is the weight of each prediction, which gives more weight to the later prediction to improve the accuracy of the final result.
[0071] Step 2, feature extraction network Res2Encoder:
[0072] The present application mainly proposes a new Res2Encoder feature encoder, which simultaneously acts on feature extraction and context network. Res2Encoder uses a unified structure design in the feature encoder and context network, which is composed of multiple RA-Bottle2neck and RB-Bottle2neck modules. Among them, the RA-Bottle2neck module mainly realizes resolution reduction through the design of stride=2, enhancing the expression ability of multi-scale features. The RB-Bottle2neck module enhances the feature selectivity and understanding ability of global information by preserving the resolution and combining GAM (global attention mechanism). This modular design makes the function realization of Res2Encoder in the two sub-networks consistent and scalable.
[0073] The task of the feature encoder is to extract rich multi-scale features for the input image pair. The specific functions include:
[0074] (1) Multi-scale feature extraction
[0075] Through the RA-Bottle2neck module, Res2Encoder can downsample the input image pair layer by layer, so as to capture features at different spatial scales. This is particularly important for the unique multi-scale flow patterns in particle images, because the feature encoder needs to provide comprehensive feature information for optical flow estimation.
[0076] (2) Enhance feature expression ability
[0077] The RB-Bottle2neck module further filters and optimizes multi-scale features through the GAM global attention mechanism, highlighting key features. This process not only enhances the model's understanding of global information but also preserves necessary detail information during the process of reducing feature resolution layer by layer.
[0078] (3) Provide high-quality features for optical flow estimation
[0079] The final output of the feature encoder is a multi-channel, high-dimensional feature map. These feature maps are input to the correlation layer and subsequent modules, providing strong support for the optical flow estimation task.
[0080] The task of the context network is to use global context information to further optimize the optical flow estimation result. The role of Res2Encoder in it mainly reflects in:
[0081] (1) Maintain high-resolution features
[0082] Res2Encoder in the context network maintains the spatial resolution of the feature map (stride = 1 in RA-Bottle2neck). This design ensures that the model can better capture local detail features and provide support for fine-grained optical flow estimation.
[0083] (2) Strengthen global information interaction
[0084] The GAM mechanism in the RB-Bottle2neck module can strengthen the global semantic relationship of feature expression through channel attention and spatial attention. The context network relies on this ability to model the potential global dependencies in the optical flow estimation process, thereby improving the estimation accuracy.
[0085] (3) Provide global optimization ability
[0086] The context network further processes the feature maps from the feature encoder, providing global optimization support for the optical flow estimation process, making the model more robust when dealing with complex flow fields.
[0087] Although the Res2Encoder structure is consistent in the feature encoder and the context network, its functional positioning is different: in the feature encoder, the core of Res2Encoder is to extract multi-scale features while gradually reducing the sampling to obtain more efficient feature representation. In the context network, Res2Encoder further optimizes high-resolution features through global attention mechanism, enhancing the model's ability to capture global dependencies and complex flow patterns. This unified but functionally complementary design makes Res2Encoder a core module connecting feature extraction and global optimization, providing an important guarantee for the overall performance improvement of the model.
[0088] In the design of the feature extraction network Res2Encoder of the present application, the network architecture is mainly composed of three levels of modules, as shown in Figure 2 Each level is composed of RA-Bottle2neck and RB-Bottle2neck modules, which are responsible for different feature processing tasks, the former is used for downsampling to maintain multi-scale information, and the latter is used to maintain resolution and introduce global attention. Thus, the multi-scale features of the particle image are effectively extracted. In the overall architecture, the input image is first processed by a 7x7 convolutional layer to generate a 64-channel initial feature map, then it is processed by three layers of RA-Bottle2neck and RB-Bottle2neck modules in turn, and finally a 128-channel output feature map is generated by a 1x1 convolution. This hierarchical structure uses the RA-Bottle2neck module for downsampling, while the RB-Bottle2neck module is responsible for maintaining resolution and introducing global attention mechanism, forming the gradual refinement and enhancement of features.
[0089] For Res2Encoder, the core design idea is to enhance the feature extraction capability of the network through multi-scale feature representation. Traditional convolutional networks usually use fixed size convolution kernels at each layer to extract features, while Res2Encoder modules introduce multi-scale feature processing in each residual block. Specifically, the input feature map is divided into multiple subsets, each of which can be independently convolved and extract features at different scales, which greatly increases the diversity in the feature space.
[0090]
[0091] As shown in the formula, the input feature map is divided into s subsets x i For the first subset x i , no convolution operation is performed, and the original input is directly retained as the output y iThe next subset x2 goes through a standard 3x3 convolution operation K2(x2) to get the output y2. From the third subset x3, the convolution operation not only depends on the input x3, but also combines the output y2 of the previous group, i.e., K3(x3+y2), so that the information from different scales is fused layer by layer. This way of layer-by-layer accumulation and recursive feature extraction continues until the processing of the last subset x s is completed.
[0092] The key advantage of this design is that each 3x3 convolution kernel not only performs a simple convolution operation on the current input features, but also incorporates the features of the previous layer, effectively expanding the receptive field. The expansion of the receptive field means that the network can capture more rich spatial information at different scales, especially when dealing with complex visual tasks. The multi-scale feature expression ability of the network makes it better adapt to different object sizes, motion patterns and structural complexity.
[0093] In addition, through the step-by-step transmission of features, multiple "equivalent receptive fields" are formed, that is, the network can capture and fuse features at different scales, and this combination effect effectively improves the modeling ability of input information. For example, when dealing with the particle image velocimetry (PIV) task in fluid mechanics, the multi-scale characteristics of the flow field (from large-scale overall flow to small-scale turbulent structure) can be captured and modeled in this way.
[0094] From the perspective of parameter optimization, the number of direct 3x3 convolution kernel operations is also reduced, significantly reducing the number of network parameters. Compared with directly using complete convolution operations at each layer, Res2Net uses a layer-by-layer feature transmission method, greatly reducing the computational burden while maintaining the integrity of the information. Due to the close relationship between feature groups, input features are efficiently converted into output features, thereby improving the overall performance of the network without increasing too much computational overhead.
[0095] scale is a new dimension proposed by Bottle2neck, which is orthogonal to traditional dimensions such as depth, width and cardinality. By processing multi-scale features at a fine-grained level, Res2Encoder not only improves the performance of the network, but also enables it to better handle complex visual tasks.
[0096] (1) RA-Bottle2neck
[0097] The design of the RA-Bottle2neck module is as follows Figure 3The scale parameter of the RA-Bottle2neck module is 4, and the baseWidth is set to 26. This module divides the input features into multiple scales (x1 to x4 in the figure) and performs parallel convolution operations on these scales. The outputs of each scale are fused after convolution, respectively, to form output features (y1 to y4), and finally these features are reassembled to the main branch. By setting stride = 2, the RA-Bottle2neck module reduces the resolution of the input feature map, so that the multi-scale information can still be preserved during downsampling. This multi-scale convolution structure can enhance the model's feature expression ability at different scales, ensuring that enough local and global information can be captured even when the resolution is reduced.
[0098] Unlike ordinary residual connections, the RA-Bottle2neck module introduces a more fine-grained feature segmentation and processing method. Through multi-scale processing, this module can maximize the model's ability to capture information of various scales in the input image while maintaining computational efficiency. This multi-scale feature extraction is very important for flow information in particle images, because flow information in particle images often exhibits multiple different scale changes.
[0099] (2) RB-Bottle2neck
[0100] The design of the RB-Bottle2neck module is shown in Figure 4 , which is characterized by introducing a GAM global attention mechanism to enhance the selectivity of features. This module extracts local features through multi-scale convolution operations while keeping the feature map resolution unchanged, and finally uses the GAM attention mechanism at the end to optimize the features. As shown in Figure 5 , the GAM mechanism first strengthens the dependence between features in the channel through channel attention, and then focuses on important spatial regions through a spatial attention module.
[0101] The introduction of the GAM mechanism in the RB-Bottle2neck improves the understanding of global information, enabling the model to better capture complex flow patterns in particle images. Since the information in the particle image flow field has complex global dependencies, the role of the GAM global attention mechanism is to suppress unimportant features through the interaction of channel and spatial dimensions, so that the model can focus on the most critical flow information in the image.
[0102] In the RA-Bottle2neck and RB-Bottle2neck modules designed in the present application, the scale parameter is set to 4 and the baseWidth is set to 26, which has sufficient theoretical basis and experimental verification support. Specifically, multi-scale feature extraction is the key to improving model performance, and the size of the scale parameter directly determines the granularity of feature segmentation and the effectiveness of multi-scale convolution operation. However, a too large scale value will significantly increase the computational complexity of the model, not only reducing the training and inference efficiency, but also possibly introducing too much redundant information, leading to overfitting problems of the model on a specific dataset. At the same time, a too small scale value cannot fully play the role of multi-scale feature extraction, so as to effectively improve the performance of the model.
[0103] After multiple experimental comparisons, we choose to set scale to 4, which strikes a good balance between the effectiveness of feature extraction and the complexity of the model. In addition, the baseWidth parameter is set to 26, which further controls the computational burden of each layer of convolution, making the model have higher efficiency while maintaining the ability of multi-scale feature extraction.
[0104] To further reduce the risk of overfitting, we introduce the Dropout mechanism in the design process and set its value to 0.2. Dropout, as a commonly used regularization method, can randomly discard part of the neuron connections to prevent the model from overfitting the training data. This setting not only meets the demand of improving the robustness of the model in theory, but also has been verified to be effective through a large number of experiments. The experimental results show that when the scale parameter is 4, the baseWidth is 26, and the Dropout value is 0.2, the model performs best on different types of datasets. This design enables the model to achieve the best balance between multi-scale feature extraction and computational efficiency, while significantly reducing the possibility of overfitting, laying a foundation for accurate prediction of complex flow fields.
[0105] Step 3, Correlation Calculation Layer
[0106] Correlation Layer (All-pair Correlation Layer): Perform all-pair correlation calculation on the extracted features to generate a four-layer correlation pyramid (as shown in the four-layer structure in Figure 1 The correlation layer plays a crucial role. It establishes the association between pixels by calculating the similarity between each pixel in two frames of images. This process is accomplished by constructing a 4D correlation volume, which stores the inner product of feature vectors between all pairs of pixels.
[0107] As shown in the following equation:
[0108] C(g θ (I1),g θ (I2))∈R H×W×H×W ,C ijkl =∑hg θ (I1)ijh·g θ (I2)klh (4)
[0109] In the equation, C is the correlation volume, g θ (I1) and g θ (I2) are the feature maps extracted by the feature encoder from the two frames of images. The equation shows that each pixel pair (i.e., (i,j) and (k,l)) in the two images gets a similarity score by computing the inner product.
[0110] The multi-scale correlation pyramid is constructed by performing a pooling operation on the correlation volume, specifically, average pooling on the last two dimensions of the volume, resulting in multiple volumes at different resolutions. In this way, both large-scale and small-scale displacement information can be captured simultaneously. During the calculation, first, the initial optical flow field (f1,f2) is estimated, which is used for local search at multiple levels of the pyramid.
[0111] The search operation can be defined as:
[0112] LC={x′+dx∣dx∈Z2,∣∣dx∣∣1≤r} (5)
[0113] In the equation, x′=(u+f1(u),v+f2(v)), this local grid will be sampled in the correlation volume according to the optical flow estimation value, resulting in a new feature map for the next step of optical flow estimation. This full-pixel correlation method significantly enhances the ability of the RAFT model to capture large displacements and complex motion scenes, especially for complex flow field prediction tasks in particle image velocimetry.
[0114] Step 4, GRU Recurrent Update Module
[0115] GRU Recurrent Update Module (Conv-GRU Update Block): This module uses a convolutional GRU to iteratively update the optical flow estimation, recursively optimizing the correlation and features to gradually approximate the true optical flow field. Each iteration can further improve the accuracy of optical flow prediction based on the estimation of the previous round.
[0116] Specifically, the optical flow prediction is optimized step by step through a gated recurrent unit (GRU). This module uses a convolutional GRU (ConvGRU) and a separable convolutional GRU (SepConvGRU) to recursively update the features, adjusting the flow field prediction step by step. In each recursion, the GRU effectively retains important historical information through the update gate and the reset gate based on the input correlation features, the current optical flow estimate, and the hidden state, while combining the current features to calculate the flow increment (delta_flow) and iteratively update the optical flow.
[0117] In addition, the action feature encoder encodes the correlation features and the initial optical flow into action features, further improving the effect of recursive updating. The SepConvGRU captures more fine-grained feature information through horizontal and vertical separable convolutions, making the optical flow estimation more accurate. Through this recursive optimization, the GRU recursive update module can refine the flow field prediction frame by frame in complex particle image velocimetry tasks, thereby gradually improving the overall accuracy of the model.
[0118] Step 5, upsampling module
[0119] In the upsampling process of the Res2RAFT model, the convex combination upsampling method is used, which is an important step in improving the resolution of optical flow.
[0120] Specifically, the model first generates low-resolution optical flow prediction results through the recursive update module. To restore these low-resolution flow fields to high-resolution flow fields with the same resolution as the original image, the convex upsampling method interpolates the low-resolution optical flow after convolution to generate higher-resolution prediction results.
[0121] In specific implementation, the convex upsampling constructs multiple masks for interpolation, and then performs convolution operations on the local regions of each flow field. These masks are normalized by the softmax function to ensure that the contribution of each pixel in the interpolation process is reasonably allocated. Finally, the low-resolution flow field after multiple convolutions is combined with the generated masks to complete the upsampling process of the optical flow. This not only preserves the details of the original optical flow, but also effectively improves the spatial resolution of the optical flow prediction, making the prediction results more accurate. The advantage of this method is that it can accurately upsample the optical flow field from low resolution to high resolution without significantly increasing the computational cost, making it suitable for complex fluid field prediction.
[0122] Step 6, output displacement field
[0123] After processing through multiple layers of modules, the predicted displacement field can be obtained, and a displacement field image can be output. For example, Figure 6As shown, the left of the figure shows the real displacement field in the particle image, where the arrows represent the direction of flow and the color represents the size of displacement. The darker the color, the greater the displacement. These real values are used to measure the accuracy of the model's prediction and are the benchmark for our evaluation of the model's performance. The middle of the figure shows the flow field predicted by the Res2RAFT model, also represented by arrows indicating the direction of flow and color indicating the size of displacement. By comparing with the real flow field on the left, we can intuitively observe the performance of the model at different Reynolds numbers. The right of the figure shows the error between the model's predicted flow field and the real flow field, using the average endpoint error (EPE) as the evaluation standard. The darker the color, the greater the error. Through these error maps, we can clearly see where the model's prediction accuracy is low, which provides a reference for further improving the model.
[0124] Experimental verification:
[0125] As shown in Table 2, the model performance is tested using multiple datasets such as Cylinder, DNS_turbulence, and SQG, and the results show that this method has high accuracy and robustness in complex flow fields, especially in Cylinder and DNS_turbulence flow fields, with an accuracy improvement of more than 20%.
[0126] Table 2 Comparison of predicted values by different methods
[0127]
[0128] It can be clearly seen that Res2RAFT significantly outperforms traditional methods in most flow fields, especially in complex flow fields such as Cylinder and DNS_Turbulence, while there are some fluctuations in error in the JHTDB Channel flow field. Overall, Res2RAFT has a clear advantage in extracting multi-scale features of complex flow fields.
[0129] Analysis of experimental results:
[0130] Overall, Res2RAFT, as an improved deep learning model, not only performs stably in simple flow fields, but also performs well in complex flow fields (such as DNS_Turbulence and SQG), such as Figure 7 , Figure 8 As shown, it has significantly improved performance compared to RAFT and other traditional methods, deep learning models, etc. In the most challenging flow field, Res2RAFT provides a powerful combination of multi-scale feature extraction and global attention mechanism, enabling it to more accurately capture complex nonlinear features in the flow field. Although there is still room for improvement in the JHTDB Channel flow field, the performance of Res2RAFT in most flow fields is sufficient to demonstrate its strong ability in the optical flow prediction task.
[0131] Experimental conclusion:
[0132] The Res2RAFT model is significantly better than the traditional method and the existing deep learning model on multiple complex flow field data sets (including Cylinder, DNS_turbulence and SQG), especially in the aspect of precision improvement, and shows more than 20% advantage. By introducing multi-scale feature extraction and global attention mechanism, the Res2RAFT model has significantly enhanced the ability to extract global and local information in complex flow fields, and has higher robustness and accuracy in processing complex nonlinear flow (such as vortex and turbulence). The Res2RAFT model of the application provides a more accurate and efficient solution for optical flow estimation in the PIV task through innovative design in feature extraction and feature selection.
[0133] The above is only part of the specific embodiments of the application, but the protection scope of the application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the application, which should be covered within the protection scope of the application.
Claims
1. A multi-scale Res2RAFT particle image velocimetry method based on global attention mechanism optimization, characterized in that, The method comprises the following steps: S1, obtaining a PIV data set and data preprocessing: Each data sample in the PIV data set comprises particle images of two consecutive frames and the real velocity of the particle corresponding to the particle images; S2, constructing a particle image velocimetry model, wherein the particle image velocimetry model is based on a RAFT optical flow network, and the RAFT optical flow network comprises a feature extraction network, a context network, a full correlation layer and a GRU recursive update module; The feature extraction network and the context network of the RAFT optical flow network are replaced by a Res2Encoder network to obtain the particle image velocimetry model. The Res2Encoder network comprises, in sequence, a convolutional layer one, an RA-Bottle2neck module, an RB-Bottle2neck module, an RA-Bottle2neck module, an RB-Bottle2neck module, an RA-Bottle2neck module, an RB-Bottle2neck module, and a convolutional layer two. The convolutional layer one uses a 7x7 convolutional layer to generate an initial feature map of 64 channels. The convolutional layer two uses a 1x1 convolution to generate an output feature map of 128 channels. The RA-Bottle2neck module is a bottleneck structure with residual learning based on the Bottle2neck structure, used for down-sampling to maintain multi-scale information, and a multi-scale feature processing method is introduced in each residual block to map the input features of the Res2Encoder module into subsets, as the following formula, and each subset independently performs convolution operation to extract features at several scales to increase the diversity of the feature space: ; For the first subset , no convolution operation is performed and the original input is directly reserved as the output ; for the second subset , a standard 3x3 convolution operation is performed to obtain the output ; and for the third subset , the convolution operation not only depends on the input of this group , but also combines the output of the previous group , i.e. , so as to gradually fuse the information from different scales layer by layer; the feature extraction mode of gradual accumulation and recursion continues until the processing of the last subset is completed. The RB-Bottle2neck module is a bottleneck structure based on a Bottle2neck structure with a global attention mechanism, and is used to maintain resolution. The global attention mechanism GAM is added after the last layer of the RA-Bottle2neck module. The global attention mechanism GAM in the RB-Bottle2neck module comprises a channel attention and a spatial attention module, and utilizes the interaction of the channel and the spatial dimension to enable the particle image velocimetry model to focus on the particle flow information in the two consecutive frames. S3, constructing a loss function, and training the particle image velocimetry model by using the PIV data set in step S1 in a supervised learning manner.
3. The multi-scale Res2RAFT PIV method based on global attention mechanism optimization according to claim 1, characterized in that, 2. The multi-scale Res2RAFT particle image velocimetry method based on a global attention mechanism according to claim 1, wherein 4. The multi-scale Res2RAFT PIV method based on global attention mechanism optimization according to claim 3, characterized in that, In step S1, the data preprocessing specifically comprises uniformly processing the particle image size in the PIV data set, setting the length and width size to 256x256, and then performing regularization processing on the image.
5. The multi-scale Res2RAFT PIV method based on global attention mechanism optimization according to claim 1, characterized in that, In step S2, s=4, the stride and baseWidth parameters of the RA-Bottle2neck module are set to 2 and 26, respectively. ; wherein n is the number of samples in the PIV dataset, is the optical flow prediction value obtained by inputting the particle images of two consecutive frames of samples in the PIV dataset into the particle image velocimetry model, is the real optical flow value reflecting the real velocity of the particles in the samples in the PIV dataset. are weights for each prediction, giving later predictions more weight to improve the accuracy of the final result. In step S2, the stride and baseWidth parameters of the RB-Bottle2neck module are set to 1 and 26, respectively. In step S3, the loss function formula is as follows:
6. The multi-scale Res2RAFT particle image velocimetry method based on a global attention mechanism according to claim 1, wherein The full-pair correlation layer is used for full-pair correlation calculation on the features extracted by the feature extraction network, and a four-layer correlation pyramid is obtained; The full-pair correlation layer establishes the correlation between pixels by calculating the similarity between each pixel in two frames of images; a 4D correlation volume is constructed according to the following formula to perform the similarity calculation, and the 4D correlation volume stores the inner product of the feature vectors between all pixel pairs: ; In the formula, is the correlation volume, and is the feature map extracted by the feature encoder for two frames of images; wherein, represents the weight coefficient in the i, j spatial dimension and the k, l channel dimension of the convolution kernel; h is the input of the convolution calculation; The four-layer correlation pyramid is constructed by performing a pooling operation on the pair correlation volume.
7. The multi-scale Res2RAFT PIV method based on global attention mechanism optimization according to claim 1, characterized in that, The GRU recursive update module uses a convolutional GRU and a separable convolutional GRU to recursively update the features and gradually adjust the flow field prediction; in each recursion, the GRU retains important historical information through the update gate and the reset gate according to the input correlation features, the current optical flow estimation and the hidden state, and combines the current features to calculate the optical flow increment and iteratively update the optical flow; The separable convolutional GRU captures more fine-grained feature information through separable convolution in the horizontal and vertical directions.
Citation Information
Patent Citations
Robust particle image velocity measurement method based on optical flow convolutional neural network
CN117333516A
Ocean particle image velocity measurement method based on unsupervised learning and attention mechanism
CN118628759A