A Deep Learning-Based Method for Validity Recognition and Velocity Measurement of Spatiotemporal Images of River Surface

By constructing a real spatiotemporal image dataset and network model with multiple scenarios, wide range, and high precision, the problem of accuracy in spatiotemporal image velocity measurement of river surfaces in complex scenarios is solved. It achieves accurate flow measurement and velocity correction without manual parameter tuning and experimental datasets, and is suitable for sites with little or no data.

CN119445370BActive Publication Date: 2025-10-28HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411474910.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-10-28
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

Existing deep learning-based spatiotemporal image-based methods for river surface velocity measurement have low accuracy in complex scenarios and lack effective means to identify and correct invalid values, resulting in large measurement errors.

Method used

We constructed a real spatiotemporal image dataset with multiple scenarios, wide measurement range, and high precision. We built a network model using EfficientNetV2 and an optimized residual module, performed STI classification and MOT regression, and augmented the data using data cleaning, flipping, and rotation methods. This enabled accurate flow measurement and velocity correction without the need for manual parameter tuning or experimental datasets.

Benefits of technology

It achieves high-precision flow velocity measurement in complex environments, reduces interference from invalid and unreliable data, improves the generalization ability of the model, is applicable to sites with no or little data, and reduces computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119445370B_ABST
    Figure CN119445370B_ABST
Patent Text Reader

Abstract

This invention discloses a deep learning-based method for identifying the validity of spatiotemporal images of river surfaces and measuring velocity. During model training, natural river videos are collected, and spatiotemporal images (STIs) of different scenes are synthesized and selected. Based on whether texture features can represent the actual flow velocity, these images are categorized into three classes: valid, unreliable, and invalid, and a classification network model is trained accordingly. Valid STIs are semi-automatically labeled with the main texture direction (MOT), and data augmentation is performed using flipping, rotation, and color dithering to construct a MOT regression dataset. A regression network model is then trained, combining grouped convolution and convolutional attention modules to optimize the residual module. During velocity measurement, the classification network model identifies valid STIs, and the regression network model detects MOTs and converts them into flow velocities. The unreliable, invalid, and blind zone velocities are interpolated and corrected using the cross-sectional velocity distribution law to obtain the cross-sectional velocity field. This invention enables accurate flow measurement and velocity correction after training, without the need for manual parameter tuning or adding experimental river datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image-based flow measurement technology, and in particular to a method for recognizing the validity of spatiotemporal images of river surfaces and measuring velocity based on deep learning. Background Technology

[0002] Timely and accurate acquisition of hydrological information such as flow velocity and flow rate is crucial for mitigating flood damage. However, during peak flood season, rivers flow rapidly, carry high sediment loads, and contain abundant floating debris, making measuring equipment susceptible to damage. Furthermore, traditional contact measurement techniques are difficult to implement effectively due to the complex turbulent characteristics of natural rivers. Therefore, in recent years, non-contact fluid motion vector measurement technology has emerged and is widely used due to its efficiency and safety. This technology records water surface video using a camera, analyzes the motion vectors of tracers, and converts them to a global coordinate system to measure flow velocity.

[0003] Currently, common image-based flow velocity measurement methods include large-scale particle tracking velocimetry, large-scale particle image velocimetry, and spatiotemporal image velocimetry (STIV). Among them, STIV directly estimates one-dimensional time-averaged flow velocity by detecting the principal directions of texture (MOTs) in the spatiotemporal image (STI), offering advantages such as high spatial resolution and strong real-time performance. MOT detection is the core step of STIV, and its detection methods can be divided into spatial domain methods and frequency domain methods. Spatial domain methods, such as the gradient tensor method, are sensitive to image region segmentation, while frequency domain methods, such as the Fast Fourier Transform (FFT-STIV), although advantageous in terms of accuracy and efficiency, are more sensitive to noise, especially in complex natural environments, where existing filtering techniques struggle to completely eliminate noise effects.

[0004] In recent years, with the rapid development of deep learning, researchers have attempted to apply it to MOT detection in STIVs to improve detection accuracy and robustness. However, although they perform well under ideal conditions, their accuracy is low in complex scenarios, requiring the addition of real-world datasets for enhanced training. These methods are primarily based on manually constructed STI datasets and classification models. Furthermore, the angles output by these classification models are discrete integer values, making it difficult to capture the continuity of the true flow velocity.

[0005] Validity identification is a key step in the STIV process, but current methods have not discussed it in depth, so there is a lack of effective means to identify and correct invalid values ​​in actual measurements. Summary of the Invention

[0006] Purpose of the invention: This invention provides a method for recognizing the validity of spatiotemporal images of river surfaces and measuring velocity based on deep learning. After training, it can achieve accurate flow measurement and velocity correction without manual parameter tuning or adding experimental river datasets.

[0007] Technical solution: The present invention provides a method for identifying the validity of spatiotemporal images of river surfaces and measuring velocity based on deep learning, comprising the following steps:

[0008] Step 1: Collect videos of natural rivers at different cross-sections, water levels, flow velocities, weather conditions, lighting, and topography. Lay out velocity measurement lines to synthesize a basic spatiotemporal image (STI) dataset. Based on whether the main texture of the STI can represent the one-dimensional time-averaged flow velocity, the basic spatiotemporal image dataset is divided into three categories: valid, unreliable, and invalid. Randomly jitter the STI with varying brightness, contrast, saturation, and hue to obtain a STI classification dataset. After semi-automatic MOT annotation of the valid STI classes, the data is cleaned according to whether it contains vertical texture and then expanded as follows: STIs without vertical texture are rotated 180 times in 1° increments, followed by maximum size cropping to maximize the retention of valid information; STIs with vertical texture are horizontally flipped, vertically flipped, or horizontally and vertically flipped. Randomly jitter the expanded valid STI classes with varying brightness, contrast, saturation, and hue to obtain a MOT regression dataset.

[0009] Step 2: Build an STI classification network model based on EfficientNetV2, input the STI classification dataset into the model for training, and save the trained model; build a MOT regression network model based on the optimized residual module, input the MOT regression dataset into the model for training, and save the trained model.

[0010] Step 3: Real-time monitoring and acquisition of river surface video through camera system installed at hydrological station, deployment of velocity measurement lines to synthesize spatiotemporal images in real time, input of the original spatiotemporal images into the trained STI classification network model to obtain three types of spatiotemporal images: valid, unreliable, and invalid. Input of the valid spatiotemporal images into the trained MOT regression network model to obtain the corresponding MOT.

[0011] Step 4: Convert the effective class of MOT into the corresponding location velocity value, and combine it with the cross-sectional velocity distribution law to interpolate and correct the unreliable, invalid and blind zone velocities to obtain the cross-sectional velocity field.

[0012] Furthermore, in step 1, the valid class is STI under conditions without strong environmental interference. In this case, MOT can correctly represent the actual flow velocity, which includes normal, glare, noise, occlusion, turbulence, static shadow, standing wave, blur, and ripple scenes. The unreliable class is STI under conditions that strongly interfere with the natural water flow or have textures that are confused with the water flow. MOT is unreliable and it is difficult to correctly represent the one-dimensional time-averaged flow velocity. It includes floodplain flow, dynamic shadow, backflow, and strong wind ripple scenes. The invalid class is STI under conditions that completely interfere with the camera's observation of the river surface. It includes strong light, dark noise, complete occlusion, strong shadow, heavy fog, and heavy rain scenes.

[0013] Furthermore, in step 1, the MOT semi-automatic annotation first uses FFT-STIV for preliminary annotation, and then verifies it with manually labeled values. When the interpolation between the two is less than 1°, the FFT-STIV labeled value is used as the dataset MOT; otherwise, it is identified as STI in complex scenarios, and the manually labeled value is used as the dataset MOT.

[0014] Furthermore, in step 1, to maximize the size after STI rotation and cropping to retain more original information, the relationship between the cropping size l, the original image size h, and the rotation angle θ is as follows:

[0015] l=h·sin45° / sin(135°-θ%90°).

[0016] Furthermore, in step 2, the STI classification network model uses EfficientNetV2 as its backbone and sets the output nodes to 3 using a fully connected FC layer, corresponding to three classifications.

[0017] Furthermore, in step 3, the MOT regression network model is based on the residual network architecture. The convolution operation in the residual module is replaced with grouped convolution, and a convolutional attention module CBAM is added before the residual connection. The output node is restricted to 1 using the fully connected FC layer, and the final activation function is modified to Sigmoid to adapt to regression convergence.

[0018] Furthermore, in step 4, the cross-sectional velocity distribution law is as follows: v(x) represents the surface velocity at a starting point distance x, x0 is the starting point distance from the nearest waterline, L is the width of the water surface in the cross-sectional direction, and k is an undetermined coefficient. The effective velocity set {v(x1), v(x2), ..., v(x0)} in the nearshore to mid-channel region is used. i )|x l ≤x i ≤x mid} and the corresponding starting point distance set {x1,x2,…,x} i |x l ≤x i ≤x mid Given the data, calculate [x] using the least squares method. l ,x mid The k value for the region, similarly, is the effective velocity set from the mid-channel region to the far shore {v(x mid ),v(x mid+1 ),…,v(x i )|x mid ≤x i ≤x r} and the corresponding starting point distance set {x mid ,x mid+1 ,…,x i |x mid ≤x i≤x r Given the data, calculate [x] using the least squares method. mid ,x r The k value of the region, where x mid x is the distance from the starting point of the middle channel. l x is the distance from the starting point of the nearshore waterline. r The distance from the starting point of the far shoreline;

[0019]

[0020] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: (1) By classifying the scenes and effectiveness of spatiotemporal images of river water surfaces, the boundary conditions for spatiotemporal image velocity measurement to be applicable to complex and changing environments in the field are clarified, providing a basis for controlling gross errors in flow velocity measurement under extreme conditions in online video flow measurement systems; (2) The lightweight EfficientNet is used. The V2 network model performs three-class validity identification of spatiotemporal images. On the one hand, it avoids the overall detection accuracy deterioration caused by the interference of invalid and unreliable samples when training the MOT regression model. On the other hand, it avoids the computational overhead of subsequent MOT detection of invalid and unreliable STIs. (3) By performing semi-automatic MOT annotation on valid STIs and comprehensively using data cleaning, flipping, rotation and color dithering to expand the data, a real spatiotemporal image MOT regression dataset of multiple scenes, wide range and high precision of river water surface is constructed for model training. Compared with the existing artificially synthesized spatiotemporal image dataset, it can truly reflect the complexity and diversity of river water surface scenes under different weather, lighting and water flow conditions, so that the model has a stronger generalization ability and can be directly applied to sites with no data or little data, without the need for model transfer learning or reinforcement training. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the method flow of the present invention.

[0022] Figure 2 A flowchart for constructing the dataset of this invention.

[0023] Figure 3 This is a structural diagram of the MOT regression network model of the present invention.

[0024] Figure 4 This is a flowchart illustrating the implementation of the method of the present invention. Detailed Implementation

[0025] like Figure 1 As shown, a method for recognizing the validity of spatiotemporal images of river surfaces and measuring velocity based on deep learning includes the following steps:

[0026] Step 1: Construct a real spatiotemporal image dataset with multiple scenes, wide measurement range, and high precision, such as Figure 2 As shown, natural river videos were collected at different cross-sections, with varying water levels, flow velocities, weather conditions, illumination, and topography. Velocity measurement lines were then laid out to synthesize the basic spatiotemporal image STI dataset.

[0027] Based on whether the main texture of the STI (Surface Indicator Texture) can represent the one-dimensional time-averaged flow velocity of water, the basic spatiotemporal image dataset is divided into three categories: valid, unreliable, and invalid. The valid category includes STI under conditions without strong environmental interference, where the Momentum of Observation (MOT) can correctly represent the actual flow velocity. This category includes scenes such as normal, glare, noise, occlusion, turbulence, static shadows, standing waves, blur, and ripples. The unreliable category includes STI under conditions that strongly interfere with the natural water flow or contain textures that are indistinguishable from the water flow. In these cases, the MOT is unreliable and cannot accurately represent the one-dimensional time-averaged flow velocity. This category includes scenes such as floodplain flows, dynamic shadows, backflows, and strong wind ripples. The invalid category includes STI under conditions that completely interfere with the camera's observation of the river surface, including scenes such as strong light, dark noise, complete occlusion, strong shadows, heavy fog, and heavy rain. Random brightness, contrast, and saturation enhancements are then applied to these invalid STI datasets to obtain the STI classification dataset.

[0028] For valid STI classes, a semi-automatic MOT annotation was performed. This involved initial annotation using FFT-STIV, followed by verification with manually labeled values. When the interpolation was less than 1°, the FFT-STIV labeled value was used as the dataset MOT; otherwise, it was considered a complex scene STI, and the manually labeled value was used as the dataset MOT. After data cleaning based on whether it contained vertical texture, the dataset was expanded as follows: STIs without vertical texture were rotated 180 times in 1° increments. To maximize the size of the rotated and cropped STIs and retain more original information, the relationship between the adaptive cropping size l, the original image size h, and the rotation angle θ is shown in Equation 1. STIs containing vertical texture were then subjected to horizontal flipping, vertical flipping, and horizontal and vertical flipping. The expanded valid STI classes were then subjected to random dithering of brightness, contrast, saturation, and hue to obtain the MOT regression dataset.

[0029] l=h·sin45° / sin(135°-θ%90°)

[0030] Step 2: Build an STI classification network model based on EfficientNetV2 and an MOT regression network model based on optimized residual modules.

[0031] The STI classification network model uses EfficientNetV2 as its backbone, employing fully connected (FC) layers to set the output nodes to 3, corresponding to three classifications. The network structure is shown in Table 1: Conv2d is the convolution operation, SiLU is the activation function, which, compared to ReLU, has the characteristics of being unbounded above, bounded below, smooth, and non-monotonic, providing better performance in deep models. Pooling is the pooling layer.

[0032] Table 1 STI Classification Network Structure

[0033]

[0034] EfficientNet V2 employs an inverse residual structure (MBConv) and a channel attention mechanism (SE). The basic process of MBConv is as follows: first, the low-dimensional input is increased in dimensionality through dilated convolution, then processed by depthwise convolution, followed by dimensionality reduction through pointwise convolution, and finally the residuals are summed. This structure effectively improves gradient propagation and significantly reduces computational cost by utilizing depthwise separable convolutions. Although depthwise convolutions are less efficient in the early stages, they are highly effective in later stages. Therefore, in the early stages of the model, ordinary 3x3 convolutions are used to replace the 1x1 dilated convolutions and 3x3 depthwise convolutions in MBConv, thus forming the Fused-MBConv structure. This design improves model efficiency while ensuring performance preservation.

[0035] The SE module explicitly establishes the dependencies between feature channels. It weights features by learning the importance of each channel, thereby highlighting important features and suppressing unimportant ones. In operation, the SE module first compresses the spatial features of each channel into a single global feature using global average pooling, thus fusing feature information. Next, the module utilizes two fully connected layers to predict the importance of each channel. The first fully connected layer reduces the number of channels and dimensionality by scaling hyperparameters, thereby reducing computational cost; the second fully connected layer restores the original dimensionality, ensuring that the importance of a channel perfectly matches the number of its feature maps. Finally, the calculated channel importance is used to weight the original feature maps, multiplying each channel's feature map by its corresponding weight to generate a new weighted feature map. This process effectively increases the network's focus on important features while suppressing unimportant ones, thereby enhancing the model's representational power and performance.

[0036] like Figure 3As shown, a MOT regression network model based on optimized residual modules is constructed. To address the challenges of deep learning training and accurate MOT detection on large datasets, a MOT regression network model combining grouped convolutions and convolutional attention mechanisms is designed. Taking the first convolutional layer in the structure diagram as an example, the kernel size is 7×7, the number of output channels is 64, and the stride is 2 / 3. Repetition refers to the number of times the module is repeated. In the BottleNeck module, Groups=32 indicates the use of 32 groups of convolutions. The last layer is a fully connected layer (FC), whose output nodes are set to 1, and the result is restricted to the (0,1) interval by the Sigmoid activation function to achieve continuous angular regression prediction.

[0037] Grouped convolution breaks down convolution into multiple groups of dimensionality-reducing convolution operations, then performs dimensionality-up concatenation before connecting the outputs. This design retains the residual module's ability to avoid gradient vanishing / exploding while increasing the module's scalability and improving computational efficiency, making it particularly suitable for resource-constrained environments and large-scale datasets.

[0038] CBAM consists of two main sub-modules: the Channel Attention Module (CAM) and the Spatial Attention Module (SAM). CAM aims to enhance the feature representation of each channel. It computes the maximum and average eigenvalues ​​for each channel using global max pooling (MaxPool) and global average pooling (AvgPool), then learns the attention weights for each channel through a shared multilayer perceptron (SharedMLP) fully connected layer (FC). Finally, it applies a sigmoid activation function to generate CAM weights, which are then applied to each channel of the original feature map. The Spatial Attention Module focuses on the spatial location of the input data. It performs global max pooling (ChannelMaxPool) and global average pooling (ChannelAvgPool) along the channel dimension, and finally applies a sigmoid activation function to generate SAM weights, ultimately producing spatial attention weights, which are then applied to the spatial dimension of the original feature map. This mechanism, which simultaneously focuses on channel and spatial information, is particularly suitable for long-texture detection tasks, effectively improving the accuracy of feature extraction and the model's performance.

[0039] The STI classification network model was trained using the STI classification dataset, and the training parameters are shown in Table 2. The MOT regression network model was trained using the MOT regression dataset, and the training parameters are shown in Table 3. The datasets were standardized to improve learning ability; the standardization formula is as follows:

[0040]

[0041] Where input(c) represents the pixel values ​​of the original data channel c, output(c) is the standardized pixel value, N represents the number of images, and H×W is the image size. Early stopping is used to avoid overfitting when accuracy stops increasing, and a batch size reduction scheme is used to monitor memory usage and avoid GPU memory overflow. The trained model is then saved.

[0042] Table 2 Training parameters of the STI classification network model

[0043]

[0044] Table 3 Training parameters of the MOT regression network model

[0045]

[0046] Step 3: As Figure 4 As shown, a camera system installed at a hydrological station acquires real-time video of the river surface, and velocity measurement lines are laid out to obtain real-time spatiotemporal images. The raw spatiotemporal images are input into a stored STI classification network model to obtain three categories of spatiotemporal images: valid, unreliable, and invalid. The valid spatiotemporal images are then input into a stored MOT regression network model to obtain the valid location MOT.

[0047] Step 4: Convert the effective class of MOT to the corresponding positional velocity value according to Equation 5. In Equation 5, δ is the MOT value. Suppose that the tracer in the object plane moves a distance D along a certain velocity measurement line in time T, which is represented by moving d pixels in the image coordinate system within frame τ. v (pixel / s) represents the optical flow motion vector, and its positive or negative value reflects the direction of motion; there is only one scaling relationship between V and v, represented by the object image scale factor Δs (m / pixel) on the velocity measurement line, the value of which is obtained by calibration using a monocular visual plane measurement method with varying height homography.

[0048]

[0049] By combining the cross-sectional velocity distribution law with interpolation corrections for unreliable, invalid, and blind zone velocities, the cross-sectional velocity field is obtained. The cross-sectional velocity distribution law is shown in Equation 6:

[0050]

[0051] v(x) represents the surface velocity at a starting point distance x, x0 is the distance from the starting point closest to the waterline, L is the width of the water surface in the cross-sectional direction, and k is an undetermined coefficient. The effective velocity set from the nearshore to the middle channel region is {v(x1), v(x2), ..., v(x...}}. i )|x l ≤x i ≤x mid} and the corresponding starting point distance set {x1,x2,…,x}i |x l ≤x i ≤x mid Given the data, calculate [x] using the least squares method. l ,x mid The k-value of the region. Similarly, the effective velocity set from the mid-channel region to the far shore {v(x mid ),v(x mid+1 ),…,v(x i )|x mid ≤x i ≤x r} and the corresponding starting point distance set {x mid ,x mid+1 ,…,x i |x mid ≤x i ≤x r Given the data, calculate [x] using the least squares method. mid ,x r The k-value of the region. Where x mid x is the distance from the starting point of the middle channel. l x is the distance from the starting point of the nearshore waterline. r The distance from the starting point of the far shoreline.

Claims

1. A method for recognizing the validity of spatiotemporal images of river surfaces and measuring velocity based on deep learning, characterized in that, Includes the following steps: Step 1: Collect videos of natural rivers at different cross-sections, water levels, flow velocities, weather conditions, lighting conditions, and topography. Lay out velocity measurement lines to synthesize a basic spatiotemporal image STI dataset. Based on whether the main texture of the STI can represent the one-dimensional time-averaged flow velocity of the water, the basic spatiotemporal image dataset is divided into three categories: valid, unreliable, and invalid. After processing, an STI classification dataset is obtained. After performing semi-automatic MOT annotation on the valid STI classes, the data is cleaned according to whether it contains vertical textures, and then expanded to obtain MOT regression datasets. Step 2: Build an STI classification network model based on EfficientNetV2, input the STI classification dataset into the model for training, and save the trained model; build a MOT regression network model based on the optimized residual module, input the MOT regression dataset into the model for training, and save the trained model. Step 3: Real-time monitoring and acquisition of river surface video through camera system installed at hydrological station, deployment of velocity measurement lines to synthesize spatiotemporal images in real time, input of the original spatiotemporal images into the trained STI classification network model to obtain three types of spatiotemporal images: valid, unreliable, and invalid. Input of the valid spatiotemporal images into the trained MOT regression network model to obtain the corresponding MOT. Step 4: Convert the effective class of MOT into the corresponding location velocity value, and combine it with the cross-sectional velocity distribution law to interpolate and correct the unreliable, invalid and blind zone velocities to obtain the cross-sectional velocity field.

2. The method for identifying the validity of spatiotemporal images of river surfaces and measuring velocity based on deep learning as described in claim 1, characterized in that, In step 1, the valid class is STI under conditions without strong environmental interference. In this case, MOT can correctly represent the actual flow velocity. It includes normal, glare, noise, occlusion, turbulence, static shadow, standing wave, blur, and ripple scenes. The unreliable class is STI under conditions that strongly interfere with the natural water flow or have textures that are confused with the water flow. MOT is unreliable and it is difficult to correctly represent the one-dimensional time-averaged flow velocity. It includes floodplain flow, dynamic shadow, backflow, and strong wind ripple scenes. The invalid class is STI under conditions that completely interfere with the camera's observation of the river surface. It includes strong light, dark noise, complete occlusion, strong shadow, heavy fog, and heavy rain scenes.

3. The method for identifying the validity of spatiotemporal images of river surfaces and measuring velocity based on deep learning as described in claim 1, characterized in that, In step 1, the semi-automatic MOT annotation first uses FFT-STIV for preliminary annotation, and then verifies it with manually labeled values. When the interpolation between the two is less than 1°, the FFT-STIV labeled value is used as the dataset MOT; otherwise, it is identified as STI in complex scenarios, and the manually labeled value is used as the dataset MOT.

4. The method for identifying the validity of spatiotemporal images of river surfaces and measuring velocity based on deep learning as described in claim 1, characterized in that, In step 1, after performing semi-automatic MOT annotation on the effective STI classes, the data is cleaned according to whether it contains vertical texture, and then expanded as follows: STIs without vertical texture are rotated 180 times in 1° increments, and then cropped to the maximum size to maximize the retention of effective information; STIs with vertical texture are flipped horizontally, vertically, and both horizontally and vertically. The expanded effective STI classes are then randomly jittered with random brightness, contrast, saturation, and hue to obtain the MOT regression dataset.

5. The method for identifying the validity of spatiotemporal images of river surfaces and measuring velocity based on deep learning as described in claim 4, characterized in that, In step 1, to maximize the size after STI rotation and cropping to retain more original information, the relationship between the cropping size l, the original image size h, and the rotation angle θ is shown in the following formula: l=h·sin45° / sin(135°-θ%90°).

6. The method for identifying the validity of spatiotemporal images of river surfaces and measuring velocity based on deep learning as described in claim 1, characterized in that, In step 2, the STI classification network model uses EfficientNetV2 as the backbone and sets the output nodes to 3 using a fully connected FC layer, corresponding to the three categories.

7. The method for identifying the validity of spatiotemporal images of river surfaces and measuring velocity based on deep learning as described in claim 1, characterized in that, In step 3, the MOT regression network model is based on the residual network architecture. The convolution operation in the residual module is replaced with grouped convolution, and a convolutional attention module CBAM is added before the residual connection. The output node is restricted to 1 by using the fully connected FC layer, and the final activation function is modified to Sigmoid to adapt to regression convergence.

8. The method for identifying the validity of spatiotemporal images of river surfaces and measuring velocity based on deep learning as described in claim 1, characterized in that, In step 4, the cross-sectional velocity distribution law is shown in the following formula, where v(x) represents the surface velocity at a starting point distance x, x0 is the starting point distance from the nearest waterline, L is the width of the water surface in the cross-sectional direction, and k is an undetermined coefficient. The effective velocity set {v(x1), v(x2), ..., v(x0)} in the nearshore to mid-channel region is used. i )|x l ≤x i ≤x mid } and the corresponding starting point distance set {x1,x2,…,x} i |x l ≤x i ≤x mid Given the data, calculate [x] using the least squares method. l ,x mid The k value for the region, similarly, is the effective velocity set from the mid-channel region to the far shore {v(x mid ),v(x mid+1 ),…,v(x i )|x mid ≤x i ≤x r } and the corresponding starting point distance set {x mid ,x mid+1 ,…,x i |x mid ≤x i ≤x r Given the data, calculate [x] using the least squares method. mid ,x r The k value of the region, where x mid x is the distance from the starting point of the middle channel. l x is the distance from the starting point of the nearshore waterline. r The distance from the starting point of the far shoreline;