River surface flow velocity measurement method, system and product based on video image
Through the river surface flow velocity measurement method based on video images, combined with the RAFT optical flow model and a multi-layer perceptron framework, the accuracy and accuracy of flow velocity measurement in large river environments are solved, achieving higher precision flow velocity prediction and equipment cost reduction.
Patent Information
- Application Number
- CN202510418415.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-11
AI Technical Summary
The existing non-contact river surface flow velocity measurement methods have accuracy and accuracy problems in large river environments, especially in complex environments where the accuracy of flow velocity prediction in remote areas is reduced.
Using a river surface flow velocity measurement method based on video images, the perspective matrix is calculated through the unmanned ship control points, combined with the RAFT optical flow model and the multi-layer perceptron (MLP) framework, the pixel-level velocity is obtained using the optical flow field, and the river real-level velocity is learned through the multi-layer perceptron, and the combined loss function is designed to enhance the physical consistency of the model.
The flow velocity measurement accuracy in large river scenarios is improved, especially in complex and large-scale river environments, which effectively solves the problem of decreasing flow velocity prediction accuracy in remote areas, while avoiding interference to the water environment, reducing operational difficulty and equipment costs.
Smart Images

Figure CN120294356A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of non-contact measurement of river surface velocity, and relates to a method, system and product for measuring river surface velocity, and particularly relates to a method, system and product for measuring large river surface velocity based on video images. Background Art
[0002] At present, the measurement methods of river velocity generally include two types: contact type and non-contact type. The contact type methods mainly rely on velocity sensors, such as acoustic Doppler current profilers (ADCP). Such devices can provide relatively accurate velocity data. However, since they need to directly contact the water surface and have a certain impact on the water environment, their applications are usually limited. In addition, traditional contact type methods also have disadvantages such as high equipment cost, complex operation, and susceptibility to environmental factors interference. In recent years, non-contact velocity measurement methods have received extensive attention, especially the methods for measuring river surface velocity through video images, including LSPTV (Large Scale Particle Tracking Velocimetry), LSPIV (Large Scale Particle Image Velocimetry), STIV (Particle Tracking Image Velocimetry) technology and optical flow method, etc. Although these non-contact velocity measurement methods have overcome the limitations of the contact type methods to a certain extent, especially having advantages in avoiding direct contact and reducing environmental impact. However, in a large-scale river environment, they still have many limitations, especially problems of accuracy, precision and adaptability to environmental factors. Summary of the Invention
[0003] In order to overcome the problem of large measurement errors of river surface velocity of large rivers through video images. The present invention provides a method, system and product for measuring large river surface velocity based on video images.
[0004] The technical solution adopted by the method of the present invention is: a method for measuring river surface velocity based on video images, comprising the following steps:
[0005] Step 1: Select N unmanned ship control points, and calculate the perspective matrix through the geographical coordinate data of the unmanned ship and the pixel coordinates of the unmanned ship in the video image; wherein, N is a preset value;
[0006] Step 2: Obtain the optical flow field of the video image and obtain the pixel-level velocity of the point to be measured for velocity;
[0007] Convert the geographical coordinates of the point to be measured for velocity into pixel coordinates through the perspective matrix; generate the pixel-level velocity of each pixel of the video image, that is, the optical flow field; obtain the pixel-level velocity of the point to be measured for velocity from the optical flow field by using the pixel coordinates of the point to be measured for velocity.
[0008] Step 3: Use a multi-layer perceptron to obtain the river real-level velocity.
[0009] Preferably, in step 1, no less than 4 control points are selected. These points should be evenly selected along the river section, and the pixel coordinates in the image must correspond to their corresponding geographic coordinates one by one. For each control point, its pixel position (x, y) in the image and its corresponding geographic location (longitude, latitude) are recorded. Given the input pixel coordinates (x pixel ,y pixel ) and the output geographic coordinates (x geo ,y geo ), use cv2.findHomography to calculate the perspective transformation matrix
[0010] As a preferred method, in step 2, the inverse perspective matrix of the perspective transformation matrix calculated by the unmanned ship control point is obtained; then the geographic coordinates (x geo ,y geo ) is converted into homogeneous coordinates and converted into a three-dimensional coordinate vector (x geo ,y geo ,1) T ; Multiply the three-dimensional coordinate vector by the inverse perspective matrix to obtain a new three-dimensional coordinate vector; then, divide the first two components of the new three-dimensional coordinate vector by the third component for normalization to obtain the coordinate value on the two-dimensional plane; finally, round the obtained coordinate value on the two-dimensional plane to an integer to obtain the corresponding pixel coordinate.
[0011] Preferably, in step 2, the RAFT optical flow model is used to generate the pixel-level speed of each pixel of the video image. First, a continuous multi-frame video image is selected as input, and the RAFT optical flow model extracts two adjacent frames from them for optical flow estimation; in each pair of adjacent frames, the RAFT optical flow model calculates the motion vector of each pixel through a deep neural network, which represents the horizontal and vertical displacement of the pixel between the two frames, that is, the u component and v component of the optical flow; this process obtains multiple optical flow fields by calculating frame by frame, and each optical flow field contains the speed information of the pixels between different frames; finally, all the optical flow field results are averaged, and the motion vector of each pixel between different frames is averaged to obtain the final optical flow field.
[0012] Preferably, in step 3, the multilayer perceptron comprises a numerical feature branch, a position encoding branch, a feature merging layer, a backbone network and an output layer;
[0013] The numerical feature branch, input feature vector x∈R n, apply a fully connected layer, batch normalization, ReLU activation function, and dropout processing to the input feature vector, and output the numerical feature features = Dropout(ReLU(LayerNorm(Linear(x, 24))), 0.2), where x = [x1, x2, x3] ∈ R 3 ; where R represents the real number field, n represents the input dimension, x1, x2, and x3 represent the pixel-level velocity, relative position, and distance to the river bank respectively, M represents the output dimension, and 0.2 represents the dropout rate;
[0014] For the position encoding processing branch, the input longitude lon, latitude lat, and pixel position (pix_x, pix_y) are converted into low-dimensional position embedding vectors lon_emb, lat_emb, pix_x_emb, and pix_y_emb through the Embedding layer; then, these low-dimensional position embedding vectors are concatenated to form a long vector x concat = Concat(lon_emb, lat_emb, pix_x_emb, pix_y_emb), and after processing through a fully connected layer, batch normalization, ReLU activation function, and dropout, output the position encoding feature pos = Dropout(ReLU(LayerNorm(Linear(x concat , 48))), 0.2), where 48 represents the output dimension and 0.2 represents the dropout rate;
[0015] For the feature merging layer, concatenate the numerical feature and the position encoding feature to form a new vector merged = Concat(features, pos); then, after passing through a fully connected layer, batch normalization, ReLU activation function, and dropout processing, output x = Dropout(ReLU(LayerNorm(Linear(merged, 96))), 0.2), where 96 represents the output dimension and 0.2 represents the dropout rate;
[0016] For the backbone network, process the data x output by the feature merging layer through two residual blocks x' = LayerNorm(Linear(ReLU(LayerNorm(Linear(identity = x, 96))), 96)), and finally, the residual block adds the original input identity = x to the output x' and outputs x” = x' + identity; where 96 represents the output dimension;
[0017] For the output layer, input the processed feature x” into the output layer to obtain the final prediction value In the formula, 48 and 1 represent the output dimensions, and 0.2 represents the dropout rate.
[0018] Preferably, in step 3, the multi-layer perceptron is a trained multi-layer perceptron; during training, the multi-layer perceptron is supervised and trained using the true river velocity measured by the ADCP carried by the unmanned ship; during training, the combined loss function final_loss is used to constrain the multi-layer perceptron.
[0019] final_loss = base_loss + physics_weight · physics_loss;
[0020] In the formula, base_loss is the basic loss function, physics_loss is the physical constraint loss function; physics_weight is the weight of the physical constraint loss.
[0021] The basic loss function base_loss is a combination of MSE loss, MAE loss, Smooth_L1 and Huber loss: base_loss = α · MSE + β · MAE + (1 - α - β) · (smooth_L1 + Huber) / 2; in the formula, α and β are the weights of MSE and MAE losses respectively.
[0022] The physical constraint loss function physics_loss is a combination of flow velocity distribution constraint, river bank constraint and velocity gradient constraint: physics_loss = profile_loss + bank_loss + 0.5 × gradient_loss;
[0023]
[0024] In the formula, profile_loss, bank_loss, and gradient_loss represent the flow velocity distribution constraint, river bank constraint, and velocity gradient constraint respectively; the parameters r and d i respectively represent the relative position [x geo , y geo of the velocity measurement point to be measured and the distance from the river bank. ε, N, respectively represent the predicted velocity, the maximum predicted velocity, the non-zero minimum constant, the total number of data points, and the i-th velocity value predicted by the model.
[0025] Preferably, in step 3, the multi-layer perceptron is a trained multi-layer perceptron; during training, first, the true river velocity measured by the ADCP carried by the unmanned boat is divided into a training set, a validation set, and a test set; secondly, the relative position features of each point to be measured for velocity are calculated, and the ratio of the distance from this point to the starting point of the cross-section to the total length of the cross-section is obtained to get the proportional value of the relative position; at the same time, the distance from this point to the river bank is calculated, and the specific method is to find the minimum distance between this point and the two end points of the cross-section; then, all features are normalized. Among them, the pixel velocity uses the standardization method, that is, subtracting the mean and dividing by the standard deviation, while the coordinate values (including longitude, latitude, and pixel coordinates) use range normalization, mapping the longitude and latitude to the integer range of [0, 1999], and the pixel coordinates are mapped to the integer range of [0, 2999]; finally, all processed features are integrated into the input format required by the model, including numerical features (standardized pixel velocity, relative position, and distance to the river bank) and position features (normalized longitude, latitude, and pixel coordinates); calculate the relative position features and the distance to the river bank, normalize all input features, use standardization for pixel velocity, and use range normalization for coordinates.
[0026] The technical solution adopted by the system of the present invention is: a river surface velocity measurement system based on video images, including:
[0027] One or more processors;
[0028] A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the river surface velocity measurement method based on video images described above.
[0029] The technical solution adopted by the method of the present invention is: a river surface velocity measurement product based on video images, including a computer program, when the computer program instructions run on a computer, enabling the computer to execute the river surface velocity measurement method based on video images described above.
[0030] The beneficial effects brought by the technical solution provided by the present invention are:
[0031] (1) By combining the optical flow model and the designed multi-layer perceptron (MLP) framework, the present invention can achieve video image surface velocity measurement in large river scenarios. The optical flow model provides accurate input for the MLP model through the estimation of pixel-level velocity, and the MLP improves the accuracy of velocity prediction by learning the mapping relationship between these pixel-level velocities and the true velocity. Especially in complex and large-scale river environments, it can effectively solve the problem of the decline in velocity prediction accuracy in the far-end area;
[0032] (2) The MLP model designed in the present invention particularly considers the encoding of location information, enabling the model to better understand spatial information. Especially through the embedding of longitude, latitude, and pixel coordinates, the learning ability of the model for geographical location and spatial distribution characteristics is enhanced. This design enables the multi-layer perceptron not only to perform speed prediction but also to more accurately capture location-related influencing factors, thereby improving the prediction ability of the model;
[0033] (3) The present invention designs a combined loss function and particularly designs a physical loss function to strengthen the physical consistency of the model. The basic loss part of the combined loss function ensures that the model can effectively fit the observed data. The physical loss introduces flow velocity distribution constraints, bank constraints, and velocity gradient constraints. These three physical constraints enable the model to follow the basic laws of river dynamics. The flow velocity distribution constraint ensures the reasonable distribution of cross-sectional flow velocity, the bank constraint guarantees the natural attenuation of near-shore flow velocity, and the velocity gradient constraint ensures the spatial continuity of the flow field. This design helps the model generate more reasonable and realistic prediction results;
[0034] (4) The present invention adopts a non-contact video image acquisition method, avoiding the interference of traditional contact methods on the water environment. At the same time, it does not require direct contact with the water surface and is applicable to more complex river environments. It also greatly simplifies the process of traditional flow velocity measurement, reduces the dependence on complex equipment, and lowers the operation difficulty and equipment cost. Description of the Drawings
[0035] The following uses examples and specific implementation manners to further illustrate the technical solution of the present invention. Additionally, some drawings are also used in the process of describing the technical solution. For those skilled in the art, without creative efforts, other drawings and the intention of the present invention can also be obtained based on these drawings.
[0036] Figure 1 is the flowchart of the method in the embodiment of the present invention;
[0037] Figure 2 is the structure diagram of the multi-layer perceptron model in the embodiment of the present invention;
[0038] Figure 3 is the schematic diagram of the comparison of surface flow velocities between the experimental RAFT and RAFT+MLP in the embodiment of the present invention. Detailed Description of the Invention
[0039] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0040] Please refer to Figure 1, This embodiment provides a method for measuring the surface velocity of a river based on video images, including the following steps:
[0041] Step 1: Select N unmanned boat control points, and calculate the perspective matrix through the unmanned boat geographic coordinate data and the pixel coordinates of the unmanned boat in the video image; where N is a preset value;
[0042] In one implementation, at least 4 control points are selected. These points should be evenly selected along the river cross-section, and the pixel coordinates in the image and their corresponding geographic coordinates must be in one-to-one correspondence. For each control point, record its pixel position (x, y) in the image and the corresponding geographic location (longitude, latitude). Use cv2.findHomography to calculate the perspective transformation matrix. This method can calculate a 3×3 perspective transformation matrix H. Specifically, given the input pixel coordinates (x pixel , y pixel ) and the output geographic coordinates (x geo , y geo ), the mathematical model of the perspective transformation is:
[0043]
[0044] In the formula, H is cv2.findHomography, is the converted geographic coordinate, and [x pixel , y pixel is the pixel coordinate in the image.
[0045] Step 2: Use the RAFT optical flow model to obtain the optical flow field of the video image and obtain the pixel-level velocity of the points to be measured;
[0046] Convert the geographic coordinates of the points to be measured into pixel coordinates through the perspective matrix; use the RAFT optical flow model to generate the pixel-level velocity of each pixel in the video image, that is, the optical flow field; use the pixel coordinates of the points to be measured to obtain the pixel-level velocity of the points to be measured from the optical flow field generated by the RAFT optical flow model;
[0047] In one implementation, the specific implementation of Step 2 includes the following sub-steps:
[0048] Step 2.1: Convert the geographic coordinates of the points to be measured into pixel coordinates through the perspective matrix:
[0049] First, it is necessary to obtain the inverse matrix of the 3×3 perspective transformation matrix that has been calculated through the control points. Then the geographic coordinates (x geo , y geo ) are converted into homogeneous coordinate form, and it is converted into homogeneous coordinates (x geo , y geo , 1) T. Multiply this three - dimensional coordinate vector by the inverse perspective matrix to obtain a new three - dimensional vector. Then, normalize the first two components of this new vector by dividing them by the third component respectively to obtain the coordinate values on the two - dimensional plane. Finally, round these coordinate values to integers to obtain the corresponding pixel coordinates.
[0050] Step 2.2: Use the RAFT optical flow model to generate the pixel - level velocity of each pixel in the video image, that is, the optical flow field:
[0051] First, select consecutive multiple frames of video images as input, and the model will extract two adjacent frames from them for optical flow estimation. In each pair of adjacent frames, the RAFT optical flow model calculates the motion vector of each pixel through a deep neural network, representing the horizontal and vertical displacements of the pixel between the two frames, that is, the u - component and v - component of the optical flow. This process obtains multiple optical flow fields by frame - by - frame calculation, and each optical flow field contains the velocity information of pixels between different frames. Finally, perform a mean operation on all the optical flow field results, and average the motion vectors of each pixel between different frames to obtain the final optical flow field.
[0052] Step 2.3: Obtain the pixel - level velocity of the speed - measuring point to be measured from the optical flow field generated by the RAFT optical flow model using the pixel coordinates of the speed - measuring point to be measured;
[0053] Step 3: Use a multi - layer perceptron to obtain the real - world - level velocity of the river.
[0054] In one implementation, see Figure 2 , the multi - layer perceptron includes a numerical feature branch, a position encoding branch, a merged feature layer, a backbone network (residual block layer), and an output layer. The specific features are as follows:
[0055] Numerical feature branch, the input feature vector \(x\in R\) n , where \(R\) represents the real number field and \(n\) represents the input dimension; apply a fully - connected layer (Linear), batch normalization (LayerNorm), ReLU activation function, and dropout processing to these input features:
[0056]
[0057] In the formula, \(x_1\), \(x_2\), \(x_3\) represent the pixel - level velocity, relative position, and distance to the river bank respectively, 24 represents the output dimension, and 0.2 represents the dropout rate;
[0058] Position encoding processing branch, the input longitude, latitude, and pixel position are converted into low - dimensional vectors through the Embedding layer:
[0059]
[0060] Then, these positional embedding vectors are concatenated to form a long vector, which is then processed through a fully connected layer, batch normalization, ReLU activation function, and dropout:
[0061]
[0062] where 48 represents the output dimension and 0.2 represents the dropout rate;
[0063] Feature merging layer, which concatenates numerical features and positional encoding features to form a new vector:
[0064] merged = Concat(features, pos);
[0065] Then, it goes through a fully connected layer, batch normalization, ReLU activation function, and dropout:
[0066] x = Dropout(ReLU(LayerNorm(Linear(merged, 96))), 0.2);
[0067] where 96 represents the output dimension and 0.2 represents the dropout rate;
[0068] Backbone network (residual block), which processes the data through two residual blocks. The structure of each residual block is as follows. Finally, the residual block adds the original input identity to the output x:
[0069]
[0070] where 96 represents the output dimension;
[0071] Output layer, which inputs the processed features into the output layer to obtain the final prediction value:
[0072]
[0073] where 48 and 1 represent the output dimensions and 0.2 represents the dropout rate;
[0074] In one embodiment, the multi-layer perceptron is a trained multi-layer perceptron. First, the true velocity of the river measured by the ADCP (Acoustic Doppler Current Profiler) carried by the unmanned boat is segmented into a training set, a validation set, and a test set. Secondly, the relative position feature of each speed measurement point to be measured is calculated. By solving the ratio of the distance from this point to the starting point of the cross-section to the total length of the cross-section, the proportional value of the relative position is obtained. At the same time, the distance from this point to the river bank is calculated. The specific method is to find the minimum distance between this point and the two endpoints of the cross-section. Then, all features are normalized. Among them, the pixel velocity uses the standardization method, that is, subtracting the mean and dividing by the standard deviation, while the coordinate values (including longitude, latitude, and pixel coordinates) use range normalization, mapping the longitude and latitude to the integer range of [0, 1999], and the pixel coordinates are mapped to the integer range of [0, 2999]. The parameters used in the normalization process will be saved for the same feature processing in the subsequent prediction process. Finally, all processed features are integrated into the input format required by the model, including numerical features (standardized pixel velocity, relative position, and distance to the river bank) and position features (normalized longitude, latitude, and pixel coordinates). Calculate the relative position feature and the distance to the river bank, normalize all input features, use standardization (mean variance) for pixel velocity, and use range normalization for coordinates.
[0075] During training, the multi-layer perceptron is supervised and trained using the true velocity of the river measured by the ADCP (Acoustic Doppler Current Profiler) carried by the unmanned boat; during training, the combined loss function final_loss is used to constrain the multi-layer perceptron;
[0076] final_loss = base_loss + physics_weight · physics_loss;
[0077] In the formula, base_loss is the basic loss function, physics_loss is the physical constraint loss function; physics_weight is the weight of the physical constraint loss;
[0078] The basic loss function base_loss is a combination of MSE loss, MAE loss, Smooth_L1, and Huber loss: base_loss = α · MSE + β · MAE + (1 - α - β) · (smooth_L1 + Huber) / 2; in the formula, α and β are the weights of the MSE and MAE losses respectively;
[0079] The physical constraint loss function physics_loss is a combination of the flow velocity distribution constraint, the riverbank constraint, and the velocity gradient constraint: physics_loss = profile_loss + bank_loss + 0.5×gradient_loss;
[0080]
[0081] In the formula, profile_loss, bank_loss, and gradient_loss respectively represent the flow velocity distribution constraint, the riverbank constraint, and the velocity gradient constraint; the parameters r and d i respectively represent the relative position [x geo , y geo of the velocity measurement point to be measured and the distance from the riverbank; ε, N, respectively represent the predicted velocity, the maximum predicted velocity, the non-zero minimum constant, the total number of data points, and the i-th velocity value predicted by the model.
[0082] This embodiment also provides a river surface flow velocity measurement system based on video images, including:
[0083] One or more processors;
[0084] A storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the river surface flow velocity measurement method based on video images.
[0085] This embodiment also provides a river surface flow velocity measurement product based on video images, including a computer program, which, when the computer program instructions run on a computer, cause the computer to execute the river surface flow velocity measurement method based on video images.
[0086] The following further elaborates on the present invention through specific experiments.
[0087] First, pixel-level prediction is performed based on the model pre-trained on the Things dataset with the RAFT architecture. Subsequently, three groups of cross-sectional data are used for the training and verification of the MLP model, and the remaining group of data is used for the final prediction. In terms of the hyperparameter configuration of the model, an optimized strategy is adopted in the experiment: the AdamW optimizer is used, the initial learning rate is set to 0.0004, and the batch size is set to 48. In terms of the loss function, a combined loss function is adopted in the experiment, where alpha = 0.6, beta = 0.2, and the physical constraint weight is 0.0015.
[0088] The experimental data used in this invention consists of high-resolution river surface video data and four sets of measured data of ADCP in transverse sections. Among them, the river video data is collected by a fixed camera on the right bank at an inclined angle, covering the entire river width completely, with a resolution of 2560×1440 pixels; the four sets of transverse section ADCP data obtained synchronously are continuously arranged along the vertical direction of the river at time intervals of 3-4 seconds, and the data recording parameters include elements such as time stamps, longitude and latitude coordinates, and surface flow velocity values. The two types of data are synchronously associated through the spatio-temporal reference to form a collaborative data set of water surface video observation and cross-section flow velocity measurement. Please see Figure 3 , which is a schematic diagram of the comparison of surface flow velocities between RAFT and RAFT+MLP. It can be seen that the data of ADCP and RAFT+MLP show highly consistent fluctuation characteristics in the overall trend. Especially in the far-end area, the flow velocity change curves of the two are highly consistent. In the proximal area, the prediction error is relatively large, mainly because the drastic change of the terrain leads to a more complex flow velocity distribution. In addition, the limitedness of the experimental data in this area makes the model restricted to a certain extent in terms of learning and generalization ability, thus further exacerbating the prediction error.
[0089] This experiment further verified the effectiveness of the MLP structure through ablation experiments. The experimental results show that the RMSE and MAE values of the pure RAFT model are 0.782 and 0.727 respectively, while the RAFT+MLP model proposed in this invention has achieved significant optimization effects under the same experimental conditions, and its RMSE and MAE values are significantly reduced to 0.242 and 0.186 respectively. Specifically, compared with the pure RAFT model, after adopting the MLP structure, the decline rate of RMSE reaches 69.0%, indicating the effective reduction effect of MLP on the prediction error; at the same time, the decline rate of MAE is further expanded to 74.3%, which not only reflects the significant suppression effect of MLP on extreme error points, but also shows its outstanding effect in overall error control and prediction accuracy improvement. It should be noted that the RAFT model can only provide pixel-level velocity and cannot be directly mapped to the actual flow velocity. In contrast, as a post-processing module, MLP successfully converts the output of RAFT into a flow velocity prediction value that is more in line with real physical laws. In the mapping from pixel-level velocity to actual velocity, MLP plays a key role, significantly reducing the prediction error and significantly improving the fitting degree.
[0090] The present invention proposes a framework combining an optical flow model and a designed multi-layer perceptron (MLP). First, an image perspective matrix is calculated using the geographical coordinates and pixel coordinates of the unmanned boat. Secondly, the optical flow field of consecutive video frames is calculated using the optical flow model to obtain the pixel-level result of the river surface velocity. Finally, a multi-layer perceptron model and a combined loss function designed according to the present invention are used to obtain the real-world velocity. The present invention uses the real velocity measured by the ADCP (Acoustic Doppler Current Profiler) carried by the unmanned boat to supervise the training of the multi-layer perceptron model, and uses the combined loss function to constrain the model, and can predict the real-world velocity by combining the river position features and numerical features. The present invention adopts a non-contact video acquisition method, which reduces the equipment cost and improves the measurement efficiency.
[0091] It should be understood that the above-described embodiments are part of the embodiments of the present invention, rather than all embodiments. In addition, the technical features in each embodiment or individual embodiment provided by the present invention can be combined with each other arbitrarily to form a feasible technical solution. This combination is not restricted by the order of steps and / or the structure composition mode, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be achieved, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0092] It should be understood that the above description of the preferred embodiments is relatively detailed and should not be considered as a limitation on the protection scope of the present invention. Under the inspiration of the present invention, those of ordinary skill in the art can also make substitutions or deformations without departing from the protection scope defined by the claims of the present invention, and all fall within the protection scope of the present invention. The protection scope claimed by the present invention shall be subject to the appended claims.
Claims
1. A method for measuring the surface velocity of a river based on video images, characterized in that, It includes the following steps: Step 1: Select N unmanned vessel control points, and calculate the perspective matrix through the geographical coordinate data of the unmanned vessel and the pixel coordinates of the unmanned vessel in the video image; where N is a preset value; Step 2: Obtain the optical flow field of the video image and obtain the pixel-level velocity of the point to be measured for speed; Convert the geographical coordinates of the point to be measured for speed into pixel coordinates through the perspective matrix; generate the pixel-level velocity of each pixel in the video image, that is, the optical flow field; obtain the pixel-level velocity of the point to be measured for speed from the optical flow field by using the pixel coordinates of the point to be measured for speed; Step 3: Use a multi-layer perceptron to obtain the real-world level velocity of the river.
2. The method for measuring the surface velocity of a river based on video images according to claim 1, characterized in that: In step 1, select no less than 4 control points. These points should be evenly selected along the river section, and the pixel coordinates in the image must correspond to their corresponding geographic coordinates one by one. For each control point, record its pixel position (x, y) in the image and its corresponding geographic location (longitude, latitude). Given the input pixel coordinates (x pixel ,y pixel ) and the output geographic coordinates (x geo ,y geo ), use cv2.findHomography to calculate the perspective transformation matrix 3. The method for measuring the surface velocity of a river based on video images according to claim 1, wherein: In step 2, obtain the inverse perspective matrix of the perspective transformation matrix that has been calculated through the control points of the unmanned ship; then the geographical coordinates (x geo , y geo ) are transformed into the form of homogeneous coordinates and converted into a three-dimensional coordinate vector (x geo , y geo , 1) T ; multiply the three-dimensional coordinate vector by the inverse perspective matrix to obtain a new three-dimensional coordinate vector; then, divide the first two components of the new three-dimensional coordinate vector by the third component for normalization to obtain the coordinate values on the two-dimensional plane; finally, round the coordinate values on the obtained two-dimensional plane to integers to obtain the corresponding pixel coordinates.
4. The method for measuring the surface velocity of a river based on video images according to claim 1, wherein: In Step 2, the RAFT optical flow model is used to generate the pixel-level velocity of each pixel in the video image. First, select a continuous multi-frame video image as the input, and the RAFT optical flow model will extract two adjacent frames from it for optical flow estimation; in each pair of adjacent frames, the RAFT optical flow model calculates the motion vector of each pixel through a deep neural network, representing the horizontal and vertical displacements of the pixel between the two frames, that is, the u component and v component of the optical flow; this process obtains multiple optical flow fields by frame-by-frame calculation, and each optical flow field contains the velocity information of the pixels between different frames; finally, perform mean processing on all the optical flow field results, and average the motion vectors of each pixel between different frames to obtain the final optical flow field.
5. The method for measuring the surface velocity of a river based on video images according to claim 1, characterized in that: In Step 3, the multi-layer perceptron includes a numerical feature branch, a position encoding branch, a feature merging layer, a backbone network, and an output layer; The numerical feature branch inputs The feature vector x ∈ R n , after applying a fully connected layer, batch normalization, ReLU activation function, and dropout processing to the input feature vector, the numerical feature features = Dropout(ReLU(LayerNorm(Linear(x, 24))), 0.2) is output, where x = [x1, x2, x3] ∈ R 3 ; where R represents the real number field, n represents the input dimension, x1, x2, and x3 respectively represent the pixel-level velocity, relative position, and distance to the riverbank, where 24 represents the output dimension, and 0.2 represents the dropout rate; For the position encoding processing branch, the input longitude lon, latitude lat, and pixel position (pix_x, pix_y) are converted into low-dimensional position embedding vectors lon_emb, lat_emb, pix_x_emb, and pix_y_emb through the Embedding layer. Then, these low-dimensional position embedding vectors are concatenated to form a long vector x concat = Concat(lon_emb, lat_emb, pix_x_emb, pix_y_emb), and after being processed by a fully connected layer, batch normalization, ReLU activation function, and dropout, the position encoding feature pos = Dropout(ReLU(LayerNorm(Linear(x concat , 48))), 0.2) is output, where 48 represents the output dimension and 0.2 represents the dropout rate; The feature merging layer concatenates the numerical feature and the position encoding feature together to form a new vector merged = Concat(features, pos); then, after passing through a fully connected layer, batch normalization, ReLU activation function, and dropout processing, it outputs x = Dropout(ReLU(LayerNorm(Linear(merged, 96))), 0.2), where 96 represents the output dimension and 0.2 represents the dropout rate; The backbone network processes the data x output by the feature merging layer through two residual blocks x' = LayerNorm(Linear(ReLU(LayerNorm(Linear(identity = x, 96))), 96)); finally, the residual block adds the original input identity = x to the output x' and then outputs x” = x' + identity; where 96 represents the output dimension; The output layer inputs the processed feature x” into the output layer to obtain the final predicted value In the formula, 48 and 1 represent the output dimensions, and 0.2 represents the dropout rate.
6. The method for measuring the surface velocity of a river based on video images according to any one of claims 1-5, characterized in that: In Step 3, the multi-layer perceptron is a trained multi-layer perceptron; during training, the multi-layer perceptron is supervised and trained by using the real velocity of the river measured by the ADCP carried by the unmanned vessel; a combined loss function final_loss is used to constrain the multi-layer perceptron during training; final_loss = base_loss + physics_weight·physics_loss; In the formula, base_loss is the basic loss function, physics_loss is the physical constraint loss function; physics_weight is the weight of the physical constraint loss; The base loss function base_loss is a combination of MSE loss, MAE loss, Smooth_L1, and Huber loss: base_loss = α·MSE + β·MAE + (1 - α - β)·(smooth_L1 + Huber) / 2; where α and β are the weights of the MSE and MAE losses respectively. The physical constraint loss function physics_loss is a combination of flow velocity distribution constraint, riverbank constraint, and velocity gradient constraint: physics_loss = profile_loss + bank_loss + 0.5×gradient_loss. Among them, profile_loss, bank_loss, and gradient_loss represent the velocity distribution constraint, the bank constraint, and the velocity gradient constraint, respectively; the parameters r and d i respectively represent the relative position of the velocity measurement point to be measured with respect to [[x geo , y geo and the distance from the river bank; ε, N, respectively represent the predicted velocity, the maximum predicted velocity, the non-zero minimum constant, the total number of data points, and the i-th velocity value predicted by the model.
7. The method for measuring the surface velocity of a river based on video images according to any one of claims 1-5, characterized in that: In step 3, the multi-layer perceptron is a trained multi-layer perceptron. During training, first, the true river velocity measured by the ADCP carried by the unmanned boat is segmented into a training set, a validation set, and a test set. Secondly, calculate the relative position features of each velocity measurement point. By solving the ratio of the distance from this point to the starting point of the cross-section to the total length of the cross-section, the proportional value of the relative position is obtained. At the same time, calculate the distance from this point to the riverbank. The specific method is to find the minimum distance between this point and the two endpoints of the cross-section. Then, normalize all features. Among them, the pixel velocity uses the standardization method, that is, subtracting the mean and dividing by the standard deviation, while the coordinate values (including longitude, latitude, and pixel coordinates) use range normalization, mapping the longitude and latitude to the integer range of [0, 1999], and the pixel coordinates are mapped to the integer range of [0, 2999]. Finally, integrate all processed features into the input format required by the model, including numerical features (standardized pixel velocity, relative position, and distance to the riverbank) and position features (normalized longitude, latitude, and pixel coordinates); calculate the relative position features and the distance to the riverbank, normalize all input features, use standardization for pixel velocity, and use range normalization for coordinates.
8. A river surface velocity measurement system based on video images, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method for measuring the surface velocity of a river based on video images as described in any one of claims 1 to 7.
9. A river surface velocity measurement product based on video images, comprising a computer program, characterized in that: When the computer program instructions run on a computer, cause the computer to execute the method for measuring the surface velocity of a river based on video images as described in any one of claims 1 to 7.
Citation Information
Cited By
Flow velocity distribution prediction method and device, storage medium and computer equipment
CN121353993A