A lightweight and adaptive river surface flow velocity monitoring method and device
Patent Information
- Application Number
- CN202610871898.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明提供一种轻量化和自适应的河流表面流速监测方法及装置,用以解决如何在环境和设备状态动态变化的情况下实现精确的河流表面流速监测的技术问题
[0019]本发明提供的轻量化和自适应的河流表面流速监测方法及装置,根据边缘终端的设备状态参数和目标水域的环境参数确定注意力头的目标数量和特征目标维度,能够自适应边缘设备算力负载及外部环境的动态变化,从而在环境和设备状态动态变化的情况下,动态调整模型结构,避免了高负载时的处理卡顿或监测精度下降的问题;同时,对多帧河流表面图像进行局部特征提取并进行因果时序融合,充分考虑了水流多帧间的时间渐变特性,有效避免了单帧方案导致的预测误差大的问题;通过目标维度的截取与最终的回归处理,在资源受限的边缘场景中成功兼顾了流速监测的准确性与实时性。
Smart Images

Figure CN122821156A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a lightweight and adaptive method and apparatus for monitoring river surface flow velocity. Background Technology
[0002] River surface velocity is a core parameter in hydrological monitoring, crucial for applications such as flood warning and water resource allocation. Currently, deep learning-based image processing techniques are commonly used for velocity monitoring. However, traditional high-precision velocity monitoring models (such as ViT-Base) typically have a massive number of parameters, exceeding 80M, placing extremely high demands on computing power. When these models are applied to edge computing devices with limited computing power, they often suffer from excessive computational burden and long inference times (greater than 200ms / frame), failing to meet the real-time monitoring requirements of 25fps in edge scenarios.
[0003] To address the challenge of running high-precision models in real-time on edge computing devices, existing technologies typically employ lightweight neural network models or feature extraction algorithms based on single-frame images to reduce computational complexity. These existing technologies primarily aim to improve image processing speed with limited edge device computing power by significantly compressing the number of model parameters and reducing the dimensionality of feature extraction, thereby attempting to achieve river surface flow velocity monitoring in edge environments.
[0004] However, the aforementioned lightweight models typically employ fixed network structures, which cannot adapt to the dynamic changes in the computing load of edge devices and the external environment. This can easily lead to processing lag or a significant decrease in monitoring accuracy under high loads. Furthermore, schemes based on single-frame images ignore the temporal variation characteristics of water flow across multiple frames (such as flow velocity inertia), resulting in large prediction errors (RMSE > 0.2 m / s) and difficulty in accurately capturing dynamic processes such as flood rise and fall. Therefore, in order to balance the accuracy and real-time performance of flow velocity monitoring in resource-constrained edge scenarios, how to achieve accurate river surface flow velocity monitoring under dynamically changing environmental and equipment conditions has become an urgent problem to be solved in this field. Summary of the Invention
[0005] This invention provides a lightweight and adaptive method and apparatus for monitoring river surface velocity, which solves the technical problem of how to achieve accurate monitoring of river surface velocity under dynamic changes in environment and equipment conditions.
[0006] This invention provides a lightweight and adaptive method for monitoring river surface velocity, comprising: The number of targets and the dimension of feature targets in the attention head are determined based on the device status parameters of the edge terminal and the environmental parameters of the target water area. Local features are extracted from multiple frames of river surface images to obtain spatial local features, and global self-attention weighting is applied to the spatial local features using the target number of attention heads to obtain an attention feature sequence. The attention feature sequence is subjected to causal temporal fusion to obtain temporal fusion features, and the channel features of the temporal fusion features in the preceding dimension are extracted based on the feature target dimension to obtain the target features; The target features are subjected to regression processing to obtain the surface velocity of the river.
[0007] According to the present invention, a lightweight and adaptive method for monitoring river surface velocity extracts spatial local features from multiple frames of river surface images, including: Initial local features of the multi-frame river surface images are extracted using depthwise separable convolution; The initial local features are fused along the channel dimension using point convolution to obtain the spatial local features.
[0008] According to a lightweight and adaptive method for monitoring river surface velocity provided by the present invention, the spatial local features are subjected to global self-attention weighting to obtain an attention feature sequence, including: The spatial local features are projected into a query matrix, a key matrix, and a value matrix using independent point convolutions; Self-attention calculation is performed on the query matrix, the key matrix, and the value matrix to obtain the attention feature sequence.
[0009] According to a lightweight and adaptive method for monitoring river surface velocity provided by the present invention, causal temporal fusion is performed on the attention feature sequence to obtain temporal fusion features, including: The attention feature sequences are stacked in chronological order to form a temporal feature sequence. The temporal feature sequence is subjected to causal convolution; the scope of the convolution kernel is the current and past historical frames. The temporal fusion features are obtained by downsampling the feature space dimension during the causal convolution process.
[0010] According to the present invention, a lightweight and adaptive method for monitoring river surface flow velocity is provided, wherein the device status parameters include computing power utilization and remaining power, and the environmental parameters include water flow turbulence and light intensity.
[0011] According to the present invention, a lightweight and adaptive method for monitoring river surface velocity determines the number of targets and feature target dimensions of the attention head based on the device status parameters of the edge terminal and the environmental parameters of the target water area, including: The joint feature vector containing the device state parameters and the environmental parameters is input into the gating network to obtain the attention head number adjustment factor and feature dimension adjustment factor output by the gating network. The number of targets is determined based on the attention head number adjustment factor and the attention head number benchmark, and the feature target dimension is determined based on the feature dimension adjustment factor and the feature dimension benchmark.
[0012] According to a lightweight and adaptive method for monitoring river surface velocity provided by the present invention, the number of targets is determined based on the attention head number adjustment factor and the attention head number benchmark, and the feature target dimension is determined based on the feature dimension adjustment factor and the feature dimension benchmark, including: The target number is obtained by multiplying the attention head number adjustment factor and the attention head number benchmark and then rounding down. The feature dimension target dimension is obtained by multiplying the feature dimension adjustment factor and the feature dimension benchmark and then rounding down.
[0013] According to a lightweight and adaptive method for monitoring river surface velocity provided by the present invention, after performing regression processing on the target features to obtain the river surface velocity, the method further includes: The dataset, including the multi-frame river surface images, the device status parameters, and the river surface flow velocity, is uploaded to the cloud, and the updated model parameters are obtained by the teacher model in the cloud through knowledge distillation based on the dataset. Update the local student model using the updated model parameters.
[0014] According to a lightweight and adaptive method for monitoring river surface velocity provided by the present invention, the updated model parameters obtained by knowledge distillation of the teacher model in the cloud based on the dataset are acquired, including: Obtain the flow velocity probability distribution obtained by the teacher model predicting the dataset; Calculate the relative entropy loss between the flow velocity probability distribution of the teacher model and the flow velocity probability distribution of the student model; Receive the updated model parameters obtained by the cloud after performing network backpropagation based on the relative entropy loss.
[0015] The present invention also provides a lightweight and adaptive river surface velocity monitoring device, comprising: The edge adaptive gating module is used to determine the number of targets and feature target dimensions of the attention head based on the device status parameters of the edge terminal and the environmental parameters of the target water area. The image feature extraction module is used to extract local features from multiple frames of river surface images to obtain spatial local features, and to perform global self-attention weighting on the spatial local features using the target number of attention heads to obtain an attention feature sequence. The causal-temporal fusion module is used to perform causal-temporal fusion on the attention feature sequence to obtain temporal fusion features, and to extract the channel features of the temporal fusion features in the preceding dimension based on the feature target dimension to obtain the target features; The feature regression processing module is used to perform regression processing on the target features to obtain the river surface flow velocity.
[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the lightweight and adaptive river surface velocity monitoring method as described above.
[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the lightweight and adaptive river surface velocity monitoring method as described above.
[0018] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the lightweight and adaptive river surface velocity monitoring method as described above.
[0019] The lightweight and adaptive river surface velocity monitoring method and device provided by this invention determines the number of targets and feature target dimensions of the attention head based on the device status parameters of the edge terminal and the environmental parameters of the target water area. It can adapt to the dynamic changes in the computing power load of the edge device and the external environment, thereby dynamically adjusting the model structure under dynamic changes in the environment and device status, avoiding the problems of processing lag or decreased monitoring accuracy under high load. At the same time, it performs local feature extraction and causal temporal fusion on multiple frames of river surface images, fully considering the temporal gradation characteristics of water flow between multiple frames, effectively avoiding the problem of large prediction errors caused by single-frame solutions. Through target dimension truncation and final regression processing, it successfully balances the accuracy and real-time performance of velocity monitoring in resource-constrained edge scenarios. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0021] Figure 1 This is a flowchart illustrating the lightweight and adaptive river surface velocity monitoring method provided by the present invention.
[0022] Figure 2 This is a system architecture diagram provided by the present invention.
[0023] Figure 3 This is a schematic diagram of the MobileViT-Block structure provided by the present invention.
[0024] Figure 4 This is a schematic diagram of causal convolutional temporal fusion provided by the present invention.
[0025] Figure 5 This is the edge adaptive gating flowchart provided by the present invention.
[0026] Figure 6 This is a comparison curve of RMSE in field testing provided by the present invention.
[0027] Figure 7 This is a schematic diagram of the lightweight and adaptive river surface velocity monitoring device provided by the present invention.
[0028] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0030] The following is combined with Figures 1 to 8 The present invention describes a lightweight and adaptive method and apparatus for monitoring river surface velocity.
[0031] Figure 1 This is a flowchart illustrating the lightweight and adaptive river surface velocity monitoring method provided by the present invention, as shown below. Figure 1 As shown, the method includes, but is not limited to, steps S1, S2, S3, and S4. This method can be applied to edge terminals, such as resource-constrained edge computing devices like smart cameras with ISP (Image Signal Processing) capabilities, embedded gateways, and mobile edge nodes.
[0032] Step S1: Determine the number of targets and feature target dimensions of the attention head based on the device status parameters of the edge terminal and the environmental parameters of the target water area.
[0033] Device status parameters can include data reflecting the internal physical constraints of the device, such as computing power utilization, remaining power, CPU (Central Processing Unit) temperature, GPU (Graphics Processing Unit) load, or memory usage.
[0034] The target water body can be a body of water such as a river, lake, reservoir, or artificial channel that requires flow velocity monitoring. Environmental parameters can include data reflecting the external environment of the target water body, such as water flow turbulence, light intensity, wind speed, rainfall, or water surface visibility.
[0035] Attention heads can be computational units, matrix mapping branches, or feature association modules used for independent representation of subspaces in multi-head self-attention mechanisms. The number of targets can be 4, 5, or 8. The dimension of the feature targets can be the number of channels of the hidden layer feature tensor, the vector length, or the width of the feature matrix, such as 58-dimensional or 64-dimensional.
[0036] The number of targets and feature dimensions of the attention head can be determined through methods such as rule mapping, querying preset lookup tables, using lightweight neural network prediction, or machine learning regression models. For example, when the edge camera detects that its CPU load is too high and the external ambient light is sufficient, the network calculates and outputs a smaller number of targets (such as reducing the number of heads from 8 to 4) and a slightly reduced feature dimension (such as reducing it from 64 dimensions to 58 dimensions).
[0037] Step S1 can provide a data-driven adaptive adjustment basis, enabling the network structure to flexibly scale according to the hardware and software environment, taking into account both computational overhead and feature representation.
[0038] Step S2: Local features are extracted from multiple frames of river surface images to obtain spatial local features. Global self-attention weighting is then applied to the spatial local features using the target number of attention heads to obtain an attention feature sequence.
[0039] Multi-frame river surface images can be a sequence of images containing water surface ripples that are acquired in real time by an edge terminal at a specific frame rate (e.g., 25fps, frames per second) and extracted sequentially in time. For example, extracting a sequence of images {I1, I2, ..., I...} that are consecutive T=5 frames. t}, where T is the length of the timing window, covering the inertial period of the water flow, and each 5 frames constitute a processing unit.
[0040] It should be noted that before feature extraction, image resolution compression is typically performed using methods such as bilinear interpolation and nearest neighbor interpolation. For example, a 1080P image is downsampled to 320×240 in real time to preserve flow velocity-related texture features such as waves and floating object trajectories, and the normalization formula I't=(I t -μ) / σ maps pixel values to [-1, 1] for processing, where I t Let t be the original image at time t, μ = 127.5 be the image mean, σ = 127.5 be the image standard deviation, and I't be the normalized image.
[0041] Spatial local features can be extracted into a three-dimensional feature tensor or a set of feature maps containing dimensions and number of channels. Attention feature sequences are temporal data arrays composed of weighted feature tensors from multiple frames.
[0042] Local features can be extracted through forward propagation of convolutional neural networks, and weighting can be achieved by assigning different weights using dot product self-attention, additive self-attention, or multi-head parallel matrix multiplication. For example, local feature maps containing wave information can be extracted from a sequence of 5 consecutive river images through multi-layer convolution, and then input into a network module that has been dynamically set to 4 attention heads for feature interaction calculation, outputting a weighted feature sequence.
[0043] Step S2 can capture local details while fully integrating the key spatial connections across the entire image.
[0044] Step S3: Perform causal temporal fusion on the attention feature sequence to obtain temporal fusion features, and extract the channel features in the preceding dimension of the temporal fusion features based on the feature target dimension to obtain the target features.
[0045] Temporal fusion features can be multi-dimensional features that combine historical information and the current state. The preceding dimension can refer to the continuous region in the feature tensor channel index starting from 0 and extending to the target feature dimension minus 1. The target feature is the core feature tensor that has been truncated and retained.
[0046] One-dimensional causal convolution can be used to fuse temporal information from each frame, and then tensor slicing can be used to extract channels of a specified dimension. For example, the attention feature sequences corresponding to 5 frames can be fused into features with inertial representation by using convolution kernels that only cover the current and past convolution operations on the time axis. Then, according to the aforementioned 58 dimensions, array slicing can be used to directly extract the first 58 channels (index 0 to 57) as the target features.
[0047] Step S3 can precisely reduce redundant feature channels while preserving the gradual change characteristics of water flow over time, thereby reducing the computational overhead of subsequent networks.
[0048] Step S4: Perform regression processing on the target features to obtain the surface velocity of the river.
[0049] The final river surface velocity can be calculated using methods such as fully connected layer mapping computation and linear regressor prediction. For example, the aforementioned 58-dimensional target features can be flattened and input into a fully connected network containing two layers of neurons (e.g., a 32-dimensional hidden layer), directly outputting the predicted river surface velocity. The formula is as follows: v_pred=ReLU(W3·ReLU(W4·F_temporal[:D']+b4)+b3); Where v_pred is the predicted river surface velocity; F_temporal is the target feature vector truncated to the preceding dimension channel according to the target dimension D'; W3 is the second-layer regression head weight matrix, with a size of 64×32; W4 is the first-layer regression head weight matrix, with a size of 32×1; b3 is the second-layer bias term, with a size of 1; b4 is the first-layer bias term, with a size of 32; ReLU is the activation function, which ensures that the final output velocity value is non-negative.
[0050] Combination Figure 2 The figure clearly shows the overall macroscopic system environment in which this monitoring method runs in the edge computing module and is linked with the cloud.
[0051] As described above, this invention determines the number of targets and feature dimensions of the attention head based on the device status parameters of the edge terminal and the environmental parameters of the target water area. It can adapt to the dynamic changes in the computing power load of the edge device and the external environment, thereby dynamically adjusting the model structure under dynamic changes in the environment and device status, avoiding the problems of processing lag or decreased monitoring accuracy under high load. At the same time, it performs local feature extraction and causal temporal fusion on multiple frames of river surface images, fully considering the temporal gradation characteristics of water flow between multiple frames, effectively avoiding the problem of large prediction errors caused by single-frame schemes. Through target dimension truncation and final regression processing, this method successfully balances the accuracy and real-time performance of flow velocity monitoring in resource-constrained edge scenarios.
[0052] In one embodiment, step S2, which involves extracting local features from multiple frames of river surface images to obtain spatial local features, may further include: Initial local features of multiple frames of river surface images are extracted using depthwise separable convolution. Spatial local features are obtained by fusing the initial local features along the channel dimension through point convolution.
[0053] MobileViT-Block can be used as the core feature extractor.
[0054] Initial local features can be the initial spatial texture of the image, wave boundary distribution, gray-level gradient matrix, or basic shape contour. Depthwise separable convolution can be a lightweight convolutional network layer that performs two stages of operation: channel-wise convolution and point-wise convolution.
[0055] Initial local features can be extracted by using spatial filtering with depthwise convolutional (DWConv) layers with kernel sizes of 3×3, 5×5, or 7×7. For example, a sliding scan can be performed on a normalized input image using a 3×3 depthwise convolutional layer with a kernel size of 3, a stride of 1, and padding of 1, processing each color channel independently to extract initial spatial local feature maps such as wavy edges. This step can significantly reduce the large number of parameters and multiply-accumulate operations required by standard convolution, while effectively extracting the spatial local geometric information of each independent channel.
[0056] Channel-dimensional fusion can involve linearly weighting spatial data from multiple independent channels, cross-channel information interaction, or recombining and compressing the depth dimension of feature maps. Point convolution can be a standard convolutional layer with a strictly 1×1 kernel or a multilayer perceptron mapping.
[0057] Channel-dimensional fusion can be achieved through methods such as cross-channel pixel-level multiplication and addition operations using point convolutional layers with a kernel size of 1×1 (PWConv). For example, point convolutions with a kernel size of 1×1 combine, linearly map, and compress the independent channel information output by depthwise convolutions, ultimately outputting a spatial local feature F_local with a size of 160×120×32 (width and height are 160 and 120 respectively, with 32 channels). This step can compensate for the deficiency of fragmented information between depthwise convolution channels, achieving effective mapping and semantic reorganization of channel dimensions.
[0058] The formulas for the above extraction and fusion process are as follows: F_local=PWConv(DWConv(I't)); Where I't is the preprocessed input normalized image, DWConv is the depthwise convolution operation for spatial extraction, PWConv is the pointwise convolution operation for channel fusion, and F_local is the spatial local feature tensor generated in the final output.
[0059] Combination Figure 3 The local feature extraction section clearly demonstrates the basic architectural position of depthwise separable convolution at the front end of feature extraction, namely, as the basic module for extracting image texture.
[0060] This invention introduces depthwise separable convolution to extract initial local features, and then uses point convolution to fuse channel dimensions. This breaks down the traditional standard 3D convolution into two independent steps: spatial extraction and channel fusion, directly reducing the computational complexity and parameter redundancy in the spatial feature extraction process. At the same time, it ensures the effective interaction and recombination of multi-channel texture information, thereby generating high-quality and low-latency spatial local features for subsequent calculations.
[0061] In one embodiment, step S2, which involves performing global self-attention weighting on the spatial local features to obtain an attention feature sequence, may further include: Independent point convolutions are used to project local spatial features into query matrices, key matrices, and value matrices, respectively. Self-attention calculation is performed on the query matrix, key matrix, and value matrix to obtain the attention feature sequence.
[0062] Projection can be a linear transformation, affine transformation, or dimensionality reshaping mapping of features. The query matrix (Q), key matrix (K), and value matrix (V) are three sets of feature matrices used in the Transformer architecture to compute the attention relevance distribution and generate a weighted output.
[0063] Projection can be achieved by using three parallel 1×1 pointwise convolutional kernels to perform matrix multiplication mapping and dimensionality compression on local spatial features. To reduce computational complexity, the number of channels is first compressed from 32 to 16 using pointwise convolutions, preserving key information while reducing the amount of subsequent matrix operations. For example, using three independent 1×1 pointwise convolutional layers with different weight parameters, the original 160×120×32 local spatial features are compressed and linearly mapped into three feature maps, each 160×120×16, which serve as the query matrix, key matrix, and value matrix for subsequent calculations. This step reduces the channel dimension of the input features and generates a triplet feature space suitable for attention-based comparison operations. The relevant mapping formula is as follows: Q = PWConv_Q(F_local); K = PWConv_K(F_local); V = PWConv_V(F_local); Wherein, PWConv_Q, PWConv_K, and PWConv_V are independent point convolution operations used for dimensionality reduction projection, F_local is the spatial local feature of the input, and Q, K, and V are the query matrix, key matrix, and value matrix generated by the output.
[0064] Self-attention computation can be achieved by calculating similarity scores using queries and keys, followed by a weighted sum of the values. This can be done through methods such as dot product attention, cosine similarity attention, or multi-head attention. The query matrix is multiplied by the transpose of the key matrix, normalized using an activation function, and then multiplied by the value matrix. For example, by calculating the dot product of Q and K, the correlation score between pixel regions of local features is obtained. This score matrix is then weighted and superimposed on the V matrix, ultimately outputting a 160×120×16 integrated attention feature sequence. This step can uncover long-range semantic dependencies between different wave regions spanning long distances in a water surface image using the global receptive field. The formula for this process is: F_attn=Softmax((QK T ) / √d_k)V; Where F_attn is the output attention feature sequence; d_k=16 is the channel dimension parameter of the key matrix, and the square root is used to scale the dot product result to prevent gradient vanishing during backpropagation; Softmax is the row normalization exponential function used to convert the score into weights that sum to 1; K T Let be the mathematical transpose of the key matrix.
[0065] The above process corresponds to Figure 3 The global relationship modeling part.
[0066] This invention utilizes independent point convolutions to project spatial local features in a dimensionality reduction manner, significantly reducing the number of feature channels and thus effectively reducing the high computational burden of subsequent large-scale matrix multiplications. Based on the generated query, key, and value matrices, self-attention calculation is performed, which can break the limitation of traditional convolutions that only focus on the local receptive field of the neighborhood. This endows the model with the ability to perceive the overall flow state of the river surface globally, so that the obtained attention feature sequence contains both local ripple details and a deep fusion of the global water flow situation.
[0067] In one embodiment, step S3, which involves performing causal-temporal fusion on the attention feature sequence to obtain temporal fusion features, may further include: The attention feature sequences are stacked in chronological order to form a temporal feature sequence. Perform causal convolution on the temporal feature sequence; the scope of the convolution kernel of the causal convolution is the current and past historical frames; During causal convolution, feature space dimensions are downsampled to obtain temporal fusion features.
[0068] Channel stacking can be achieved by merging tensors along the time dimension, batch dimension, or channel dimension using the framework's underlying concatenation function. The temporal feature sequence is a feature tensor matrix that aggregates visual information from multiple consecutive time steps.
[0069] Stacking can be achieved by merging independent features from multiple frames along the time axis using functions in deep learning frameworks. For example, the attention-weighted features F_attn corresponding to the images at T=5 frames can be stacked sequentially. 1 F_attn 2 , ..., F_attn T The features are physically stacked in time along the time dimension according to their chronological order, forming a feature sequence with an overall size of 5×160×120×16 (representing the number of time steps × height × width × number of channels, respectively). This step can organize the discrete single-frame features into a continuous spatiotemporal data structure that can be processed by time-series operators.
[0070] Causal convolution can be a special one-dimensional or multi-dimensional convolution operation that ensures the output response at time step t depends only on time t and the inputs before time t. The convolution kernel's scope is the region of the original input image that the kernel can cover and extract as it slides. History frames are video image frames that have been acquired and extracted before the current processing time.
[0071] Causal convolution can be implemented by fusing temporal features, such as one-dimensional convolution with asymmetric padding or causal dilated convolution, using strict computational graph control to utilize only current and past frame information. For example, a one-dimensional causal convolution (CausalConv1D) with a kernel size of 3 can be used to slide along the time dimension across the stacked sequence of the above 5 frames. This ensures that the output value at time t is calculated only from the pixels of the current frame (t), the previous frame (t-1), and the two frames before that (t-2), thus physically preventing the leakage of future frames. This step strictly prevents information from future video frames from participating in the current state calculation in advance, closely aligning with the objective physical reality of real-time online flow rate monitoring, where only historical and current images can be obtained.
[0072] The feature space dimension is the length and width resolution of the feature tensor. Downsampling can be achieved by setting a convolutional kernel with a stride greater than 1, max pooling, or average pooling.
[0073] Downsampling can be achieved by directly setting a stride greater than 1 when configuring the causal convolution parameters, or by immediately following it with a pooling layer. For example, while the causal convolution operates along the time dimension, a downsampling operation with a stride of 2 halves the spatial dimensions (from 160×120 to 80×60) and expands the number of output channels to 64, resulting in a more compact feature F_temporal with a final output size of 5×80×60×64. This step can further compress the spatial size of the feature matrix while extracting and fusing temporally varying features, reducing the node pressure on subsequent fully connected layers.
[0074] The formula for the above time-series fusion process is: F_temporal=CausalConv1D([F_attn 1 F_attn 2 , ..., F_attn T ]); Where F_temporal represents the final temporal fusion feature; CausalConv1D is the one-dimensional causal convolution operation for processing, and the actual parameters can be kernel size=3, dilation=1, padding=1, stride=2; [F_attn 1 F_attn 2 , ..., F_attn T T-frame attention feature sequences are spliced and stacked along the time axis.
[0075] Combination Figure 3 Causal-temporal fusion part and Figure 4 This clearly demonstrates how convolutional computation absolutely guarantees temporal causality.
[0076] This invention fully preserves the temporal dependence and physical inertia of the dynamic evolution of water flow by stacking features in real time sequence and processing them using causal convolution with strictly limited kernel scope. At the same time, it completely eliminates the leakage and interference of future data information from the structure. In addition, the addition of stride control in the causal convolution scanning realizes the downsampling of the feature space dimension, which further compresses the spatial redundancy size of the feature matrix by a factor of two, significantly improving the computational efficiency and processing speed of generating temporal fusion features.
[0077] In one embodiment, the device status parameters of the present invention may include computing power utilization and remaining power, and the environmental parameters may include water flow turbulence and light intensity.
[0078] Computing power utilization can be the current CPU core utilization percentage of the edge terminal device, the GPU processing queue load rate, or the computing load of the dedicated NPU (Neural Processing Unit). Remaining power can be the percentage of charge currently stored in the edge terminal's battery relative to its total capacity, the SOC (State of Charge) returned by the intelligent battery management system, or the remaining available voltage. Water flow turbulence intensity is a quantitative indicator measuring the irregularity of surface flow, the intensity of eddies, or the complexity of ripple textures in a water body.
[0079] Device status and environmental parameters can be obtained by calling the underlying system interfaces of the edge device's operating system, reading hardware logs, or by collecting analog signals from external sensors and performing A / D conversion. For example, the computing power utilization rate C can be obtained by querying the native Linux system's `top` command, ensuring C ∈ [0, 1]; the remaining power B can be obtained by the ACPI power management command, ensuring B ∈ [0, 100]; the light intensity L can be obtained by collecting the light sensor built into the camera's motherboard, ensuring L ∈ [0, 1000] lux (lux is the unit of light intensity); and the water flow turbulence T can be obtained by calculating the pixel gradient variance of single or multiple frames of images captured by the camera, ensuring T ∈ [0, 1]. The purpose of this acquisition and parameter extraction calculation is to transform the abstract system operating environment and complex natural constraints into objective, standardized, and directly quantifiable input conditions that can be used for tensor operations.
[0080] The turbulence intensity T of water flow can be evaluated and obtained using the formula T=(1 / (H×W))Σ|∇I(i,j)|. Here, ∇ is the Sobel edge operator used to calculate the gradient, H and W are the height and width of the image being measured, respectively, I(i,j) is the brightness value of the pixel at the corresponding coordinate position (i,j), |∇I(i,j)| is the absolute value of the gradient at that point, and Σ represents the summation of the absolute values of the gradients of all pixels in the image. This yields a scientifically accurate turbulence intensity T that reflects the complexity and disorder of the water flow surface texture.
[0081] This invention defines the device status parameters as objective computing power utilization and remaining power, and the environmental parameters as measurable water flow turbulence and light intensity. It directly establishes the physical representation variables of the four core dimensions that affect the inference performance of edge nodes and the quality of input images. This enables the subsequent adaptive dynamic adjustment of the model network structure of the system to be based on accurate, comprehensive data that closely matches the actual pain points of edge computing in the field.
[0082] In one embodiment, step S1 may further include: The joint feature vector containing device state parameters and environmental parameters is input into the gating network to obtain the attention head number adjustment factor and feature dimension adjustment factor of the gating network output; The number of targets is determined based on the attention head number adjustment factor and the attention head number benchmark, and the feature target dimension is determined based on the feature dimension adjustment factor and the feature dimension benchmark.
[0083] The joint feature vector can be a tensor, a one-dimensional array, or a list of feature combinations, formed by concatenating and linking state and environment parameters from multiple different physical dimensions. The gating network can be an auxiliary neural network model containing a multilayer perceptron (MLP), fully connected layers, decision trees, or lightweight recurrent networks. The values of the attention head number adjustment factor and the feature dimension adjustment factor can be restricted to specific ranges, such as [0.5, 1], to dynamically scale the target network size proportionally.
[0084] Factors can be obtained by constructing a lightweight fully connected network structure with hidden layers and performing nonlinear forward propagation calculations, or by using support vector machines for mapping. For example, a lightweight gated network with two fully connected layers and a hidden layer dimension of 16 can be constructed at the edge terminal. A one-dimensional joint feature vector containing four specific values [C, B, L, T] arranged in sequence is input into the first layer of this network, processed sequentially by activation functions, and outputs an attention head number adjustment factor α∈[0.5, 1] and a feature dimension adjustment factor β∈[0.5, 1] at the end of the network. Utilizing the powerful fitting ability of lightweight neural networks, the complex nonlinear correspondence between the physical state of the parameters and the optimal execution hyperparameters of the model can be automatically learned. The relevant gated network mapping calculation formula is as follows: α=σ(W1·[C,B,L,T]+b1)×0.5+0.5; β=σ(W2·[C,B,L,T]+b2)×0.5+0.5; Where α and β are the attention head number adjustment factor and feature dimension adjustment factor generated in the final output, respectively; σ is the sigmoid activation function that smoothly maps the original linear output of the network to the interval [0, 1]; W1 and W2 are the corresponding weight matrices inside the gated network, both with a size of 4×16; [C, B, L, T] is the concatenated joint feature vector; b1 and b2 are the corresponding network bias terms, both with a size of 16. The weight matrices W1 and W2 of the fully connected layer of the gated network can be pre-trained by collecting one week of historical operating data from edge devices. The training objective function is set to minimize the overall error of flow rate prediction while meeting real-time constraints (e.g., inference time < 40ms).
[0085] The attention head count baseline and feature dimension baseline can be the default, fixed-size hyperparameters of the original Transformer model when it is not compressed or pruned. The number of targets or the feature dimension can be determined by directly multiplying the calculated adjustment factors with the initial hard-coded baseline in the model code, scaling the correlation, or querying a proximity threshold. For example, when the preset attention head count baseline H0 = 8 and the feature dimension baseline D0 = 64, the dynamic factors α and β obtained from the aforementioned forward computation of the network are read, and the baselines are scaled proportionally according to their values to ultimately determine the number of targets and the number of feature dimensions to be used in the current video frame. In actual edge system operation, to adapt to dynamic changes, α and β can be set to be re-collected and updated every 5 frames.
[0086] Combination Figure 5 The figure visually illustrates how the joint physical parameters are mapped step by step through two layers of fully connected gating networks and are ultimately reliably converted into an adjustment strategy for model structure adaptation.
[0087] This invention directly inputs a joint feature vector, composed of physical device status and external environmental parameters, into a lightweight gating network. This enables the network to intelligently and automatically learn the optimal mapping between external dynamic physical factors and the complexity of the model's inference structure, thereby accurately outputting the optimal adjustment factor under different conditions. By combining the system's preset baseline size to dynamically determine the final number and dimensions of targets, the invention ensures that the model can achieve the optimal balance between terminal computing power requirements and flow rate feature expression capabilities based entirely on real-time data when dealing with complex and ever-changing edge scenarios, thus realizing a real-time balance between monitoring accuracy and inference speed.
[0088] In one embodiment, determining the target number based on an attention head number adjustment factor and an attention head number benchmark, and determining the feature target dimension based on a feature dimension adjustment factor and a feature dimension benchmark, may further include: Multiply the attention head number adjustment factor and the attention head number baseline, then round down to obtain the target number; The feature dimension target dimension is obtained by multiplying the feature dimension adjustment factor and the feature dimension baseline and then rounding down.
[0089] The formula for calculating the target quantity is: H' = floor(α × H0); Where H' is the target number and H0 is the preset baseline for the number of attention heads.
[0090] In practical implementations within deep learning frameworks, the physical reduction of the model's inference structure can be achieved by using a masking mechanism to block redundant attention head outputs in network modules (i.e., disconnecting their computation graph branches or resetting their weights to zero). For example, when CPU utilization exceeds 80% and the device is under severe high load (e.g., C=0.9), the gating network automatically reduces the output α to between 0.5 and 0.6 (e.g., 0.6). Calculating 0.6 × 8 = 4.8 and rounding it down yields the target number H' = 4. At this point, the number of heads controlling the self-attention computation components is reduced from 8 to 4-5. This allows the device's single-frame inference time to be rapidly reduced from 35ms to 25ms. Although the flow rate error increases from 0.13m / s to 0.16m / s (an increase of 23%), the continuity of video stream processing is successfully ensured, and fatal stuttering is avoided.
[0091] The formula for calculating the feature target dimension is: D'=floor(β×D0); Where D' is the calculated feature target dimension, and D0 is the fixed feature dimension benchmark.
[0092] In the actual code implementation of the framework, tensor index slicing can be used on the generated feature vectors to extract and retain only the first D' dimensions of the tensor. For example, when the external ambient light is sufficient (e.g., L=800 lux, or below 200 lux), the β output of the gated network is kept at a high level of 0.9-1.0 (e.g., 0.9) to retain more effective visual features. In this case, 0.9×64=57.6 is calculated, and after floor rounding, the feature target dimension D'=57 is obtained. This appropriately reduces the amount of computation while ensuring inference accuracy, so that the RMSE (root mean square error) of the monitoring results is stably maintained at an extremely low 0.11m / s.
[0093] This invention multiplies continuous adjustment factors by a preset benchmark and then rounds down, converting the adjustment weights within the continuous value range into discrete integers that satisfy the hard constraints of the tensor dimension of the deep learning model. Combined with physical engineering operations that shield redundant attention head outputs in real time in the actual network structure and directly extract the pre-dimensional features, this ensures that the number of targets and the feature target dimensions obtained after dynamic adjustment are structurally valid in actual software and hardware inference operations, and maintains the continuous stability of the system under various combinations of equipment load pressure and external lighting conditions.
[0094] In one embodiment, after step S4, the method of the present invention may further include: The dataset, which includes multiple frames of river surface images, equipment status parameters, and river surface flow velocity, is uploaded to the cloud. The updated model parameters are obtained by the cloud-based teacher model through knowledge distillation based on the dataset. Update the local student model using updated model parameters.
[0095] The teacher model is a high-precision network model deployed and permanently residing on a cloud server, with an extremely large number of internal parameters and a very thorough pre-training process. For example, the cloud-based model ViT-Base has over 86M parameters and was pre-trained on millions of river video datasets. Knowledge distillation is an advanced training paradigm that allows a student model with very few network parameters to attempt to learn and fit the probability distribution, feature map-level response, or soft labels output by the large-parameter teacher model, thereby completing the transfer of prior knowledge. Updating the model parameters involves generating a completely new set of neural network weight files after iterative optimization in the cloud.
[0096] Lightweight IoT communication protocols such as MQTT (Message Queuing Telemetry Transport), HTTP POST requests, or WebSocket long connections can be used via wireless networks to proactively send locally cached data from the edge to the cloud server. Subsequently, an asynchronous distillation training script is launched in the cloud, and after training, the edge device receives the updated model parameters. For example, the edge computing module transmits the predicted flow rate result v_pred and corresponding slice data to the cloud in real time via the MQTT protocol with extremely low latency (<100ms) (it can be configured to extract one packet every 5 frames and upload it to the cloud), while simultaneously caching detailed data from the most recent hour locally on the edge for subsequent analysis. The cloud platform collects this type of real-world monitoring data (including image sequences, flow rate labels, and device status) uploaded from various edge devices over a long period, and then performs a knowledge distillation iteration task every two weeks. A powerful teacher model re-examines this data and guides the generation of updated model parameters suitable for the edge.
[0097] The student model is a lightweight parametric model deployed on edge terminals to perform daily low-power monitoring tasks. The model has only 12M parameters and an average inference time of only 34.2ms on the device, which can perfectly meet the 25fps requirement.
[0098] The newly generated parameter weight file can be pulled from the cloud to the edge local file system by calling the OTA (Over-The-Air) component, and the local student model can be updated by hot-loading and replacing the network weight variables in the original physical memory through the framework API. For example, after the cloud completes the bi-weekly distillation update calculation, the updated model parameters, which are 12MB in size, are directly pushed to the corresponding edge terminal device via OTA. The edge system seamlessly reloads the weights in memory, thereby realizing a complete data loop of "cloud training - edge inference".
[0099] This invention packages real-time image sequences, device status, and local flow rate results collected at the edge into a dataset and uploads it to the cloud. This allows the cloud-based, high-precision teacher model, with its unlimited computing power, to utilize this rich real-world edge data to perform knowledge distillation on the student model with a very small number of parameters, thereby producing updated model parameters with better generalization. Finally, these updated model parameters are used to update the local student model at the edge. This cloud-edge collaborative iterative mechanism allows the resource-constrained edge terminal to continuously absorb the macroscopic prior knowledge of the cloud-based large model, significantly improving the long-term accuracy and environmental resistance of the flow rate monitoring system without any additional cost of manually labeled data.
[0100] In one embodiment, obtaining the updated model parameters obtained by knowledge distillation of the dataset from the teacher model in the cloud may further include: Obtain the flow velocity probability distribution obtained by the teacher model predicting the dataset; Calculate the relative entropy loss between the flow velocity probability distribution of the teacher model and the flow velocity probability distribution of the student model; Receive updated model parameters obtained from the cloud after backpropagation of the network based on relative entropy loss.
[0101] The velocity probability distribution is a set vector of probabilities that the predicted target velocity falls within the range of the divided discrete intervals after the estimated continuous velocity physical value range is discretized at equal intervals according to a specific set step size.
[0102] For the collected real image and video streams, the continuous prediction range from 0 to the maximum possible flow velocity can be discretized into a series of intervals with a step size of 0.1 m / s, such as dividing it into a set of intervals v_i∈[0, 0.1), [0.1, 0.2), ... Then, the predicted probability distribution output by the executed teacher model on each of these divided intervals is obtained and denoted as P_teacher.
[0103] The relative entropy loss, also known as the KL divergence loss function, is calculated using the following formula: L_KD=KL(P_teacher||P_student)=ΣP_teacher(v_i)log(P_teacher(v_i) / P_student(v_i)); Where L_KD is the calculated relative entropy loss, KL represents the Kullback-Leibler divergence calculation function, P_teacher represents the velocity probability distribution output by the high-precision teacher model, P_student represents the velocity probability distribution output by the student model to be optimized, v_i is the specific velocity interval divided by the discretization interval, Σ represents the summation operation over all discretized velocity intervals, and log is the natural logarithm function.
[0104] Backpropagation in the network involves calculating the error gradient of the weights of each layer's parameters using the chain rule of calculus, based on the calculated loss function value, and then calling the optimizer to update and modify the student model weights in reverse.
[0105] The student model weight matrix in cloud memory can be iteratively updated using gradient backpropagation based on derivatives via a cloud-based deep learning framework (such as Adam or SGD optimizer). Then, the edge device initiates a communication request to download and apply the latest parameter update file. Driven by this cloud-edge closed-loop system, it exhibits exceptional iterative evolution performance, demonstrating a continuously improving monitoring accuracy during 30 days of continuous field testing.
[0106] Combination Figure 6 The diagram visually demonstrates the effectiveness of the invention. From... Figure 6 The experimental data clearly show that, because causal convolution effectively preserves the dynamic temporal dependence of water flow, the causal model reduces the RMSE index of flow velocity measurement by about 30% compared to the single-frame model based solely on images, with the error significantly decreasing from 0.20 m / s to 0.14 m / s. Furthermore, after the model of this invention is superimposed and continuously iterated using the cloud-based knowledge distillation mechanism, the edge model can continuously absorb deep knowledge from the cloud. In the following three months (or during long-term operation), its RMSE error index eventually decreased further from 0.15 m / s at the initial deployment stage to 0.12 m / s, and the accuracy was improved by about 20%. Its final monitoring accuracy is extremely close to the 0.10 m / s of the massive cloud-based teacher model.
[0107] This invention obtains the probability distribution of flow velocity predicted by the output of a large teacher model after discretizing the continuous flow velocity, and then rigorously calculates the KL divergence between it and the output distribution of a small-parameter student model according to the formula. This allows the student model to not only fit the correct target prediction value during cloud training, but also to deeply absorb and understand the evaluation knowledge and confidence rules of the teacher model for non-target intervals. Based on this loss feedback containing rich nondeterministic physical information, the network backpropagation is performed to update the parameters, which effectively accelerates the gradient convergence speed of the lightweight student model. This allows the overall error of its flow velocity prediction to gradually and stably approach that of the teacher network with massive parameters, significantly improving the upper limit of the dynamic flow velocity prediction accuracy of edge terminal devices without relying on large-scale manually labeled data.
[0108] The edge terminal of this invention uses a smart camera with ISP functionality, and the cloud server is configured with a GPU. The overall process includes: The camera captures 1080P video at 25fps, downsamples it to 320×240 in real time, and normalizes it according to I'=(I-127.5) / 127.5. Every 5 frames constitute a processing unit.
[0109] MobileViT-Block uses a 3×3 depthwise separable convolution kernel with a stride of 1 and padding of 1; and a 1×1 pointwise convolution kernel. Self-attention Q, K, and V generation employs independent 1×1 convolutions with 16 output channels. Causal convolutions use 1D convolutions in the temporal dimension with a kernel size of 3, a stride of 2, padding of 1, and 64 output channels, reducing the spatial dimension to 80×60.
[0110] The weights of the fully connected layers in the gated network are pre-trained using one week's worth of operational data from edge devices. The training objective is to minimize the flow rate prediction error while meeting real-time constraints (inference time < 40ms). During actual operation, α and β are updated every 5 frames. When CPU utilization exceeds 80%, α automatically decreases to 0.5-0.6, and the number of attention heads decreases from 8 to 4-5. When the light intensity is below 200 lux, β remains at 0.9-1.0 to retain more features.
[0111] The regression header receives the β-truncated F_temporal and outputs the flow rate through two fully connected layers (32-dimensional hidden layer). The flow rate value is uploaded to the cloud every 5 frames via the MQTT protocol. Data from the most recent hour is cached locally for subsequent analysis.
[0112] The cloud platform performs a knowledge distillation iteration every two weeks, distributing the updated model parameters via OTA. During distillation, the KL divergence between the probability distribution output by the teacher model and the output by the student model is calculated, and the student model weights are updated via backpropagation.
[0113] After 30 days of continuous testing, the average inference time of this invention was 34.2ms, the frame rate was 29.2fps, and the RMSE was 0.14m / s. It automatically reduced the load under high CPU load and had no stuttering throughout the process, verifying the effectiveness of this invention.
[0114] The lightweight and adaptive river surface velocity monitoring device provided by the present invention will be described below. The lightweight and adaptive river surface velocity monitoring device described below can be referred to in correspondence with the lightweight and adaptive river surface velocity monitoring method described above.
[0115] like Figure 7 As shown, the lightweight and adaptive river surface velocity monitoring device provided by the present invention includes: The edge adaptive gating module is used to determine the number of targets and feature target dimensions of the attention head based on the device status parameters of the edge terminal and the environmental parameters of the target water area. The image feature extraction module is used to extract local features from multiple frames of river surface images to obtain spatial local features, and then perform global self-attention weighting on the spatial local features using a target number of attention heads to obtain an attention feature sequence. The causal-temporal fusion module is used to perform causal-temporal fusion on the attention feature sequence to obtain temporal fusion features, and to extract the channel features in the preceding dimension of the temporal fusion features based on the feature target dimension to obtain the target features; The feature regression processing module is used to perform regression processing on the target features to obtain the surface flow velocity of the river.
[0116] Figure 8 A schematic diagram of the physical structure of an electronic device is provided. This device may include a processor, a communications interface, memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions from the memory to execute a lightweight and adaptive method for monitoring river surface flow velocity.
[0117] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0118] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to perform the lightweight and adaptive river surface velocity monitoring methods provided by the above methods.
[0119] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the lightweight and adaptive river surface velocity monitoring methods provided by the methods described above.
[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0121] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A lightweight and adaptive method for monitoring river surface velocity, characterized in that, include: The number of targets and the dimension of feature targets in the attention head are determined based on the device status parameters of the edge terminal and the environmental parameters of the target water area. Local features are extracted from multiple frames of river surface images to obtain spatial local features, and global self-attention weighting is applied to the spatial local features using the target number of attention heads to obtain an attention feature sequence. The attention feature sequence is subjected to causal temporal fusion to obtain temporal fusion features, and the channel features of the temporal fusion features in the preceding dimension are extracted based on the feature target dimension to obtain the target features; The target features are subjected to regression processing to obtain the surface velocity of the river.
2. The lightweight and adaptive river surface velocity monitoring method according to claim 1, characterized in that, Spatial local features are obtained by extracting local features from multiple frames of river surface images, including: Initial local features of the multi-frame river surface images are extracted using depthwise separable convolution; The initial local features are fused along the channel dimension using point convolution to obtain the spatial local features.
3. The lightweight and adaptive river surface velocity monitoring method according to claim 1, characterized in that, Global self-attention weighting is applied to the spatial local features to obtain an attention feature sequence, including: The spatial local features are projected into a query matrix, a key matrix, and a value matrix using independent point convolutions; Self-attention calculation is performed on the query matrix, the key matrix, and the value matrix to obtain the attention feature sequence.
4. The lightweight and adaptive river surface velocity monitoring method according to claim 1, characterized in that, The attention feature sequence is subjected to causal temporal fusion to obtain temporal fusion features, including: The attention feature sequences are stacked in chronological order to form a temporal feature sequence. The temporal feature sequence is subjected to causal convolution; the scope of the convolution kernel is the current and past historical frames. The temporal fusion features are obtained by downsampling the feature space dimension during the causal convolution process.
5. The lightweight and adaptive river surface velocity monitoring method according to claim 1, characterized in that, The equipment status parameters include computing power utilization and remaining power, and the environmental parameters include water flow turbulence and light intensity.
6. The lightweight and adaptive river surface velocity monitoring method according to claim 1 or 5, characterized in that, The number of targets and feature target dimensions of the attention head are determined based on the device status parameters of the edge terminal and the environmental parameters of the target water area, including: The joint feature vector containing the device state parameters and the environmental parameters is input into the gating network to obtain the attention head number adjustment factor and feature dimension adjustment factor output by the gating network. The number of targets is determined based on the attention head number adjustment factor and the attention head number benchmark, and the feature target dimension is determined based on the feature dimension adjustment factor and the feature dimension benchmark.
7. The lightweight and adaptive river surface velocity monitoring method according to claim 6, characterized in that, The target quantity is determined based on the attention head number adjustment factor and the attention head number benchmark, and the feature target dimension is determined based on the feature dimension adjustment factor and the feature dimension benchmark, including: The target number is obtained by multiplying the attention head number adjustment factor and the attention head number benchmark and then rounding down. The feature dimension target dimension is obtained by multiplying the feature dimension adjustment factor and the feature dimension benchmark and then rounding down.
8. The lightweight and adaptive river surface velocity monitoring method according to claim 1, characterized in that, After performing regression processing on the target features to obtain the river surface velocity, the process further includes: The dataset, including the multi-frame river surface images, the device status parameters, and the river surface flow velocity, is uploaded to the cloud, and the updated model parameters are obtained by the teacher model in the cloud through knowledge distillation based on the dataset. Update the local student model using the updated model parameters.
9. The lightweight and adaptive river surface velocity monitoring method according to claim 8, characterized in that, Obtain the updated model parameters obtained by knowledge distillation of the teacher model in the cloud based on the dataset, including: Obtain the flow velocity probability distribution obtained by the teacher model predicting the dataset; Calculate the relative entropy loss between the flow velocity probability distribution of the teacher model and the flow velocity probability distribution of the student model; Receive the updated model parameters obtained by the cloud after performing network backpropagation based on the relative entropy loss.
10. A lightweight and adaptive river surface velocity monitoring device, characterized in that, include: The edge adaptive gating module is used to determine the number of targets and feature target dimensions of the attention head based on the device status parameters of the edge terminal and the environmental parameters of the target water area. The image feature extraction module is used to extract local features from multiple frames of river surface images to obtain spatial local features, and to perform global self-attention weighting on the spatial local features using the target number of attention heads to obtain an attention feature sequence. The causal-temporal fusion module is used to perform causal-temporal fusion on the attention feature sequence to obtain temporal fusion features, and to extract the channel features of the temporal fusion features in the preceding dimension based on the feature target dimension to obtain the target features; The feature regression processing module is used to perform regression processing on the target features to obtain the river surface velocity.