Hydraulic engineering slope deformation monitoring method and system based on video stream
By acquiring video stream data and automatically extracting feature points, the problem of full coverage and real-time continuous monitoring of slope deformation in water conservancy projects has been solved, achieving accurate identification of slope deformation and comprehensive data support, while reducing equipment and labor costs.
Patent Information
- Application Number
- CN202511854685.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-17
AI Technical Summary
Existing methods for monitoring slope deformation in water conservancy projects suffer from several problems, including limited monitoring points, difficulty in achieving full coverage, high labor costs, inability to achieve real-time continuous monitoring, poor applicability in areas with signal obstruction, impact of lighting changes on the stability of feature points, lack of effective feature point clustering analysis mechanisms, inability to accurately identify areas of coordinated deformation, deviations in displacement calculations, overemphasis on local point displacement monitoring, and insufficient ability to identify the overall deformation trend of slopes.
By acquiring video stream data, natural feature points are automatically extracted, deformation areas are identified based on the spatial distribution characteristics of pixel displacement vectors, cluster displacement vectors are generated, and the physical spatial displacement is calculated using a coordinate transformation model, thus constructing a complete slope deformation monitoring technology system.
It enables continuous monitoring around the clock, reduces equipment costs, accurately identifies slope deformation patterns, provides timely and accurate data support, and provides a reliable basis for engineering safety assessment.
Smart Images

Figure CN121685646A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of water conservancy project safety monitoring technology, and in particular to a method and system for monitoring slope deformation in water conservancy projects based on video streams. Background Technology
[0002] The stability of slopes in water conservancy projects is directly related to the safe operation of water conservancy facilities. Traditional deformation monitoring mainly uses contact measurement methods such as total stations and inclinometers. Although these methods have high measurement accuracy, they have inherent limitations such as limited monitoring points, difficulty in achieving full coverage, high labor costs, and inability to achieve real-time continuous monitoring. At the same time, their applicability is limited in areas with signal obstruction, such as canyons, and the equipment cost and maintenance requirements are still relatively high.
[0003] With the development of computer vision technology, existing video monitoring methods mainly track displacement changes by manually setting up markers or relying on simple image feature matching. However, these methods face significant challenges in practical applications: changes in lighting and weather conditions in natural environments can affect the stability of feature points; there is a lack of effective feature point clustering analysis mechanisms, making it difficult to accurately identify areas of coordinated deformation; and in the process of converting pixel displacement to physical displacement, factors such as camera pose, lens distortion, and slope topography are often ignored, leading to deviations in displacement calculation.
[0004] In addition, existing methods mostly focus on displacement monitoring at local points, and are insufficient in identifying the overall deformation trend of slopes. They cannot effectively distinguish the deformation characteristics of different blocks and are difficult to provide comprehensive deformation field information for slope stability assessment. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art, the embodiments of this application provide a method for monitoring slope deformation in water conservancy projects based on video streams to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, this application provides a method for monitoring slope deformation in hydraulic engineering based on video streams, comprising:
[0007] The video acquisition equipment is deployed in a stable area of the slope to acquire video stream data containing the slope monitoring area, which is continuously captured by the video acquisition equipment.
[0008] Extract multiple natural feature points from the initial frame of the video stream data;
[0009] Based on the positional changes of the natural feature points in subsequent video frames, the pixel displacement vector of each natural feature point is calculated.
[0010] Based on the spatial distribution characteristics of the pixel displacement vector, natural feature points with consistent motion trends are identified as the same deformation region, and cluster displacement vectors are generated.
[0011] The physical spatial displacement of the deformed region is calculated based on the cluster displacement vector and coordinate transformation model.
[0012] To address the aforementioned problems, this application also provides a video stream-based slope deformation monitoring system for hydraulic engineering projects, the system comprising:
[0013] The video acquisition and monitoring module is used to deploy video acquisition equipment in a stable area of the slope and acquire video stream data containing the slope monitoring area continuously captured by the video acquisition equipment.
[0014] An initial feature point extraction module is used to extract multiple natural feature points from the initial frame of the video stream data;
[0015] The pixel displacement calculation module is used to calculate the pixel displacement vector of each natural feature point based on the position change of the natural feature point in subsequent video frames;
[0016] The deformation region identification module is used to identify natural feature points with consistent motion trends as the same deformation region based on the spatial distribution characteristics of the pixel displacement vector, and generate cluster displacement vectors.
[0017] The physical displacement conversion module is used to calculate the physical spatial displacement of the deformed region based on the cluster displacement vector and coordinate transformation model.
[0018] This invention continuously acquires slope images via video streams, automatically extracting natural feature points with stable optical characteristics. This replaces traditional manual inspections and sensor deployment, significantly reducing equipment costs and enabling continuous monitoring around the clock. By analyzing the spatial distribution characteristics of pixel displacement vectors, the system intelligently identifies clusters of natural feature points with consistent movement trends, accurately delineates areas of coordinated deformation, and generates cluster displacement vectors representing the overall movement trend. This overcomes the limitations of single-point monitoring and provides a comprehensive understanding of slope deformation. Based on a precise coordinate transformation model, image pixel displacements are converted into physical spatial displacements, fully considering factors such as the spatial position and shooting angle of the camera equipment to ensure the spatial accuracy of the monitoring results. This allows video monitoring data to be directly used for engineering safety assessments. From data acquisition, feature extraction, and area identification to displacement calculation, a complete slope deformation monitoring technology system is constructed, providing timely, accurate, and comprehensive data support for slope stability assessment in water conservancy projects and offering a reliable basis for engineering decisions. Attached Figure Description
[0019] Figure 1 A flowchart illustrating a video stream-based method for monitoring slope deformation in hydraulic engineering, as provided in an embodiment of this application.
[0020] Figure 2A functional block diagram of a video stream-based slope deformation monitoring system for hydraulic engineering provided in an embodiment of this application;
[0021] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0023] This application provides a method for monitoring slope deformation in hydraulic engineering projects based on video streams. The execution entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method for monitoring slope deformation in hydraulic engineering projects based on video streams can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.
[0024] Reference Figure 1 The diagram shown is a flowchart illustrating a video stream-based method for monitoring slope deformation in hydraulic engineering, according to an embodiment of this application. In this embodiment, the video stream-based method for monitoring slope deformation in hydraulic engineering includes:
[0025] S1. Deploy the video acquisition device in a stable area of the slope to acquire video stream data containing the slope monitoring area, which is continuously captured by the video acquisition device.
[0026] Video acquisition equipment is a hardware device capable of continuously capturing images and forming dynamic video sequences, enabling real-time image acquisition and output of target areas; a stable slope area refers to a slope in a water conservancy project where there is no risk of displacement, settlement, or collapse, and the spatial location of this area will not change during the monitoring period; a slope monitoring area refers to an area in a water conservancy project slope that requires close attention to deformation, as this area may have displacement risks due to changes in geological conditions, water erosion, or load effects, and is the core target area for deformation monitoring.
[0027] Video stream data refers to a collection of images that are continuously captured by video acquisition equipment and arranged in chronological order. Each image frame in this collection carries capture time information and can reflect the appearance of the target area at different points in time.
[0028] In some embodiments, deploying the video acquisition device in a stable slope area to acquire video stream data containing the slope monitoring area continuously captured by the video acquisition device includes: fixing the video acquisition device on stable bedrock facing the slope monitoring area; adjusting the shooting parameters of the video acquisition device to ensure coverage of the entire slope monitoring area; starting the video acquisition device to perform continuous image acquisition; and acquiring the video stream data containing time sequence information output by the video acquisition device.
[0029] Stable bedrock refers to a rock layer within the stable area of a slope that has a complete geological structure, high hardness, and is free from weathering or fissures, providing a stable foundation for video acquisition equipment. Shooting parameters refer to the settings used by the video acquisition equipment to adjust the image acquisition effect, mainly including lens focal length, shooting angle, image resolution, frame rate, and exposure. These parameters directly affect the image clarity and coverage. Timing information refers to the shooting time identifier corresponding to each image frame in the video stream data, usually existing in the form of a timestamp, which can be used to determine the temporal order of different image frames.
[0030] This step first involves deploying the video acquisition equipment in a stable area of the slope where it will not shift, thus avoiding interference from the equipment's own displacement on the monitoring results. Then, a video stream containing time-series information is acquired through continuous shooting, providing basic data support for subsequent tracking of feature point position changes and analysis of slope deformation.
[0031] Deploying video acquisition equipment in a stable slope area involves: first, determining the location of the stable slope area through geological survey, prioritizing areas close to the slope monitoring area and with stable bedrock geological structure; then, mechanically fixing the video acquisition equipment to the stable bedrock, specifically using expansion bolts to connect the equipment bracket to the stable bedrock, ensuring no looseness between the bracket and the bedrock; after installation, calibrating the equipment with a level to ensure that the shooting direction of the equipment does not shift due to installation tilt; finally, mounting the high-definition network camera on the bracket to achieve the deployment of the equipment in the stable slope area.
[0032] Acquiring video stream data containing the slope monitoring area continuously captured by video acquisition equipment includes: first adjusting the shooting parameters to ensure coverage of the entire slope monitoring area, then starting the equipment for continuous image acquisition, and finally acquiring video stream data with time sequence information from the equipment output.
[0033] The video acquisition equipment is fixed on stable bedrock facing the slope monitoring area. This includes: first, determining the boundary coordinates of the slope monitoring area through on-site measurement; then, using a total station to measure the azimuth angles of the four vertices of the monitoring area relative to the installation point on the stable bedrock; next, adjusting the shooting direction of the video acquisition equipment according to the azimuth angles to ensure that the line connecting the center line of the equipment lens and the center of the monitoring area is horizontal, thus ensuring that the equipment is facing the monitoring area; finally, checking the equipment's fixation status again to prevent the equipment from shifting due to vibration during long-term use.
[0034] The slope monitoring area is a rectangular area in the middle of the slope, ranging from 200m to 300m above sea level. The stable bedrock installation point is located at an altitude of 320m. The azimuth angles of the four vertices of the monitoring area relative to the installation point were measured using a total station and found to be 185°, 195°, 275°, and 285°, respectively. Then, the camera lens direction was adjusted so that the center line of the lens pointed to the center of the monitoring area (at an altitude of 250m). After the adjustment, the expansion bolts of the fixing bracket were checked with a torque wrench to ensure that the tightening torque of each bolt reached 30N·m, thus achieving the fixing of the camera facing the monitoring area.
[0035] Adjusting the shooting parameters of the video acquisition equipment to ensure coverage of the entire slope monitoring area includes: First, calculating the required shooting angle and lens focal length based on the actual size of the slope monitoring area and the distance between the equipment installation point and the monitoring area. The calculation formula is: focal length = (image sensor size × installation distance) / monitoring area size. Then, adjusting the lens focal length based on the calculation results, while also adjusting the shooting angle, image resolution, and frame rate to avoid difficulty in extracting feature points due to low resolution or omissions due to changes in feature point positions due to low frame rate. During the adjustment process, the image should be observed in real time through the equipment preview function to ensure that all four vertices of the monitoring area are within the image range, with no edges exceeding or being missed.
[0036] For example, the horizontal distance between the video acquisition device installation point and the slope monitoring area is 80m. The monitoring area is a rectangular area with a length of 50m and a width of 30m. The device's image sensor size is 1 / 2.7 inches (6.17mm long and 4.55mm wide). The required focal length is calculated using the formula (6.17mm × 80m) / 50m ≈ 9.87mm. The lens focal length is adjusted to 10mm, and the image resolution is set to 1920 × 1080, the frame rate is set to 30fps, and the exposure is adjusted to 1 / 500s. By observing the device preview screen, it is confirmed that all four vertices of the monitoring area are completely displayed in the image.
[0037] Start the video acquisition device for continuous image acquisition, including: checking the power supply status and data transmission link of the video acquisition device; then sending a start command through the device control software to set the continuous acquisition mode, which must be set to uninterrupted continuous shooting; after starting, the device acquisition status needs to be checked in real time.
[0038] The video acquisition equipment is powered by a 24V DC power supply, with a voltage regulator module installed at the power input. Data transmission is connected to the monitoring center server via a gigabit network cable. In the equipment control software at the monitoring center, the continuous acquisition mode is selected, and the acquisition duration is set to 7×24 hours. After startup, the preview screen can be viewed in real time through the software to ensure uninterrupted continuous image acquisition.
[0039] The process of acquiring video stream data containing timing information from a video capture device includes: first, enabling the timing information generation function in the video capture device and setting the timestamp format to "year-month-day hour:minute:second.millisecond" to ensure that each image frame carries a unique timestamp; then, transmitting the video stream data output by the device to the backend storage server via a data transmission link; after receiving the video stream data, the storage server parses the data, extracts each image frame and its corresponding timestamp, and stores them in the database in chronological order; finally, the data verification module verifies the stored video stream data to check for missing image frames or incorrect timestamps, and if errors are found, a retransmission mechanism is triggered.
[0040] After the video capture device enables the timestamp function, the timestamp of each image frame is accurate to the millisecond level. The device transmits the video stream to the backend 8-bay storage server via gigabit network cable using the RTSP protocol. The server parses the video stream into single-frame images (JPEG format) and corresponding timestamp files, storing them in a directory structure of "date / device number". The data verification module counts the number of image frames every hour. If it finds that the number of frames in a certain hour is less than 90,000 (30fps×3600s), it automatically sends a retransmission command to the device to ensure that the acquired video stream data is complete and contains timing information.
[0041] This step addresses the issue of monitoring errors caused by equipment displacement by deploying the video acquisition equipment on stable bedrock in the stable area of the slope, ensuring that the equipment's position remains unchanged and laying the foundation for the accuracy of subsequent monitoring data. Adjusting the shooting parameters ensures coverage of the entire slope monitoring area, resolving the problem of missed deformation detection due to omissions in the monitoring area, and achieving comprehensive monitoring of the slope's risk areas. Furthermore, acquiring continuous video stream data containing temporal information solves the problem of monitoring data lacking a time dimension, making it impossible to analyze deformation trends, and provides continuous time-series data support for subsequent tracking of feature point position changes and displacement calculation.
[0042] S2. Extract multiple natural feature points from the initial frame of the video stream data.
[0043] The initial frame refers to the first image frame selected from the video stream data for subsequent feature point extraction and analysis. This image frame carries the earliest temporal information and is usually used as a reference frame for comparing the changes in feature point positions in other video frames.
[0044] Natural feature points refer to points extracted from the initial frame that reflect the unique appearance features of the slope monitoring area. These points are usually located in areas with rich slope surface texture or clear outlines. Their optical characteristics are not easily changed significantly by changes in lighting and shadows in subsequent video frames, and can be used as marker points to track slope deformation.
[0045] In some embodiments, extracting multiple natural feature points from the initial frame of the video stream data includes: converting the initial frame of the video stream data into a grayscale image; performing corner detection on the grayscale image to obtain multiple initial feature points; and selecting the natural feature points with stable optical characteristics from the initial feature points based on the texture richness of the initial feature points.
[0046] A grayscale image is an image in which the initial frame of a color image is converted using a specific algorithm, and each pixel is represented by a grayscale value (usually ranging from 0 to 255) to indicate its brightness. The larger the grayscale value, the brighter the pixel, and the smaller the grayscale value, the darker the pixel.
[0047] Corner detection refers to the process of identifying points in a grayscale image whose pixel grayscale values change drastically and have obvious grayscale changes in two orthogonal directions. These points are usually the intersections of object edges or the turning points of contours in the image, and have the potential to become feature points. Initial feature points refer to points that are initially identified from grayscale images through corner detection and may have tracking value. These points have not been screened and may include some points with unstable optical features or insufficient texture information.
[0048] Texture richness is an indicator used to measure the complexity of the image texture in the region surrounding an initial feature point. It is usually determined by calculating the frequency of change and gradient difference of pixel gray values in a fixed-size region around the feature point. The higher the texture richness, the richer the image details in that region, and the stronger the uniqueness and stability of the feature point.
[0049] Stable optical features refer to the characteristic that the optical properties of a natural feature point, such as the grayscale distribution and gradient features of its surrounding area, remain relatively stable under different lighting conditions and shooting angles. Feature points with this characteristic can still be accurately identified and tracked in subsequent video frames.
[0050] The core purpose of this step is to select stable natural feature points from the initial frames of the video stream data, providing reliable marker objects for subsequent calculation of feature point position changes and analysis of slope deformation. First, the initial frame is converted into a grayscale image to reduce the interference of color information on feature extraction. Then, initial feature points are initially obtained through corner detection. Finally, natural feature points with stable optical features are selected based on texture richness to ensure that feature points are not easily lost or mismatched during subsequent tracking.
[0051] Converting the initial frame of the video stream data into a grayscale image includes: extracting the initial frame from the video stream data obtained in step S1. The initial frame must be selected to meet the conditions of clear image and no obvious occlusion (such as no large number of trees or weeds obscuring the slope monitoring area); then using a weighted average method to convert the color image of the initial frame into a grayscale image. The specific calculation formula is grayscale value = 0.299×R + 0.587×G + 0.114×B, where R represents the red channel value of the pixel in the initial frame, G represents the green channel value, and B represents the blue channel value. The grayscale value of each pixel is calculated using this formula, and finally a grayscale image containing only grayscale information is generated.
[0052] The first image frame is extracted from the video stream as the initial frame. The resolution of the initial frame is 1920×1080. There is no obvious occlusion in the slope monitoring area in the image. The color image of the initial frame is converted into a grayscale image using a weighted average method. The BGR color image of the initial frame is converted into a grayscale image. After conversion, the details such as rock texture and crack outline of the slope surface are clearly visible in the grayscale image. The basis of this implementation is that the human eye is most sensitive to green light, followed by red light, and least sensitive to blue light. In the weighted average method, the weight of the G channel (0.587) is the highest, followed by the R channel (0.299), and the weight of the B channel (0.114) is the lowest. This is consistent with the characteristics of human vision. It can remove the redundancy of color information while retaining the key details of the initial frame, reduce the data processing volume of subsequent corner detection, and improve the efficiency of feature extraction.
[0053] Corner detection is performed on the grayscale image to obtain multiple initial feature points. The Shi-Tomasi corner detection algorithm is then applied, with the following specific steps: First, the grayscale image is Gaussian filtered and smoothed using convolution operations to reduce image noise interference with the detection results. Second, the gradient values of each pixel in the grayscale image in the x and y directions are calculated. The Sobel operator is used to convolve the grayscale image in the x and y directions respectively to obtain the gradient matrix in the x direction of each pixel. and the gradient matrix in the y-direction The third step is to calculate the autocorrelation matrix M for each pixel based on the gradient matrix. The formula for calculating the autocorrelation matrix M is as follows: During the calculation process, it is necessary to... The first step is to perform a summation operation within a 3×3 neighborhood. The second step is to solve for the two eigenvalues λ1 and λ2 of the autocorrelation matrix M. According to the Shi-Tomasi algorithm's judgment rule, when the smaller of the two eigenvalues is greater than the preset eigenvalue threshold, the pixel is judged as a corner point. The third step is to perform non-maximum suppression processing on the detected corner points, retaining only the corner point with the largest eigenvalue within the 3×3 neighborhood to avoid the concentrated distribution of corner points, and finally obtaining multiple discretely distributed initial feature points.
[0054] The Gaussian template size was set to 5×5 with a standard deviation of 1.0, the Sobel operator's convolution kernel size was set to 3×3, the eigenvalue threshold was set to 200, and the non-maximum suppression neighborhood size was set to 3×3. During the detection process, the noise in the grayscale image was significantly reduced after Gaussian filtering, while the fine texture of the rock surface remained clear. The gradient matrix calculated by the Sobel operator accurately reflects the trend of pixel grayscale changes. After solving the autocorrelation matrix, 1200 pixels with smaller eigenvalues than 200 were selected. After non-maximum suppression processing, 350 initial feature points were finally obtained. These initial feature points are mainly distributed at the rock edges and crack intersections on the slope surface.
[0055] Gaussian filtering can effectively smooth image noise, which can lead to false corners. Smoothing can improve the accuracy of corner detection. The Sobel operator can efficiently calculate pixel gradients, accurately reflecting the direction and intensity of gray-level changes in the image, providing reliable data for subsequent autocorrelation matrix calculation. The Shi-Tomasi algorithm can more accurately identify corners with stable tracking potential by judging whether the smaller value of two feature values exceeds a threshold. Non-maximum suppression processing can avoid excessive density of initial feature points in local areas, ensuring uniform distribution of feature points during subsequent screening and tracking.
[0056] Based on the texture richness of the initial feature points, natural feature points with stable optical characteristics are selected from the initial feature points. This includes: first, determining the texture richness calculation region for each initial feature point, selecting an 11×11 square region as the texture analysis region centered on each initial feature point; then calculating the texture richness within this region, specifically by measuring the variance of the pixel grayscale values within the region, using the variance calculation formula as follows: variance... ,in, The variance represents the grayscale value of each pixel within the region, μ represents the average grayscale value of all pixels within the region, and n represents the total number of pixels within the region (n is 121 for an 11×11 region). The larger the variance, the greater the difference in grayscale values among pixels within the region, and the higher the texture richness. Then, a texture richness threshold is set, which needs to be determined based on the overall grayscale distribution of the grayscale image. Usually, the threshold is set to 1.2 times the average texture richness of all initial feature points. Finally, initial feature points with texture richness greater than the threshold are selected. The surrounding areas of these feature points have rich texture, and their optical characteristics are not easily changed by environmental changes in subsequent video frames. These feature points are identified as natural feature points with stable optical characteristics.
[0057] Centered on each initial feature point, an 11×11 image region was extracted. The mean and variance of pixel grayscale values within each region were calculated using Python's NumPy library. The texture richness (variance) of 350 initial feature points was statistically analyzed, yielding an average value of 85. The texture richness threshold was set to 85×1.2=102, resulting in 220 initial feature points with a variance greater than 102. Within the texture analysis regions of these feature points, the alternation of light and dark rock textures was obvious. For example, some feature points were located at the boundaries of rock bedding, with grayscale values ranging from 50 to 200 and a variance of 130, indicating high texture richness. Among the initial feature points that were removed, some were located in gentle areas of the slope surface, with grayscale values concentrated between 120 and 140 within the texture analysis regions and a variance of only 60, indicating low texture richness. These 220 selected feature points were ultimately used as natural feature points for subsequent steps.
[0058] Regions with rich textures typically contain more unique image details, and their corresponding feature points have higher recognizability. Even with changes in lighting (such as an increase in overall brightness from cloudy to sunny) or slight occlusion (such as a few fallen leaves drifting by), the texture structure around these feature points can still be identified, resulting in stronger optical feature stability. In contrast, regions with sparse textures have poor feature point uniqueness and are easily confused with surrounding pixels, leading to feature point loss or mismatch during subsequent tracking. Calculating texture richness using variance can quantify the severity of grayscale changes within a region, providing an objective basis for selecting stable feature points. Setting a threshold based on the average value ensures that the selection criteria are adapted to the actual situation of the grayscale image, avoiding an excessively high threshold that results in an insufficient number of natural feature points, or an excessively low threshold that results in unstable feature points remaining after selection.
[0059] This step addresses the problem of low feature extraction efficiency caused by the large amount of data and redundant information in color images. It uses the Shi-Tomasi corner detection algorithm to obtain initial feature points, solving the problem of insufficient initial feature points due to the weak recognition ability of traditional feature extraction algorithms in low-contrast areas of slope surfaces. This algorithm can accurately identify corner points in low-contrast areas such as rock edges and cracks on slope surfaces, ensuring that the number of initial feature points meets the needs of subsequent analysis. Furthermore, by selecting natural feature points based on texture richness, it solves the problem of unstable optical features of extracted feature points, which are easily lost or mismatched in subsequent tracking, leading to large errors in deformation analysis.
[0060] S3. Based on the positional changes of the natural feature points in subsequent video frames, calculate the pixel displacement vector of each natural feature point.
[0061] The current video frame refers to a single image frame selected from subsequent video frames that is temporally adjacent to the initial frame. Selecting this frame can minimize the interference of non-deformation factors of the slope (such as slow changes in illumination) on the determination of the location of feature points within the time interval. The correspondence of natural feature points refers to finding points in the current video frame that match the location and have the same optical characteristics as each natural feature point in the initial frame. Through this relationship, the location of the same natural feature point in different frames can be determined.
[0062] In some embodiments, calculating the pixel displacement vector of each natural feature point based on the position changes of the natural feature points in subsequent video frames includes: selecting a current video frame adjacent to the initial frame in the video stream data; tracking the position of the natural feature points in the current video frame and establishing a correspondence between the natural feature points; calculating the coordinate difference between the natural feature points in the initial frame and the current video frame; and generating the pixel displacement vector of each natural feature point based on the coordinate difference.
[0063] Subsequent video frames refer to all image frames in the video stream data that are in chronological order after the initial frame; coordinate difference refers to the numerical difference between the coordinates of the same natural feature point in the initial frame and its coordinates in the current video frame, including the coordinate difference in the x-axis direction and the coordinate difference in the y-axis direction, which can intuitively reflect the change in the position of the feature point in the image plane.
[0064] A pixel displacement vector is a vector representation of the displacement of a natural feature point in the image plane. This vector contains the magnitude of the displacement (calculated from the coordinate difference) and the direction of the displacement (determined by the sign of the difference between the x-axis and y-axis coordinates), and can completely describe the motion state of the feature point at the pixel level.
[0065] Image coordinates refer to a two-dimensional coordinate system used to locate the position of pixels in an image. The origin is usually the top left corner of the image, with the positive x-axis pointing to the right and the positive y-axis pointing downwards. The coordinate values are in pixels. For example, the coordinates of a pixel in the image can be represented as (x, y), where x represents the pixel number in the horizontal direction and y represents the pixel number in the vertical direction.
[0066] The core purpose of this step is to transform the positional changes of natural feature points into quantifiable pixel displacement vectors, providing pixel-level basic data for subsequent calculation of slope physical spatial displacement. First, the current video frame adjacent to the initial frame is selected to reduce interference. Then, the correspondence between feature points is established through tracking. Next, the coordinate difference is calculated. Finally, a pixel displacement vector containing magnitude and direction is generated to ensure that the motion state of each natural feature point can be accurately described.
[0067] Selecting the current video frame adjacent to the initial frame in the video stream data includes: first, obtaining the timing information of all image frames in the video stream data, sorting the image frames by timestamp from smallest to largest, and determining the position of the initial frame in the sorted sequence; then, determining the time interval between adjacent frames based on the frame rate of the video stream. If the frame rate is 30fps, the time interval between adjacent frames is 1 / 30 second; finally, selecting the next image frame after the initial frame from the sorted sequence as the current video frame, ensuring that the current video frame is temporally continuous with the initial frame and there are no other image frames in between.
[0068] The video stream data obtained in step S1 has a frame rate of 30fps, and the timestamp of the initial frame is "2025-11-22 10:00:00.000". After sorting all the image frames by timestamp, the image frame with the timestamp "2025-11-22 10:00:00.033" is found (this timestamp is 1 / 30 of a second from the initial frame, about 0.033 seconds), and it is determined as the current video frame adjacent to the initial frame.
[0069] The time interval between adjacent frames is extremely short. If the slope deforms within this time, the amount of deformation is small and the influence of non-deformation factors (such as changes in light and wind blowing weeds) can be ignored, which can minimize the false changes in the position of feature points caused by non-deformation factors.
[0070] The location of the natural feature points in the current video frame is tracked, and the correspondence between the natural feature points is established using the Lucas-Kanade optical flow method. The specific steps are as follows:
[0071] The first step is to construct an 8×8 neighborhood window for each natural feature point in the initial frame. This window contains the feature point and the surrounding pixels, and the pixel grayscale distribution within the window serves as the optical feature template for that feature point.
[0072] The second step is to set a 16×16 search window in the current video frame, centered on the coordinates of the natural feature points in the initial frame. The search window is larger than the neighborhood window to ensure that it can cover the area that the feature points may move to.
[0073] The third step is to calculate the gray-level similarity between the neighborhood window and all possible 8×8 sub-windows within the search window. The normalized cross-correlation coefficient (NCC) is used as the similarity measure. The NCC value ranges from -1 to 1. The closer the NCC value is to 1, the more similar the gray-level distributions of the two windows are.
[0074] The fourth step is to determine the center coordinates of the sub-window with the largest NCC value that is greater than the preset similarity threshold within the search window as the position in the current video frame that corresponds to the natural feature point of the initial frame.
[0075] The fifth step is to repeat the above operation for each natural feature point, and finally establish a one-to-one correspondence between the natural feature points in the initial frame and the current video frame.
[0076] Based on step S2, 220 natural feature points were selected. An 8×8 neighborhood window was constructed for each natural feature point, and a 16×16 search window was set in the current video frame. The NCC similarity threshold was set to 0.85. During the tracking process, for the natural feature point with coordinates (500, 300) in the initial frame, the gray values in its neighborhood window were mainly distributed between 80 and 150. In the search window of the current video frame, a sub-window with an NCC value of 0.92 (greater than the threshold of 0.85) was found. The center coordinates of this sub-window were (501, 302). This coordinate was determined as the corresponding position of the natural feature point in the current video frame. Finally, among the 220 natural feature points, 215 successfully established a correspondence, and 5 were not matched because the NCC value was lower than the threshold (they will be removed later).
[0077] The Lucas-Kanade optical flow method is based on the assumption of grayscale invariance, meaning that the grayscale distribution of the same feature point remains unchanged in adjacent frames. By constructing optical feature templates and searching similar sub-windows, it can accurately find the position of feature points in the current frame. An 8×8 neighborhood window can contain enough grayscale information to ensure feature uniqueness, a 16×16 search window can cover the maximum possible displacement range of feature points in a short time, and an NCC threshold of 0.85 can effectively eliminate false matches caused by noise and small changes in illumination, ensuring the accuracy of the correspondence.
[0078] Calculating the coordinate difference between the natural feature point in the initial frame and the current video frame includes: first, obtaining the coordinates (x0, y0) of the natural feature point in the initial frame and the coordinates (x1, y1) in the current video frame, with the coordinate values determined by the image coordinate system; then, calculating the coordinate difference in the x-axis direction and the y-axis direction respectively. The formula for calculating the x-axis coordinate difference Δx is Δx = x1 - x0, and the formula for calculating the y-axis coordinate difference Δy is Δy = y1 - y0. If Δx is positive, it means that the feature point has moved in the positive x-axis direction (horizontally to the right) relative to the initial frame in the current frame; if Δx is negative, it means that it has moved in the negative x-axis direction (horizontally to the left). Similarly, a positive Δy value indicates movement in the positive y-axis direction (vertically downward), and a negative Δy value indicates movement in the negative y-axis direction (vertically upward).
[0079] A natural feature point has coordinates of (500, 300) in the initial frame and (501, 302) in the current video frame. According to the formula, Δx = 501 - 500 = 1 pixel and Δy = 302 - 300 = 2 pixels, indicating that the feature point has moved 1 pixel horizontally to the right and 2 pixels vertically downward relative to the initial frame in the current frame. Another natural feature point has initial coordinates of (800, 450) and current coordinates of (799, 451). Δx = -1 pixel and Δy = 1 pixel are calculated, indicating that the feature point has moved 1 pixel horizontally to the left and 1 pixel vertically downward.
[0080] Image coordinate systems can accurately locate pixel positions. By calculating coordinate differences, the displacement of feature points in the x and y axes can be directly quantified. The sign of the coordinate difference can intuitively reflect the direction of displacement. This calculation method is simple and efficient, and can quickly process coordinate data of a large number of natural feature points.
[0081] Generating the pixel displacement vector for each of the natural feature points based on the coordinate differences includes: first determining the representation of the pixel displacement vector, using a two-dimensional vector (Δx, Δy), where Δx is the x-axis coordinate difference and Δy is the y-axis coordinate difference; then calculating the magnitude of the vector, using the formula: The unit is pixels, and this size reflects the total displacement of the feature point in the image plane; next, the direction of the vector is determined by calculating the angle θ between the vector and the positive x-axis. The formula for calculating the angle θ is: The angle range is arrive Finally, the Δx, Δy, vector magnitude, and vector direction of each natural feature point are integrated to form a complete pixel displacement vector.
[0082] For a given natural feature point, Δx = 1 pixel and Δy = 2 pixels. Calculate the vector size using the formula. Pixels, included angle The pixel displacement vector of this feature point is represented as (1, 2), with a displacement of approximately 2.24 pixels and a direction perpendicular to the positive x-axis. Angle; another feature point Δx = -1 pixel, Δy = 1 pixel, calculate the vector size | Pixels, included angle Its pixel displacement vector is represented as (-1, 1), with a displacement magnitude of approximately 1.41 pixels and a direction perpendicular to the positive x-axis. horn.
[0083] Two-dimensional vectors can simultaneously contain the magnitude and direction of displacement, providing a more complete description of the motion state of feature points compared to individual coordinate differences. The calculation of vector magnitude and angle is based on plane geometry principles, resulting in accurate and reliable results that can provide clear vector data support for subsequent analysis of the motion trend of deformed regions.
[0084] This step addresses the issues of excessive interference from non-deformation factors and numerous pseudo-changes in feature point positions caused by long time intervals by selecting the current video frame adjacent to the initial frame. It also solves the problems of low matching accuracy and easy loss of feature points in traditional tracking algorithms when tracking feature points and establishing correspondences using the Lucas-Kanade optical flow method. Furthermore, by calculating coordinate differences and generating pixel displacement vectors, it addresses the problem of unquantifiable feature point position changes and difficulty in describing motion states, transforming position changes into vectors containing both magnitude and direction, providing quantifiable data support for subsequent identification of deformable areas and calculation of physical displacement.
[0085] S4. Based on the spatial distribution characteristics of the pixel displacement vector, natural feature points with consistent motion trends are identified as the same deformation region, and cluster displacement vectors are generated.
[0086] Spatial distribution characteristics refer to the distribution pattern of pixel displacement vectors in the image plane, mainly including the consistency of vector direction, the correlation of vector magnitude, and the spatial position correlation of corresponding natural feature points. This feature can reflect the motion coordination of natural feature points in different regions. Consistent motion trend means that the pixel displacement vectors of multiple natural feature points are close in direction and similar in magnitude, and these feature points are adjacent in spatial position in the image, indicating that they are driven by the deformation behavior of the same slope.
[0087] The deformation zone refers to the physical area corresponding to natural feature points with the same movement trend in the slope monitoring area. The slope rock or soil in this area undergoes synchronous displacement during the monitoring period. The cluster displacement vector is a vector that can represent the overall movement state of all natural feature points in a certain deformation zone, and can reflect the overall displacement magnitude and direction of the deformation zone.
[0088] In some embodiments, the step of identifying natural feature points with consistent motion trends as the same deformation region based on the spatial distribution characteristics of the pixel displacement vectors and generating cluster displacement vectors includes: analyzing the directional consistency of the pixel displacement vectors to form a displacement vector set; dividing natural feature points with consistent motion trends into the same deformation region based on the directional similarity of each pixel displacement vector in the displacement vector set; performing a composite calculation on the pixel displacement vectors of all natural feature points within the deformation region; and generating a cluster displacement vector representing the overall motion trend of the deformation region based on the composite calculation result.
[0089] The displacement vector set refers to the dataset formed by organizing all valid pixel displacement vectors (excluding feature point vectors without established correspondence) obtained in step S3 according to the natural feature point number. This set contains the vector magnitude, direction and corresponding coordinate information of each feature point. Directional similarity refers to the degree of closeness between two pixel displacement vectors in the direction. It is usually measured by calculating the difference between the direction angles of the two vectors. The smaller the difference, the higher the directional similarity.
[0090] In some embodiments, dividing natural feature points with consistent motion trends into the same deformation region based on the directional similarity of each pixel displacement vector in the displacement vector set includes: calculating the direction angle of each pixel displacement vector in the displacement vector set to form a direction angle set; constructing a direction similarity matrix based on the direction angle set; constructing a graph model with the natural feature points as nodes and direction similarity and spatiotemporal proximity as edges; clustering the natural feature points using a graph theory-based community detection algorithm based on the graph model to form an initial deformation region, wherein the edge weights are jointly determined by a weighted function of direction similarity and spatial distance; and confirming the final deformation region based on the spatial distribution continuity of the natural feature points in the initial deformation region.
[0091] The orientation angle refers to the angle between the pixel displacement vector and the positive x-axis of the image coordinate system, ranging from -180° to 180°. It can be calculated from the x and y-axis components of the vector using the arctangent function. The orientation angle set is a dataset formed by arranging the orientation angles of each pixel displacement vector in the displacement vector set according to the feature point number. The orientation similarity matrix is a two-dimensional matrix with natural feature points as rows and columns, and matrix elements representing the orientation similarity of the pixel displacement vectors of corresponding two feature points. The matrix size is N×N (N is the number of effective natural feature points), which can intuitively reflect the degree of orientation correlation between all feature points.
[0092] A graph model is a mathematical model used to describe the relationships between natural feature points. In this model, each natural feature point is defined as a "node", and the relationship between two nodes is defined as an "edge". The attributes (weights) of the edges are determined by the directional similarity and spatiotemporal proximity between the nodes. Spatiotemporal proximity refers to the degree of closeness between two natural feature points in the image space and their correlation in the time dimension. In this embodiment, only spatial proximity is considered (since the time dimension is the interval between the initial frame and the adjacent frame). It is measured by calculating the Euclidean distance between feature points. The smaller the distance, the higher the spatial proximity.
[0093] Community detection algorithms based on graph theory refer to algorithms that divide closely connected nodes into the same "community" (i.e., the initial deformed region) by analyzing the weights of edges between nodes in a graph model. This embodiment uses the Louvain algorithm, which can efficiently process a large number of nodes while ensuring the accuracy of the division.
[0094] The weight of an edge refers to the numerical value of the edge connecting two nodes in a graph model. In this embodiment, the weight is calculated by a weighted function of directional similarity and spatial distance. The formula is: weight = weight coefficient × directional similarity + (1 - weight coefficient) × (1 / spatial distance). The weight coefficient is 0.6, directional similarity is given priority, and the unit of spatial distance is pixels.
[0095] In some embodiments, confirming the final deformed region based on the spatial distribution continuity of natural feature points in the initial deformed region includes: calculating the spatial distance between each natural feature point in the initial deformed region and generating a distance matrix; determining whether the natural feature points form a connected region based on the distance matrix and a preset spatial distance threshold; and confirming the connected region that satisfies the condition of spatial distribution continuity as the final deformed region.
[0096] In some embodiments, generating a cluster displacement vector representing the overall motion trend of the deformed region based on the synthetic calculation results includes: calculating corresponding weighting coefficients based on the texture richness of each natural feature point in the deformed region; using the weighting coefficients to perform weighted synthesis on the pixel displacement vectors of each natural feature point in the deformed region; and generating the cluster displacement vector representing the overall motion trend of the deformed region based on the weighted synthesis results.
[0097] The initial deformation region refers to a temporary region obtained by clustering natural feature points using a community detection algorithm. This region only satisfies the directional similarity condition, and its spatial distribution continuity has not yet been verified. Spatial distribution continuity means that the natural feature points in the initial deformation region are continuously distributed in the image space without obvious spatial gaps. This can be determined by whether the spatial distance between feature points is less than a preset threshold. The distance matrix is a two-dimensional matrix with the natural feature points in the initial deformation region as rows and columns, and the matrix elements as the Euclidean distance between corresponding two feature points. It is used to quantify the spatial positional relationship between feature points.
[0098] The spatial distance threshold is a critical distance used to determine whether natural feature points are continuously distributed within the initial deformation area. In this embodiment, the size of the texture richness calculation area (11×11 pixels) in step S2 is set to 22 pixels (i.e., twice the side length of the texture analysis area) to ensure that the analysis areas of adjacent feature points overlap and meet the continuity requirement.
[0099] A connected region is a region in which the spatial distance between all natural feature points within the initial deformation region is less than the spatial distance threshold, and any two feature points can be indirectly connected through other feature points (i.e., there are no isolated points).
[0100] The weighting coefficient is a numerical value used to measure the importance of the pixel displacement vector of each natural feature point within the deformation area. This coefficient is calculated based on the texture richness of the feature point. The higher the texture richness, the larger the weighting coefficient, indicating that the displacement data of the feature point is more reliable. Weighted synthesis refers to the process of summing the pixel displacement vectors of all natural feature points within the deformation area according to their respective weighting coefficients, and then dividing by the sum of the weighting coefficients to obtain a vector representing the overall movement trend of the area.
[0101] The core objective of this step is to identify regions driven by the same deformation from scattered natural feature point displacement data and generate their overall displacement vectors, providing regional-level data for subsequent calculations of physical spatial displacement. First, the initial deformation regions are divided by directional similarity to exclude isolated points with large directional differences; then, the spatial continuity of feature points within the region is verified to ensure that the region corresponds to the actual slope deformation region; finally, cluster displacement vectors are synthesized based on texture richness weighting to highlight the contribution of reliable feature points.
[0102] The consistency of pixel displacement vector direction is analyzed to form a displacement vector set, including: first, selecting valid pixel displacement vectors from the results of step S3 and removing natural feature points that have not established a corresponding relationship; then, establishing association data for each valid feature point as "number-vector x component-vector y component-feature point image coordinates-texture richness", where the vector x and y components come from the coordinate difference in step S3, the feature point image coordinates are the coordinates recorded when extracting natural feature points, and the texture richness is the variance value calculated in step S2; finally, all association data are sorted according to the feature point number to form a displacement vector set.
[0103] Step S3 yields 215 valid pixel displacement vectors. A dataset is created using an Excel spreadsheet. The record for feature point number 1 is “1-1-2-(500,300)-130” (x component 1, y component 2, coordinates (500,300), texture richness 130), and the record for number 2 is “2-(-1)-1-(800,450)-110”. This process is repeated to complete the organization of the 215 records, forming a set of displacement vectors.
[0104] The displacement vector set needs to integrate vector data, spatial location data, and reliability data (texture richness) to provide a complete data source for subsequent orientation analysis, spatial verification, and weighted synthesis. Filtering valid vectors can avoid invalid data (such as unmatched feature points) from interfering with the clustering results.
[0105] Calculate the direction angle of each pixel displacement vector in the displacement vector set to form a direction angle set. This includes: first, calculating the direction angle θ using the arctangent function for the x-component (Δx) and y-component (Δy) of each vector in the displacement vector set, as shown in the formula. The calculation results are rounded to one decimal place and the unit is degrees. Then, the number of each feature point is associated with the corresponding orientation angle, and the orientation angles are sorted by number to form a set of orientation angles.
[0106] For feature point 1, Δx=1 and Δy=2 are calculated as follows: For feature point 2, Δx = -1 and Δy = 1, the calculation yields... The θ values of 215 feature points are associated with their numbers to form a set of direction angles. This is based on the fact that the direction angle is the only indicator of the direction of a quantized vector, and the consistency of the direction angle can be ensured by calculating it using a unified arctangent function.
[0107] Constructing a direction similarity matrix based on the set of direction angles includes: First, determining the calculation rules for direction similarity: If the absolute value of the difference in direction angles between two feature points i and j is... If the directional similarity is 1, then the similarity is 1; if If the directional similarity is 0.5, then the similarity is 0.5; if Then the directional similarity is 0 ( The threshold is set based on the characteristic that slope deformation is usually a gentle movement with small directional differences; then, all feature point pairs (i,j) in the direction angle set are traversed, and the direction similarity of each pair is calculated according to the rules; finally, the calculation results are filled into the corresponding matrix elements with the feature point number as the row and column, forming an N×N direction similarity matrix (N=215); for example, in the above project, feature point 1 ( With feature point 3 The difference is The similarity is 1; feature point 1 and feature point 5 The difference is The similarity is 0.5; feature point 1 and feature point 2 The difference is The similarity is 0, and the 215×215 matrix is constructed according to this rule.
[0108] A graph model is constructed using natural feature points as nodes and directional similarity and spatiotemporal proximity as edges. The process includes: first, defining each natural feature point as a node in the graph model, with node attributes including feature point number and image coordinates; then, calculating the spatial distance d between each pair of nodes (i,j) using the Euclidean distance formula; next, calculating the weight of the edge between each pair of nodes using the weight formula: weight = 0.6 × directional similarity + 0.4 × (1 / spatial distance) (α = 0.6, prioritizing directional similarity). If the spatial distance d = 0 (same node), the weight is set to 0; if the directional similarity = 0, the weight is also set to 0 (directly excluding node pairs with large directional differences); finally, integrating the nodes and weighted edges to form the graph model, retaining only edges with weights > 0.
[0109] The directional similarity between feature point 1 (coordinates (500, 300)) and feature point 4 (coordinates (505, 303)) is 1, and their spatial distance is... For each pixel, the weight is calculated to be 0.6×1+0.4×(1 / 5.83)≈0.668, and an edge with a weight of 0.668 is constructed between the two nodes. The directional similarity between feature point 1 and feature point 2 is 0, and the weight is 0, so no edge is constructed.
[0110] The graphical model needs to consider both directional coordination and spatial correlation to avoid classifying feature points with the same direction but isolated in space into the same region. Setting α=0.6 can ensure that the direction, the core deformation index, plays a dominant role, while spatial distance can help verify the continuity of the region.
[0111] Based on the graph model, a graph theory-based community detection algorithm is used to cluster natural feature points to form initial deformable regions. The Louvain algorithm is then used to partition the graph model into communities. The implementation steps of this algorithm are as follows: First, for each node, calculate the modularity gain for assigning it to an adjacent community or retaining it in the original community. Modularity is an indicator of the quality of community partitioning; the higher the value, the better the partitioning. Second, according to the principle of maximizing modularity gain, assign each node to the optimal community to form an initial community partition. Third, treat each initial community as a "super node" and calculate the edge weights between super nodes (i.e., the sum of edge weights between nodes in the original community and nodes in another community) to construct a new graph model. Fourth, repeat steps one through three until the modularity no longer increases. Then, from the final community partitioning results, select communities containing ≥5 feature points (removing isolated small communities to avoid noise interference). Each community that meets the criteria is an initial deformable region.
[0112] The Louvain algorithm was used to cluster 215 nodes, resulting in 8 communities. Six communities contained 32, 28, 45, 38, 22, and 18 feature points, respectively, while two communities contained only 2 and 3 feature points (removed). These six communities were used as the initial deformation regions, and the feature point numbers in each region were compiled into a list (e.g., initial deformation region 1 contained feature points numbered 1-32).
[0113] The Louvain algorithm can efficiently process large-scale nodes and ensure that nodes within a community are closely connected and nodes between communities are sparsely connected by maximizing modularity. This meets the characteristic of "strong correlation of feature points in the same deformable region". By screening communities with ≥5 feature points, false deformable regions caused by a single abnormal feature point can be excluded.
[0114] Based on the spatial continuity of the natural feature points in the initial deformation region, the final deformation region is determined, including: The first step is to calculate the spatial distance between each natural feature point in the initial deformation region, generating a distance matrix: For each of the M feature points in the initial deformation region, all feature point pairs (i,j) are traversed, and the spatial distance d_ij is calculated using the Euclidean distance formula. This distance is then filled into the matrix by rows and columns, forming an M×M distance matrix. The second step, based on the distance matrix and a preset spatial distance threshold (22 pixels), determines whether the natural feature points form a connected region: For each element in the distance matrix, if d_ij ≤ 22 pixels, it is marked... If a feature point is marked as "connected", it is marked as "disconnected". Next, check if there are any isolated points in the region (i.e., the d_ij of a certain feature point and all other feature points are greater than 22 pixels). If there are no isolated points and any two feature points can be indirectly connected through the "connectivity" relationship (e.g., if i and k are connected, and k and j are connected, then i and j are connected), then the region is a connected region. The third step is to confirm the connected regions that meet the spatial distribution continuity condition as the final deformable regions. If the initial deformable region has isolated points or cannot form a connected region, then the isolated points in the region are removed and the region is re-verified. If the number of remaining feature points is less than 5, then the region is abandoned.
[0115] The initial deformation region 1 contains 32 feature points. After calculating its distance matrix, it was found that the d_ij of all feature point pairs is ≤18 pixels (≤22 pixels), and any two feature points are connected through other feature points (e.g., feature point 1 and feature point 32 are indirectly connected through feature points 15 and 22). There are no isolated points, so it is confirmed as the final deformation region 1. The initial deformation region 3 contains 45 feature points. The distance matrix shows that the d_ij of feature point number 105 is >30 pixels (isolated point) with all other feature points. After removing this point, 44 feature points remain. The distance matrix is recalculated, and all d_ij are ≤20 pixels, forming a connected region, which is confirmed as the final deformation region 3.
[0116] Slope deformation typically manifests as displacement over a continuous area, without spatially isolated clusters of feature points aligned in the same direction. Therefore, verifying continuity through a spatial distance threshold ensures that the final deformation area corresponds to the actual physical deformation area of the slope, and eliminating isolated points avoids interference from anomalous feature points.
[0117] The pixel displacement vectors of all natural feature points within the deformation region are synthesized and calculated, including: firstly, extracting the "vector x component (Δx_i), vector y component (Δy_i), and texture richness (R_i)" data of all feature points within the final deformation region from the displacement vector set; then, calculating the weighting coefficient w_i for each feature point based on the texture richness, using the formula: ,in, The sum of texture richness of all feature points within the deformation region is used, with a weighting coefficient sum of 1; finally, the x and y components of the vector are weighted and summed separately to calculate the x component of the synthesized vector. ,y component ΔX and ΔY are the results of the composite calculation.
[0118] The final deformed region 1 contains 32 feature points, with a total texture richness ΣR_i=3200 (e.g., feature point 1 has R_i=130, w_i=130 / 3200≈0.0406; feature point 5 has R_i=150, w_i=150 / 3200≈0.0469); the weighted sum of Δx_i of all feature points in this region. Δy_i weighted sum Therefore, the result of the composite calculation is (1.2, 1.9).
[0119] The higher the texture richness of a feature point, the more stable its optical features, and the more reliable the displacement vector calculated in step S3. By highlighting its contribution through weighting coefficients, the influence of sparse and easily disturbed feature points on the synthesis result can be reduced, making the synthesized vector more reflective of the true displacement of the deformed area.
[0120] Based on the results of the synthetic calculation, a cluster displacement vector representing the overall motion trend of the deformed region is generated, including: first, using ΔX (x-component) and ΔY (y-component) from the synthetic calculation results as the fundamental components of the cluster displacement vector; then calculating the magnitude and direction of the cluster displacement vector, with the magnitude formula as follows: (Unit: pixels), direction formula is: (Unit: degrees); Finally, the "deformation region number - ΔX - ΔY - vector size - vector direction" are integrated into structured data to form the cluster displacement vector of the deformation region.
[0121] The final composite calculation result of deformed region 1 is (1.2, 1.9), and the calculated vector size is... Pixels, orientation Therefore, the cluster displacement vector of this region is recorded as " ".
[0122] This step addresses the problems of traditional clustering algorithms (such as K-means) requiring a preset number of clusters and easily misclassifying different deformation regions by constructing an directional similarity matrix and a graphical model, and using the Louvain algorithm for clustering. The Louvain algorithm can automatically identify the optimal number of clusters, and by combining directional similarity and spatial proximity, the segmentation accuracy is significantly improved. By verifying the continuity of spatial distribution, the problem of false deformation regions caused by clustering only by direction is solved, ensuring that the final deformation regions correspond to the real physical regions of the slope. By using weighted synthesis based on texture richness, the problem of treating all feature points equally and the large interference of abnormal data is solved, highlighting the contribution of reliable feature points and making the cluster displacement vector more accurate. Overall, the transformation from scattered feature points to regional vectors is realized.
[0123] S5. Calculate the physical spatial displacement of the deformed region based on the cluster displacement vector and coordinate transformation model.
[0124] Physical spatial displacement refers to the magnitude and direction of displacement of the deformed area in real physical space. It can directly reflect the actual degree of deformation of the slope rock or soil and is a key indicator for assessing slope stability. The coordinate transformation model is a mathematical model used to convert cluster displacement vectors in the image plane into physical spatial displacement. This model is built based on the spatial position parameters of the video acquisition device and can establish a mapping relationship between image coordinates and physical coordinates.
[0125] In some embodiments, calculating the physical spatial displacement of the deformed region based on the cluster displacement vector and coordinate transformation model includes: establishing the coordinate transformation model based on the spatial position parameters of the video acquisition device; inputting the cluster displacement vector into the coordinate transformation model to perform coordinate transformation calculation; and obtaining the physical spatial displacement of the deformed region based on the output result of the coordinate transformation calculation.
[0126] The spatial position parameters of video acquisition equipment refer to the parameters that describe the installation position and orientation of the video acquisition equipment in physical space. These mainly include the three-dimensional coordinates (Xc, Yc, Zc) of the equipment, the angle between the lens optical axis and the horizontal plane, the angle between the projection of the lens optical axis in the horizontal plane and the due north direction, and the lens focal length. These parameters are obtained through on-site measurement and are the basis for constructing the coordinate transformation model.
[0127] Image coordinates refer to a two-dimensional coordinate system used to locate the position of pixels in an image. Its origin is the upper left corner of the image, the x-axis is horizontal to the right, and the y-axis is vertical to the bottom. The unit is pixels. Physical coordinates refer to a three-dimensional coordinate system used to describe the position of the slope monitoring area in real physical space. In this embodiment, the geodetic coordinate system is used. The origin is the benchmark point preset by the water conservancy project. The x-axis points east, the y-axis points north, and the z-axis is vertical to the ground and upward. The unit is meters.
[0128] Coordinate transformation calculation refers to the process of inputting the pixel components (ΔX_pix, ΔY_pix) of the cluster displacement vector into the coordinate transformation model, and calculating the displacement components (ΔX_phy, ΔY_phy, ΔZ_phy) of the deformed region in physical coordinates through the model formula; the model output is the set of physical coordinate displacement components obtained after coordinate transformation calculation, which includes displacement values in three directions: x-axis (eastward), y-axis (northward), and z-axis (vertical), and can be directly used to represent the amount of physical spatial displacement.
[0129] The core purpose of this step is to convert the pixel-level cluster displacement vector obtained in step S4 into a physical spatial displacement that can be directly used for slope stability assessment. First, by measuring the spatial position parameters of the video acquisition device, the transformation relationship between image coordinates and physical coordinates is established to solve the problem that pixel displacement cannot directly correspond to physical deformation. Then, the cluster displacement vector is substituted into the model calculation to ensure the accuracy of the transformation process. Finally, the physical spatial displacement is output to provide the final quantitative result for slope monitoring.
[0130] The process of establishing a coordinate transformation model based on the spatial position parameters of the video acquisition device includes: first, obtaining the spatial position parameters of the video acquisition device, and then constructing a coordinate transformation model based on the principle of perspective projection, which is divided into two steps:
[0131] The first step is to use a total station to measure the three-dimensional coordinates (Xc, Yc, Zc) of the equipment. The total station is set up on the reference point of the geodetic coordinate system of the water conservancy project. The center point of the equipment mounting bracket is aimed and measured to obtain the three-dimensional coordinates of the equipment in physical space. The pitch angle α of the equipment is measured using an electronic inclinometer, and the azimuth angle β of the equipment is measured using a compass. The lens focal length f is obtained from the equipment parameter manual.
[0132] The second step is to construct a coordinate transformation model: based on the principle of perspective projection, establish the mapping relationship between image coordinates and physical coordinates. The core formula of the model is:
[0133]
[0134] Wherein, ΔX_phy is the eastward displacement in physical space (meters), ΔY_phy is the northward displacement (meters), and ΔZ_phy is the vertical displacement (meters); ΔX_pix and ΔY_pix are the pixel components of the cluster displacement vector; Z_phy is the average elevation of the deformation area in physical space (obtained through previous slope surveys); S_x and S_y are the pixel sizes of the image sensor; f is the lens focal length (meters); and α is the device pitch angle.
[0135] The parameters of the video acquisition equipment deployed in step S1 are as follows: The three-dimensional coordinates of the equipment were measured using a total station. The pitch angle was measured by an electronic inclinometer. The azimuth angle is measured by the compass. The device manual indicates that the lens focal length is f=8mm and the pixel size is... The average elevation of the deformation area, Z_phy, was found to be 180.00m in the preliminary survey. Substituting these parameters into the model formula, the coordinate transformation model for this project was obtained.
[0136] .
[0137] Perspective projection is a fundamental principle of optical imaging, accurately describing the geometric relationship between image coordinates and physical coordinates. The spatial position parameters of the video acquisition device directly determine the imaging angle and mapping scale; measuring these parameters with specialized instruments ensures the accuracy of the model input. Introducing the average elevation Z_phy of the deformed region corrects for differences in imaging scale across different elevation areas, avoiding conversion errors caused by variations in slope elevation. (Pixel size...) The introduction of this technology allows the number of pixels to be converted into actual physical size.
[0138] The cluster displacement vector is input into the coordinate transformation model for coordinate transformation calculation, including: first, extracting the cluster displacement vector of each final deformation region from the results of step S4, and obtaining its pixel components ΔX_pix (x-axis pixel displacement) and ΔY_pix (y-axis pixel displacement); then, substituting ΔX_pix and ΔY_pix into the corresponding formulas of the established coordinate transformation model to calculate the eastward displacement ΔX_phy, northward displacement ΔY_phy, and vertical displacement ΔZ_phy in physical space; during the calculation process, attention should be paid to the consistency of units, ensuring that all parameters are converted to international units such as meters and radians (if angle calculation is involved), to avoid unit confusion leading to errors; finally, the calculation results are preliminarily verified. If the absolute value of the displacement in a certain direction far exceeds the normal deformation range of the slope (e.g., ΔX_phy > 1 meter, short-term slope deformation is usually less than 0.1 meters), it is necessary to check whether there are errors in the cluster displacement vector or model parameters and exclude abnormal data.
[0139] The pixel components of the cluster displacement vector are the displacement quantization results in the image plane, which can only be converted into physical space displacement through a coordinate transformation model; unit consistency is the basis for ensuring calculation accuracy. Unifying the units of optical and physical parameters can avoid order-of-magnitude deviations caused by unit conversion errors; routine deformation range verification is a key link in data quality control. It can quickly identify input data errors and ensure the reliability of subsequent evaluations, because the short-term deformation of slopes is usually small when there are no extreme geological disasters (such as landslides), and results that far exceed this range are likely to be data anomalies.
[0140] The physical spatial displacement of the deformed region is obtained based on the output results of coordinate transformation calculations. This includes: first, integrating ΔX_phy, ΔY_phy, and ΔZ_phy obtained from coordinate transformation calculations to form a set of physical spatial displacement components for the deformed region; then, calculating the magnitude and direction of the physical spatial displacement. The formula for calculating the displacement magnitude is as follows: (Unit: meters) This value reflects the total displacement of the deformed area; the displacement direction is determined by the proportion of the eastward, northward, and vertical components. The ratio of each component to the total displacement can be calculated to describe the main direction of displacement. For example, if ΔX_phy has the largest proportion, it indicates that the displacement is mainly eastward; finally, the "deformed area number - eastward displacement - northward displacement - vertical displacement - total displacement magnitude - main displacement direction" are compiled into a structured report to form the final physical space displacement record of the deformed area.
[0141] The physical displacement components of the corrected deformation region 1 are rice, rice, meters; calculate the total displacement as Meters; Calculate the proportion of each component: the proportion of eastward is 0.077 / 0.148≈52%, the proportion of northward is 0.122 / 0.148≈82%, and the proportion of vertical is 0.033 / 0.148≈22%. Therefore, the main displacement direction is northeast. The final structured report records deformation area 1 as follows: 0.077m eastward, 0.122m northward, 0.033m vertically, total displacement 0.148m, with the main direction being northeast.
[0142] The physical space displacement components can reflect the deformation in different directions. For example, vertical displacement reflects slope settlement or heave, while horizontal displacement reflects slope sliding. The total displacement magnitude and main direction can provide the overall deformation degree and trend, meeting the different needs of slope stability assessment.
[0143] This step addresses the problem that traditional video monitoring can only acquire pixel displacements and cannot be converted into physical displacements, thus making it unsuitable for direct slope stability assessment. By constructing a coordinate transformation model based on the principle of perspective projection and the spatial position parameters of the equipment, it solves the problem of large transformation errors caused by ignoring slope elevation differences. Furthermore, by calculating the three-dimensional physical displacement components and the total displacement, it addresses the issue that it can only monitor horizontal deformation and cannot fully reflect the three-dimensional deformation state of the slope, thus providing a more comprehensive quantitative basis for slope stability assessment.
[0144] like Figure 2 The diagram shown is a functional block diagram of a video stream-based slope deformation monitoring system for hydraulic engineering provided in an embodiment of this application.
[0145] The video stream-based slope deformation monitoring system 100 for hydraulic engineering described in this application can be installed in an electronic device. Depending on the functions implemented, the video stream-based slope deformation monitoring system 100 may include a video acquisition and monitoring module 101, an initial feature point extraction module 102, a pixel displacement calculation module 103, a deformation area identification module 104, and a physical displacement conversion module 105. The module described in this application can also be referred to as a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0146] In this embodiment, the functions of each module / unit are as follows:
[0147] The video acquisition and monitoring module 101 is used to deploy the video acquisition device in the stable area of the slope and acquire video stream data containing the slope monitoring area continuously captured by the video acquisition device.
[0148] The initial feature point extraction module 102 is used to extract multiple natural feature points from the initial frame of the video stream data;
[0149] The pixel displacement calculation module 103 is used to calculate the pixel displacement vector of each natural feature point based on the position change of the natural feature point in subsequent video frames.
[0150] The deformation region identification module 104 is used to identify natural feature points with consistent motion trends as the same deformation region based on the spatial distribution characteristics of the pixel displacement vector, and generate cluster displacement vectors.
[0151] The physical displacement conversion module 105 is used to calculate the physical spatial displacement of the deformed region based on the cluster displacement vector and coordinate conversion model.
[0152] In the embodiments provided in this application, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0153] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0154] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0155] It is obvious to those in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application.
[0156] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, it should be understood by those skilled in the art that modifications or equivalent substitutions can be made to the technical solutions of this application without departing from the spirit and scope of the technical solutions of this application.
Claims
1. A hydraulic engineering slope deformation monitoring method based on video stream, characterized in that, The method comprises: deploying a video acquisition device in a slope stability area, and acquiring video stream data containing a slope monitoring area continuously shot by the video acquisition device; extracting a plurality of natural feature points from an initial frame of the video stream data; calculating a pixel displacement vector of each natural feature point based on a position change of the natural feature point in a subsequent video frame; identifying natural feature points with consistent motion trends as the same deformation region based on the spatial distribution characteristics of the pixel displacement vector, and generating a cluster displacement vector; calculating a physical space displacement amount of the deformation region according to the cluster displacement vector and a coordinate conversion model.
2. The method for monitoring the deformation of the water conservancy slope based on the video stream according to claim 1, wherein, The method comprises: fixing the video acquisition device on a stable bedrock directly opposite the slope monitoring area; adjusting the shooting parameters of the video acquisition device to ensure covering the entire slope monitoring area; starting the video acquisition device for continuous image acquisition; acquiring the video stream data containing time sequence information output by the video acquisition device. 3.The method of claim 1, wherein, The method comprises: converting the initial frame of the video stream data into a gray-scale image; performing corner detection on the gray-scale image to obtain a plurality of initial feature points; selecting the natural feature points with stable optical characteristics from the initial feature points based on the texture richness of the initial feature points. 4.The method of claim 1, wherein, The method comprises: selecting a current video frame adjacent to the initial frame in the video stream data; tracking the position of the natural feature point in the current video frame to establish a corresponding relationship of the natural feature point; calculating the coordinate difference value of the natural feature point between the initial frame and the current video frame; generating the pixel displacement vector of each natural feature point based on the coordinate difference value. 5.The method of claim 1, wherein, The method comprises: analyzing the direction consistency of the pixel displacement vector to form a displacement vector set; dividing the natural feature points with consistent motion trends into the same deformation region based on the direction similarity of each pixel displacement vector in the displacement vector set; performing synthetic calculation on the pixel displacement vectors of all natural feature points in the deformation region; generating a cluster displacement vector representing the overall motion trend of the deformation region based on the synthetic calculation result.
6. The method for hydraulic engineering slope deformation monitoring based on video stream according to claim 5, characterized in that, The method comprises: calculating the direction angle of each pixel displacement vector in the displacement vector set to form a direction angle set; constructing a direction similarity matrix based on the direction angle set; constructing a graph model taking the natural feature points as nodes and the direction similarity and space-time proximity as edges; Based on the graph model, a graph theory-based community discovery algorithm is used to cluster the natural feature points to form an initial deformation region, wherein the weight of an edge is determined by a weighted function of the direction similarity and the spatial distance; Based on the spatial distribution continuity of the natural feature points in the initial deformation region, a final deformation region is confirmed.
7. The method for hydraulic engineering slope deformation monitoring based on video stream according to claim 6, characterized in that, The confirmation of the final deformation region based on the spatial distribution continuity of the natural feature points in the initial deformation region includes: Calculate the spatial distance between each natural feature point in the initial deformation region to generate a distance matrix; Based on the distance matrix and a preset spatial distance threshold, it is judged whether the natural feature points form a connected region; The connected region that meets the spatial distribution continuity condition is confirmed as the final deformation region. 8.The method of claim 5, wherein, The generation of the cluster displacement vector representing the overall motion trend of the deformation region based on the synthesis calculation result includes: Based on the texture richness of each natural feature point in the deformation region, a corresponding weighting coefficient is calculated; The pixel displacement vector of each natural feature point in the deformation region is weighted and synthesized using the weighting coefficient; Based on the weighted synthesis result, the cluster displacement vector representing the overall motion trend of the deformation region is generated. 9.The water conservancy project slope deformation monitoring method based on video stream of claim 1, wherein, The calculation of the physical space displacement amount of the deformation region according to the cluster displacement vector and a coordinate conversion model includes: Based on the spatial position parameters of the video acquisition device, the coordinate conversion model is established; The cluster displacement vector is input into the coordinate conversion model for coordinate conversion calculation; Based on the output result of the coordinate conversion calculation, the physical space displacement amount of the deformation region is obtained.
10. A video stream-based hydraulic engineering slope deformation monitoring system for implementing the video stream-based hydraulic engineering slope deformation monitoring method of any one of claims 1-9, characterized in that, The system includes: A video acquisition and monitoring module is used to deploy a video acquisition device in a slope stability region to obtain video stream data containing a slope monitoring region continuously shot by the video acquisition device; An initial feature point extraction module is used to extract a plurality of natural feature points from the initial frame of the video stream data; A pixel displacement calculation module is used to calculate the pixel displacement vector of each natural feature point based on the position change of the natural feature point in the subsequent video frame; A deformation region identification module is used to identify natural feature points with consistent motion trends as the same deformation region based on the spatial distribution characteristics of the pixel displacement vector to generate a cluster displacement vector; A physical displacement conversion module is used to calculate the physical space displacement amount of the deformation region according to the cluster displacement vector and a coordinate conversion model.
Citation Information
Cited By
Evaluation model construction method for road slope disaster prevention and reduction
CN121982456A