Unmanned aerial vehicle video stream analysis method based on high-frequency digital information analysis
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INNER MONGOLIA POLICE COLLEGE
- Filing Date
- 2025-02-08
- Publication Date
- 2026-05-12
Smart Images

Figure CN120071198B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of unmanned aerial vehicle video stream analysis and geological disaster monitoring, and particularly relates to an unmanned aerial vehicle video stream analysis method based on high-frequency digital information analysis. BACKGROUND
[0002] Unmanned aerial vehicle video stream analysis is of great significance in the field of unmanned aerial vehicle monitoring. It uses high-frequency video stream acquisition and intelligent analysis technology to process dynamic change scenes in real time, especially in the fields of geological disaster monitoring, emergency response, and environmental protection, providing high-precision data support. Through high-frequency digital information analysis, abnormal changes can be identified more quickly, and an efficient early warning mechanism can be provided.
[0003] In the prior art (Chinese invention patent, publication number: CN117854256B, name: Geological disaster monitoring method based on unmanned aerial vehicle video stream analysis), unmanned aerial vehicles are used to collect video streams and monitor geological disasters through optical flow calculation methods. Its main technical means include real-time collection of image frames, grayscale processing, optical flow calculation, and motion amplitude distribution map generation. However, its technical defects mainly manifest as:
[0004] In the prior art, optical flow calculation is performed by processing grayscale images frame by frame, which is difficult to capture high-frequency change characteristics in complex environments, especially in areas with complex crack expansion or humidity anomalies, and optical flow calculation has a precision bottleneck; the prior art is mainly based on single video stream data and cannot fuse humidity, thermal radiation, and other modal characteristics, resulting in insufficient description of complex geological disasters; the prior art lacks a real-time path adjustment strategy based on risk priority, making it difficult to optimize the allocation of unmanned aerial vehicle monitoring tasks, leading to resource waste or monitoring blind spots. SUMMARY
[0005] To address the above-mentioned problems in the prior art, the present application provides an unmanned aerial vehicle video stream analysis method based on high-frequency digital information analysis. Based on multi-modal data fusion and high-frequency digital information analysis, the present application constructs a set of geological disaster monitoring methods that integrate dynamic risk heat map generation, time series modeling, and reinforcement learning path planning. By fusing humidity, thermal radiation, and geometric features, the risk priority is dynamically calculated and the unmanned aerial vehicle path planning is optimized to achieve accurate monitoring and risk warning of the target area. The present application significantly improves the real-time performance, flexibility, and accuracy of unmanned aerial vehicle monitoring, and can quickly mark high-risk areas in complex environments, providing an efficient solution for geological disaster warning and response.
[0006] An unmanned aerial vehicle video stream analysis method based on high-frequency digital information analysis, comprising the following steps:
[0007] Based on the dynamic partition strategy, the unmanned aerial vehicle is used to collect multi-modal data in the target area, the multi-modal data including multi-spectral video frame data and point cloud depth data, and data denoising and spatial alignment processing are performed to generate preliminary three-dimensional model data;
[0008] Based on the preliminary three-dimensional model data, a three-dimensional modeling algorithm is used to generate a low-resolution volume data model, and a high-resolution volume data model is generated by refining the key areas; multi-modal data fusion technology is used to combine the multi-spectral video frame data and the point cloud depth data, map the material characteristics of the target area to the three-dimensional model, and extract high-frequency feature data and low-frequency contour data of the target area through a feature extraction algorithm;
[0009] Based on the high-resolution volume data model and the high-frequency feature data, a causal network is constructed, and a causal graph node is defined; a causal reasoning method is used to quantify the causal relationship between nodes to generate causal contribution data; time series modeling is used to model dynamic feature data to predict the future trend of the target area and generate extended prediction data;
[0010] The extended prediction data is input, and a dynamic risk heat map is generated through a feature fusion algorithm; a generative adversarial network is used to optimize the risk labeling of the target area, and a reinforcement learning path planning algorithm is used to optimize the flight path of the unmanned aerial vehicle in combination with the dynamic risk heat map to preferentially monitor the high-risk part of the target area; the newly added multi-modal acquisition data is fed back to dynamically update the volume data model and the risk heat map, realizing a monitoring closed loop.
[0011] Preferably, the multi-spectral video frame data in the multi-modal data includes a visible light band for capturing surface texture features of the target area, a near-infrared band for monitoring humidity changes of soil and vegetation, and a short-wave infrared band for detecting local anomalies of ground thermal radiation intensity.
[0012] Preferably, the multi-spectral video frame data in the multi-modal data is denoised by an adaptive filter, which dynamically adjusts the filter parameters according to the band characteristics of the multi-spectral video frame data to eliminate noise caused by changes in light; the point cloud depth data is cleared of abnormal points by statistical filtering, which removes depth data points deviating from the expected range based on the local distribution characteristics of depth information.
[0013] Preferably, the three-dimensional modeling algorithm includes generating a low-resolution volume data model using a neural radiance field model, which distributes the geometry and spectral information of the preliminary three-dimensional model data through ray sampling to generate a low-resolution three-dimensional representation; based on the low-resolution volume data model, a deep super-resolution network is used to refine the cracks and settlement areas in the key areas, which reconstructs the detail information layer by layer through convolution layers to generate a high-resolution volume data model.
[0014] Preferably, the multi-modal data fusion technology comprises mapping the humidity features and thermal radiation features of the multi-spectral video frame data into the three-dimensional model generated by the point cloud depth data using a non-linear interpolation method, the non-linear interpolation method associates the material features with the geometric features according to the coordinate relationship of the multi-spectral data and the point cloud depth data in the three-dimensional space, and generates a multi-dimensional material three-dimensional model with humidity gradient and thermal radiation intensity information.
[0015] Preferably, the feature extraction algorithm comprises performing two-dimensional fast Fourier transform on the humidity and thermal radiation features of the high-resolution volume data model, extracting high-frequency components for capturing the details of crack propagation and local settlement in the target area; and further processing the extracted high-frequency components using continuous wavelet transform, the wavelet transform refines the spatial representation of the crack edge and settlement features by decomposing the local changes of the high-frequency components.
[0016] Preferably, the causal network constructs a causal graph based on the humidity changes, thermal radiation intensity, crack propagation amplitude and ground settlement in the high-resolution volume data model; the causal relationship of the causal graph is quantified by a Bayesian network inference method, the Bayesian network calculates the causal influence weight between nodes according to the prior probability and conditional probability of each node, and generates causal contribution data.
[0017] Preferably, the time series modeling models the dynamic feature data through a time convolution network, the convolution kernel of the time convolution network is dynamically adjusted according to the length of the time series and the feature change rate, and the time-dependent relationship between the humidity change, crack propagation and thermal radiation intensity is captured; the output result of the time convolution network is used to predict the future trend of the target area, and generate extension prediction data.
[0018] Preferably, the dynamic risk heat map fuses the humidity features, thermal radiation features and geometric features from the high-resolution volume data model through a multi-modal interactive attention mechanism; the interactive attention mechanism dynamically adjusts the contribution of the humidity features and the thermal radiation features in the high-risk area labeling by calculating the weight matrix between the modal features, and generates a dynamic risk heat map with enhanced resolution and highlighted focus.
[0019] Preferably, the reinforcement learning path planning algorithm calculates the optimal flight path based on the high-risk area priority of the dynamic risk heat map, in combination with the current position and flight parameters of the unmanned aerial vehicle; the path planning algorithm optimizes the path covering the high-risk area through a reward function, adjusts the flight task of the unmanned aerial vehicle in real time, and generates task feedback data for updating the dynamic risk heat map.
[0020] Compared with the prior art, the advantages and beneficial effects of the present application are that:
[0021] The application realizes accurate labeling of high-risk areas by fusing humidity, thermal radiation and geometric features through a multi-modal interactive attention mechanism, and overcomes the problem of insufficient description of complex geological disasters by single modal data in the prior art.
[0022] The application realizes path optimization of efficient coverage of high-risk areas by dynamically adjusting the flight task of the unmanned aerial vehicle through the reinforcement learning path planning algorithm, and solves the defect of insufficient flexibility of path planning in the prior art.
[0023] The application realizes accurate early warning of geological disasters by modeling dynamic features through the time convolution network, capturing the time dependence of humidity change, crack expansion and thermal radiation intensity, and generating extended prediction data, which makes up for the short board of insufficient analysis of time series features in the prior art.
[0024] The application significantly improves the resolution and labeling accuracy of the heat map by optimizing the dynamic risk heat map through the generative adversarial network, and enhances the visualization effect of the high-risk area. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 It is a flowchart of the method of the application;
[0026] Figure 2 It is a multi-modal data fusion process schematic diagram in the application;
[0027] Figure 3 It is a dynamic risk heat map generation process schematic diagram in the application;
[0028] Figure 4 It is a schematic diagram of the unmanned aerial vehicle path optimization process in the application;
[0029] Figure 5 It is a schematic diagram of the monitoring closed-loop updating process in the application. DETAILED DESCRIPTION
[0030] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, many specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure.
[0031] As shown in Figure 1 A method for analyzing unmanned aerial vehicle video streams based on high-frequency digital information analysis includes the following steps:
[0032] Based on a dynamic partitioning strategy, multi-modal data is collected in the target area by the unmanned aerial vehicle, the multi-modal data includes multi-spectral video frame data and point cloud depth data, and data denoising and spatial alignment processing are performed to generate preliminary three-dimensional model data;
[0033] In the present application, the core principle of the dynamic partitioning strategy is to divide the target area into multiple dynamically adjusted sub-regions to achieve efficient task allocation and coverage optimization of the UAVs. This strategy uses the Voronoi partitioning method to adjust the partitioning results in real time according to the current position of the UAVs and the terrain characteristics of the target area, ensuring that the UAV tasks are evenly distributed in space and avoiding monitoring blind spots and data redundancy. Each sub-region corresponds to the task range of a UAV, and the partitioning results are updated by a dynamic partitioning algorithm to adapt to real-time changes in environmental conditions such as changes in UAV power, obstacles, or risk levels of the target area.
[0034] The UAVs collect multi-modal data within their respective partitioned ranges, including multi-spectral video frame data and point cloud depth data. The multi-spectral video frame data includes visible light bands for capturing surface texture, near-infrared bands for monitoring humidity changes, and short-wave infrared bands for detecting local anomalies in thermal radiation intensity; the point cloud depth data is collected by a laser radar to provide three-dimensional geometric information of the terrain. Through a time synchronization mechanism, the multi-modal data is time-aligned frame by frame to ensure the spatial consistency of the multi-spectral frames and the point cloud depth frames.
[0035] The collected multi-modal data needs to be processed for data denoising and spatial alignment to improve the reliability and fusion accuracy of the data. Data denoising uses an adaptive filtering method to remove light interference and random noise for different waveband characteristics of the multi-spectral frame data by dynamically adjusting the filtering parameters. Point cloud depth data is filtered to remove outliers by statistical filtering, and noise points outside the threshold range are removed according to the local distribution characteristics of the depth information. Spatial alignment uses a global registration method based on the Iterative Closest Point (ICP) algorithm to align the multi-spectral frames and point cloud depth data in three-dimensional space, ensuring that multi-modal information is fused in the same spatial coordinate system.
[0036] Through the dynamic partitioning strategy, the UAVs can efficiently cover the target area, maximizing the use of monitoring resources, especially in the case of a large and complex terrain target area. This strategy effectively avoids repetition or omission of UAV tasks, improving the efficiency and completeness of data collection. Multi-modal data collection combines spectral information and three-dimensional geometric information to provide comprehensive data support for subsequent modeling and analysis. Data denoising and spatial alignment processing significantly improve the quality of multi-modal data, ensuring the accuracy of subsequent volume data modeling and feature extraction.
[0037] Preferably, the multi-spectral video frame data in the multi-modal data includes visible light bands for capturing surface texture features of the target area, near-infrared bands for monitoring humidity changes in soil and vegetation, and short-wave infrared bands for detecting local anomalies in surface thermal radiation intensity.
[0038] The visible light band is mainly used to capture the surface texture features of the target area. Its principle is to generate image frames by receiving the visible light signals reflected by the ground surface through a multispectral camera. Since the intensity of visible light reflection is closely related to the surface material (such as soil, rock, vegetation) and terrain features (such as cracks, subsidence), the frame data of this band can clearly show the geometric texture and details of the surface. In the multispectral camera carried by the UAV, an RGB (red, green, blue) three-channel sensor is used to collect visible light band data. The resolution of each frame is set to be higher than 1 cm / pixel, so that high-precision texture features can be extracted in subsequent processing.
[0039] The near-infrared band is used to monitor the humidity changes of soil and vegetation. Its principle is based on the absorption characteristics of near-infrared light in water molecules and plant leaves. By analyzing the strength changes of the band signal, the humidity distribution of the target area can be evaluated. The high penetration characteristics of the near-infrared band enable it to detect humidity changes below the ground surface to some extent. The original data collected by the near-infrared band is calculated by spectral reflectance analysis to calculate the humidity change. The formula of spectral reflectance R(λ) is:
[0040]
[0041] where R(λ) is the reflectivity at wavelength λ, I r (λ) is the reflected light intensity received by the sensor, I i (λ) is the incident light intensity. The decrease in reflectivity usually corresponds to an increase in humidity. Combined with the humidity distribution chart, potential landslide risk areas can be marked.
[0042] The short-wave infrared band is used to detect local anomalies in the intensity of surface thermal radiation. Its principle is to analyze the surface temperature distribution through thermal radiation characteristics. The short-wave infrared band is extremely sensitive to changes in surface thermal signals and can be used to identify local temperature increases in soil and rock caused by sliding and stress concentration.
[0043] Specific implementation: Data collection in the short-wave infrared band is carried out through infrared sensors, and the data contains the temperature distribution of the target area. The thermal radiation intensity EEE of each pixel point is calculated according to the simplified formula of Planck's law:
[0044] E = ∈·σ·T 4
[0045] where E is the radiation intensity, ∈ is the emissivity of the surface, σ is the Stefan-Boltzmann constant, and T is the absolute temperature. By comparing the temperature anomaly points in the region, high-risk landslide precursor areas can be marked.
[0046] The multi-band combination of multispectral video frame data provides comprehensive support for high-frequency digital information analysis of target areas. The visible light band makes the texture features of the ground surface (such as cracks and subsidence) clear, the near-infrared band enhances the monitoring ability of the ground surface humidity distribution, and the short-wave infrared band is used to identify thermal radiation anomalies of landslide precursors. Through the fusion of three-band data, the potential risk points of the target area can be accurately located, providing multi-dimensional support for subsequent three-dimensional modeling and causal network analysis.
[0047] In an embodiment, in a landslide high-risk mountainous area, three unmanned aerial vehicles equipped with multispectral cameras are arranged in the target area to collect visible light, near-infrared, and short-wave infrared band data from different angles. The flight height of the unmanned aerial vehicle is set to 100 meters, and the camera synchronously collects three-band data at a frame rate of 30 Hz.
[0048] The collected RGB images clearly show the shape and distribution of surface cracks. Through image enhancement algorithms, the crack edges are refined, and the trend of crack width change is quantified, providing a basis for preliminary labeling of high-risk points.
[0049] In the soil and vegetation coverage area of the target area, analysis of spectral reflectance finds that the humidity in some areas is abnormally high. These areas are further associated with crack distribution, indicating that humidity anomalies may cause local subsidence.
[0050] In the landslide precursor area, thermal radiation analysis finds that the temperature of some points is significantly higher than that of the surrounding area. Combined with crack distribution and humidity data, it is confirmed that these thermal anomaly points are stress concentration areas induced by landslides.
[0051] The three-band data are integrated to generate a multi-modal risk distribution map, and the cracks, humidity anomaly points, and high-temperature points of thermal radiation in the target area are accurately labeled, providing accurate data input for subsequent dynamic modeling and path optimization.
[0052] Preferably, the multispectral video frame data in the multi-modal data is denoised by an adaptive filter, which dynamically adjusts the filter parameters according to the band characteristics of the multispectral video frame data to eliminate noise caused by changes in light; the point cloud depth data is cleared of abnormal points by statistical filtering, which removes depth data points that deviate from the expected range based on the local distribution characteristics of depth information.
[0053] The multispectral video frame data may be affected by changes in light (such as fluctuations in natural light intensity and local shadows) and sensor noise during the collection process, which may cause a decrease in image quality. The adaptive filter dynamically adjusts the filter parameters according to the band characteristics of the multispectral data and the local light environment, adaptively enhances the effective signal, and suppresses the invalid noise. The adaptive filter calculates the statistical characteristics of the image based on a local window and dynamically adjusts the filter weights. The filter weight calculation formula is:
[0054]
[0055] Where W(x,y) represents the filtering weight of pixel (x,y). For the local variance of the signal, This represents the variance estimated for noise. This formula ensures that regions with drastic signal changes (such as edges or texture features) retain more detail, while flat regions suppress noise more effectively. Spectral data in different bands (visible, near-infrared, and short-wave infrared) have different signal characteristics; therefore, filter parameters need to be calculated independently based on band characteristics. For example, the short-wave infrared band is sensitive to thermal radiation, and its noise primarily originates from photothermal interference; the filter will preferentially smooth noise signals in low-frequency regions.
[0056] Point cloud depth data may contain outliers during acquisition due to laser scattering, terrain occlusion, or sensor errors. These outliers can affect the accuracy of terrain modeling. Statistical filtering improves the quality of point cloud data by analyzing the local distribution characteristics of the depth data and removing outliers that significantly deviate from their surroundings. Within the neighborhood of each depth point, the mean μ and standard deviation σ of the depth values are calculated, and outliers meeting the following criteria are removed:
[0057] |D i -μ|>k·σ
[0058] Among them, D i is the depth value of the current depth point, and k is the filtering threshold coefficient, used to control the strictness of outlier removal. Local distribution characteristics are quickly calculated by building a neighborhood search tree in the point cloud data to reduce processing time. In point cloud data with complex terrain (such as landslide areas), outliers may be concentrated in edge areas or vegetation-covered areas. Statistical filtering can effectively remove these outliers, ensuring the geometric accuracy of subsequent 3D modeling.
[0059] Adaptive filtering of multispectral video frame data significantly improves the signal-to-noise ratio of images in different bands, especially in environments with drastic changes in light, preserving more detailed textures and regional features, making the annotation of humidity gradients and thermal radiation anomalies more accurate.
[0060] Statistical filtering of point cloud depth data eliminates depth deviations caused by noise in lidar acquisition, significantly improving the spatial consistency of point cloud data and laying a high-quality data foundation for 3D modeling and multimodal data fusion.
[0061] In an example, in a high-risk landslide area, multispectral video frame data and point cloud depth data were collected by drone, and denoising was performed by applying adaptive filters and statistical filters respectively.
[0062] Multispectral video frame data was acquired using an onboard multispectral camera, including visible, near-infrared, and short-wave infrared bands. The flight altitude was 120 meters, and the frame rate was 30Hz. Point cloud depth data was acquired by a lidar system, generating a point cloud density of 1000 points per square meter.
[0063] In the shortwave infrared band, signal fluctuations are significant due to surface thermal radiation and light interference. The adaptive filter dynamically adjusts weights based on local signal changes, filtering out noise signals in the low-frequency region while enhancing the detail and texture of cracks and areas with abnormal humidity. In vegetated areas, there are many depth anomalies caused by laser reflection. Statistical filtering removes points that deviate from the local mean by three times the standard deviation, and these points are usually inconsistent with the surrounding terrain.
[0064] After adaptive filtering, the multispectral video frame data shows clearer texture details in areas of thermal radiation anomalies and more distinct boundaries of humidity changes. The statistically filtered point cloud depth data exhibits greater coherence in edge regions and more accurate terrain geometry, providing high-quality input for subsequent modeling.
[0065] Based on the preliminary 3D model data, a low-resolution volumetric data model is generated using 3D modeling algorithms, and a high-resolution volumetric data model is generated by refining key areas. Multi-modal data fusion technology is used to combine multispectral video frame data with point cloud depth data to map the material features of the target area into the 3D model. High-frequency feature data and low-frequency contour data of the target area are extracted using feature extraction algorithms.
[0066] The preliminary 3D model data is the product of denoising and registration processing of multimodal data collected by the UAV, containing the geometric structure and basic material features of the target area. The 3D modeling algorithm employs the Neural Radiation Field (NeRF) model, which reconstructs a low-resolution volumetric data model of the region by sampling the volume distribution of each ray in space. The core of the NeRF model lies in integrating the color and density of each ray r(t):
[0067]
[0068] Where C(r) represents the color of the light, σ(r(t)) represents the volume density, c(r(t)) is the color, T(t) is the cumulative transmittance of the light, and t is the cumulative transmittance of the light. n and t f These represent the starting and ending points of light rays. This model can quickly generate low-resolution 3D representations, providing the basic structure of the global region.
[0069] Based on the low-resolution volumetric data model, the marked key regions (such as cracks and subsidence) are refined. Refinement employs a deep super-resolution network (SR-Net), which enhances the features of the regions layer by layer through multiple convolutions to restore details. Each convolution operation is based on the following formula:
[0070] f l =ReLU(W l ·f l-1 +b l )
[0071] Among them, f l W represents the feature output of the l-th layer. l and b l Here, represents the weights and biases, respectively, and ReLU is the activation function. Through multiple feature extractions and refinements, a 3D model with high spatial resolution is generated.
[0072] To enhance the material properties of the 3D model, multimodal data fusion technology was used to map humidity and thermal radiation features from multispectral video frame data into the 3D model. Nonlinear interpolation technology was then used to fuse multispectral frame data with point cloud depth data through 3D coordinate relationships, achieving a precise correlation between geometric and material information.
[0073] Feature decomposition is performed on the high-resolution volumetric data model to extract high-frequency feature data (such as crack edges and minor changes in subsidence areas) and low-frequency contour data (such as overall topographic changes) of the target region. High-frequency features are extracted using a two-dimensional fast Fourier transform (2D-FFT), with the following formula:
[0074]
[0075] Here, F(u,v) represents the frequency domain signal, f(x,y) is the spatial signal, and M and N are the spatial dimensions. Texture details are extracted using high-frequency components in the frequency domain, and edge features are refined using continuous wavelet transform. The low-frequency contour preserves the overall shape and variation trend of the target region and is used for global modeling.
[0076] Through 3D modeling algorithms and multimodal data fusion technology, the geometric and material information of the target area was represented with high precision, and the high-resolution volumetric data model provided a solid foundation for subsequent analysis. High-frequency feature extraction enhanced the detailed representation of the target area, especially in landslide monitoring, where cracks, settlement, and local anomalies were accurately labeled, while low-frequency contour data provided overall trend information, facilitating subsequent dynamic expansion and prediction.
[0077] In one example, in a mountainous area, the initial 3D model data of the target region was generated from multimodal data collected by a drone. Cracks and subsidence within the region were the key monitoring targets, requiring high-resolution modeling.
[0078] A neural radiation field model was used to globally model the preliminary 3D model data, generating a low-resolution volumetric data model. The model shows the basic topography of the region, including major ridges, slopes, and depressions.
[0079] For cracks and settlement areas, a deep super-resolution network is used to enhance features layer by layer to generate a high-resolution volumetric data model. After processing, the width variations of cracks and the boundaries of settlement areas are clearly visible.
[0080] The humidity gradient and thermal radiation intensity of the multispectral frame data were mapped into a 3D model using nonlinear interpolation. The crack region showed humidity anomalies, while the thermal radiation intensity of the subsidence area was higher than the surrounding area.
[0081] A two-dimensional fast Fourier transform is performed on the high-resolution volumetric data model to extract high-frequency features of crack edges and settlement areas, and continuous wavelet transform is used to further refine the edge details. Low-frequency contour data is extracted globally to show the overall trend of the region.
[0082] The results of 3D modeling and feature extraction clearly show the distribution and propagation direction of cracks, and the humidity and thermal anomalies in the settlement area are accurately marked, providing high-precision input data for dynamic risk assessment and propagation prediction.
[0083] Preferably, the 3D modeling algorithm includes generating a low-resolution volumetric data model using a neural radiation field model. The neural radiation field model performs distributed encoding of the geometric and spectral information of the preliminary 3D model data through ray sampling to generate a low-resolution 3D representation. Based on the low-resolution volumetric data model, a deep super-resolution network is used to refine the cracks and subsidence areas in key regions. The deep super-resolution network reconstructs detailed information layer by layer through convolutional layers to generate a high-resolution volumetric data model.
[0084] This invention uses the Neural Radiation Field (NeRF) model to generate a low-resolution volumetric data model, and then uses a deep super-resolution network to refine key parts of the target area (such as cracks and settlement) to finally generate a high-resolution volumetric data model, laying the foundation for subsequent high-frequency digital information analysis.
[0085] The neural radiation field model is an algorithm that uses ray sampling to distribute and encode spectral and geometric information in three-dimensional space. In this invention, the initial three-dimensional model data is mapped to a continuous voxel distribution in space. The neural radiation field generates a low-resolution three-dimensional representation of the target region by sampling the color and density of ray r(t). The ray integral formula is as follows:
[0086]
[0087] C(r) represents the color of light, indicating the light emitted from t. n (Light source) to t f Color accumulation at (the endpoint of the ray); σ(r(t)) represents the voxel density, describing the geometric information of a point in space; c(r(t)) represents the color value of the ray at position r(t); T(t) represents the cumulative transmittance of the ray, indicating the color value from t... n The transmittance to t.
[0088] The neural radiation field model uses a multilayer perceptron (MLP) to fit σ(r(t)) and c(r(t)) to generate a low-resolution three-dimensional volume data model, providing a global representation of the target region.
[0089] Based on the low-resolution volumetric data model, key areas such as cracks and settlement are refined by employing a deep super-resolution network to enhance detailed features layer by layer. This network extracts, refines, and reconstructs the input 3D features through convolutional operations. Each convolutional layer calculates its output features using the following formula:
[0090] f l =ReLU(W l ·f l-1 +b l )
[0091] f l W represents the output feature of the l-th layer; l f represents the weight matrix; l-1 Indicates the features of the previous layer; b l This represents the bias term. The final layer of the network outputs a high-resolution 3D model, where the texture of the crack edges and the geometric details of the settlement area are significantly enhanced.
[0092] The neural radiation field model rapidly generates a low-resolution volumetric data model, fully revealing the basic geometric features of the target region and providing a reliable global framework for subsequent refinement. The deep super-resolution network refines details in crack and subsidence areas, improving the model's spatial resolution and enabling clear representation of high-frequency features (such as discontinuous changes at crack edges and local protrusions in subsidence areas). The high-resolution volumetric data model provides high-precision input data for subsequent high-frequency feature extraction and dynamic expansion prediction, significantly improving the accuracy and reliability of the entire approach.
[0093] In an example, in a high-risk landslide area, preliminary 3D model data collected by a drone was used to generate a high-resolution volumetric data model by combining a neural radiation field model and a deep super-resolution network.
[0094] The preliminary 3D model data contains multimodal information of the target region, including point cloud depth data and multispectral video frame data. After denoising and registration processing, it is input into the neural radiation field model.
[0095] The neural radiation field model uses light sampling to distribute and encode the geometric and spectral information of the target region, generating a low-resolution volumetric data model. This model displays the global terrain features of the target region, such as the slope of hills and the distribution of depressions.
[0096] Based on the low-resolution model, cracks and subsidence areas are selected as the focus of refinement. The deep super-resolution network enhances the features of these areas layer by layer, and the texture variations of cracks and the boundary details of subsidence are clearly expressed in the high-resolution model.
[0097] The generated high-resolution volumetric data model shows the width, depth, and direction of the cracks, while providing a detailed description of the boundaries and internal morphology of the settlement area, thus providing high-precision input data for landslide risk assessment.
[0098] Preferred, such as Figure 2 As shown, the multimodal data fusion technology includes using a nonlinear interpolation method to map the humidity features and thermal radiation features of multispectral video frame data to a three-dimensional model generated from point cloud depth data. The nonlinear interpolation method associates material features with geometric features based on the coordinate relationship between multispectral data and point cloud depth data in three-dimensional space, generating a multidimensional material three-dimensional model with humidity gradient and thermal radiation intensity information.
[0099] The multimodal data fusion technology in this invention uses a nonlinear interpolation method to map the humidity and thermal radiation characteristics of multispectral video frame data onto a 3D model generated from point cloud depth data, thus producing a multidimensional material 3D model with a tight integration of material and geometric features. This model not only provides the geometric shape of the target area but also encompasses the spatial distribution information of humidity gradient and thermal radiation intensity, thereby providing comprehensive data support for high-frequency digital information analysis.
[0100] Nonlinear interpolation methods model the coordinate relationship between multispectral video frame data and point cloud depth data in 3D space, accurately projecting material features (humidity and thermal radiation features) onto the 3D geometric model represented by the point cloud data. Compared to linear interpolation methods, nonlinear interpolation can more accurately handle irregularly distributed data points while preserving detailed features. Nonlinear interpolation uses a weighted inverse distance algorithm to map material features, and the interpolation function can be expressed as:
[0101]
[0102] F(p) represents the interpolation result at point p, and f(p) represents the material characteristic value at that point; i ) represents the input data point p i Material characteristic value; w i Represents the weight, from point p to p i The distance d(p,p) i The distance attenuation is determined by k; k represents the parameter controlling the distance attenuation, usually taken as k=2 to balance the weighted attenuation rate. Through the above interpolation calculation, the material features in the multispectral data are mapped point by point to the three-dimensional coordinates generated from the point cloud data.
[0103] Point cloud depth data provides geometric morphology information in three-dimensional space, while multispectral video frame data contains humidity and thermal radiation characteristics. Through nonlinear interpolation techniques, these material features are assigned to each three-dimensional coordinate point in the point cloud data, thereby generating a multidimensional 3D material model that includes both geometric structure and material information. In practice, the interpolated material features are bound to the point cloud data, forming 3D points with the following attributes:
[0104] Geometric properties: the three-dimensional coordinates (x, y, z) of the point;
[0105] Material properties: humidity characteristic value, thermal radiation characteristic value.
[0106] The final output multi-dimensional material 3D model retains the spatial structure of the point cloud data, and at the same time superimposes material feature distribution information on each point, providing complete input for subsequent modeling and high-frequency feature analysis.
[0107] Nonlinear interpolation technology accurately combines the spatial distribution of humidity gradients and thermal radiation intensity with a 3D geometric model, generating a multi-dimensional material 3D model that simultaneously displays the morphological features and material changes of the target area. Through the mapping and fusion of material features, points of humidity anomaly and thermal radiation anomaly can be clearly marked in scenarios such as landslide monitoring, providing a scientific basis for identifying landslide precursor areas. The nonlinear characteristics of the interpolation method ensure that the mapping of material features maintains high accuracy and detail even when data distribution is uneven or local details exist, avoiding feature ambiguity caused by linear interpolation.
[0108] In one example, in a high-risk landslide area, a drone was used to collect multispectral video frame data and point cloud depth data, and nonlinear interpolation technology was used to perform multimodal data fusion.
[0109] Input data: Point cloud depth data, collected by LiDAR, with a point cloud density of 1000 points per square meter; multispectral video frame data, with the near-infrared band reflecting the humidity gradient distribution and the short-wave infrared band showing the changes in thermal radiation intensity.
[0110] For each point cloud data point p iThe corresponding region is found in the multispectral frame data using three-dimensional coordinates, and the humidity feature value and thermal radiation feature value are calculated. The interpolation function assigns the material feature value to each point cloud coordinate point according to the principle of weighted inverse distance.
[0111] Output: A multi-dimensional material 3D model, where each point contains 3D coordinates and material characteristics (humidity gradient value and thermal radiation intensity value). The model demonstrates the superposition of the geometry and material properties of a high-risk landslide area.
[0112] In the generated model, the humidity gradient around the crack area shows significant anomalies, while the shortwave infrared thermal radiation intensity forms high points in the local subsidence area. By overlaying material and geometric features, the model clearly marks high-risk points of landslide precursors, providing complete data for subsequent dynamic expansion prediction and risk assessment.
[0113] Preferably, the feature extraction algorithm includes performing a two-dimensional fast Fourier transform on the humidity and thermal radiation features of the high-resolution volume data model to extract high-frequency components, which are used to capture details of crack propagation and local settlement in the target area; the extracted high-frequency components are further processed by continuous wavelet transform, which refines the spatial representation of crack edges and settlement features by decomposing the local changes of high-frequency components.
[0114] The feature extraction algorithm in this invention takes a high-resolution volumetric data model as input and extracts high-frequency components of humidity and thermal radiation features using a two-dimensional fast Fourier transform (2D-FFT) to capture detailed features of the target area (such as crack propagation and local settlement). Subsequently, the extracted high-frequency components are further processed using continuous wavelet transform (CWT) to decompose the local variations of the high-frequency components, thereby refining the spatial representation of crack edges and settlement features.
[0115] 2D-FFT is used to transform spatial domain data to the frequency domain, extracting high-frequency components corresponding to detailed changes in humidity and thermal radiation features. Changes in humidity gradients and thermal radiation intensity are represented as high-frequency signals in the frequency domain; therefore, Fourier transform can significantly amplify these detailed features. The humidity and thermal radiation features of a high-resolution volumetric data model are stored in the form of a two-dimensional matrix, and the Fourier transform formula is as follows:
[0116]
[0117] F(u,v) represents the frequency domain signal; f(x,y) represents the input spatial signal; M and N represent the row and column sizes of the matrix; (u,v) represents the frequency domain coordinates. By performing high-pass filtering on the frequency domain signal, only high-frequency components are retained to enhance the detailed features of crack propagation and local settlement.
[0118] Based on the extracted high-frequency components, continuous wavelet transform is used to decompose and refine local variations. Wavelet transform constructs a mother wavelet function with localization properties to decompose the high-frequency signal, making the characteristics of cracks and settlement edges clearer. The wavelet transform formula is:
[0119]
[0120] W(a,b) represents the wavelet transform coefficients, indicating the component intensity of the signal under scaling a and translation b; f(t) represents the input signal; ψ(t) represents the mother wavelet function; a represents the scaling parameter; and b represents the translation parameter. The magnitude of the wavelet coefficients indicates the degree of signal variation at different scales. By selecting an appropriate mother wavelet function (such as the Morlet wavelet), the details of crack edges and settlement areas can be refined more accurately.
[0121] The high-frequency features output by wavelet transform are used to re-annotate the spatial features of cracks and settlement areas. Combined with the geometric information of the high-resolution volume data model, a refined feature representation is finally formed.
[0122] Extracting high-frequency components using 2D-FFT significantly enhances the details of cracks and settlement corresponding to changes in humidity gradients and thermal radiation. These details, which may be masked by low-frequency information in the spatial domain, are clearly expressed through frequency domain analysis. Continuous wavelet transform, through multi-scale decomposition, further amplifies the local features of high-frequency components, significantly refining the discontinuous changes at crack edges and the boundary features of settlement areas, making it particularly suitable for detailed analysis of landslide precursor areas. Combining high-frequency features with the geometric information of a high-resolution volumetric data model generates a feature representation that includes both the geometric structure of the target area and refines local details, providing high-precision input for subsequent dynamic expansion prediction.
[0123] In an example, in a high-risk landslide area, a drone was used to collect multimodal data and generate a high-resolution volumetric data model. Feature extraction algorithms were then used to perform high-frequency analysis of humidity gradients and thermal radiation characteristics, capturing details of crack propagation and settlement.
[0124] 2D-FFT was applied to the humidity gradient and thermal radiation feature matrices respectively to preserve high-frequency components in the frequency domain and enhance texture details in crack propagation and settlement areas. A high-pass filter was used to remove low-frequency signals, highlighting the edges of humidity changes and local anomalies in thermal radiation.
[0125] The extracted high-frequency components were decomposed into multiple scales using continuous wavelet transform. Appropriate wavelet scaling parameter 'a' and translation parameter 'b' were selected to finely annotate the width, direction, and edge protrusions of the cracks. The boundary variations of the settlement area were further optimized using the spatial distribution of wavelet coefficients.
[0126] Crack characteristics: The width and depth of the cracks are clearly visible after wavelet refinement, showing the crack propagation trend and local details. Settlement characteristics: The high-frequency components of the thermal radiation characteristics form a concentrated distribution in the settlement area, and the boundary details are clearly marked through wavelet analysis.
[0127] Based on high-resolution volumetric data models and high-frequency feature data, a causal network is constructed and causal graph nodes are defined. Causal inference methods are used to quantify the causal relationships between nodes and generate causal contribution data. Time series modeling is used to model dynamic feature data, predict the future change trend of the target area, and generate extended prediction data.
[0128] This invention utilizes high-resolution volumetric data models and high-frequency feature data to construct causal networks, quantifies the causal relationships between nodes through causal inference methods, and predicts the dynamic change trends of target regions based on time series modeling, generating extended prediction data.
[0129] A causal network is a directed acyclic graph (DAG) that reflects the causal relationships between variables. In this invention, the nodes of the causal network include humidity changes, thermal radiation intensity, crack propagation magnitude, surface subsidence, and future trends of the target area. These nodes represent causal relationships through edges, and the direction of the edges reflects the direction of the causal relationship. The values of the nodes are calculated using a high-resolution volumetric data model and high-frequency feature data. For example, humidity changes are calculated from the near-infrared reflectance of multispectral video frame data, and crack propagation magnitude is refined from high-frequency feature components.
[0130] Causal relationship modeling uses structure learning algorithms (such as PC algorithm or GRA-GCN model) to generate causal graph structure, and combines domain knowledge and data correlation analysis to adjust the edges and directions of the graph.
[0131] The core of causal inference is quantifying the contribution of each causal relationship, that is, analyzing the direct or indirect impact of a change in one node on changes in other nodes. Through Bayesian network inference, conditional probability distributions are calculated and causal contributions are quantified. Through inference, weights (contribution values) for each causal path are generated and stored as causal contribution data, used to explain the interactions between variables.
[0132] Dynamic feature data reflects changes in humidity, thermal radiation, crack propagation, and other factors over time. Time series modeling predicts future trends in a target area by capturing the temporal dependencies of these changes. Temporal Convolutional Networks (TCNs) are used to model dynamic feature data, their core being the use of one-dimensional convolution to capture long-term dependencies. The time series prediction results output by TCNs are combined with causal contribution data to generate extended prediction data, marking future high-risk points and their extent of expansion in the target area.
[0133] Causal networks clearly define the causal relationships between variables. Particularly in landslide risk monitoring, they can quantify the impact of factors such as humidity changes, thermal radiation intensity and crack propagation, and surface subsidence on landslide triggering, providing a scientific basis for risk prediction. Causal contribution data provides the influence weights between variables, helping to identify the dominant role of key triggers (such as abnormal humidity or thermal radiation) in regional changes. Dynamic characteristic data, through time series modeling, predicts the future expansion trend of the target area, providing a forward-looking reference for the dynamic monitoring of high-risk areas.
[0134] Preferably, the causal network constructs a causal graph based on humidity changes, thermal radiation intensity, crack propagation amplitude, and surface subsidence in a high-resolution volume data model; the causal relationships in the causal graph are quantified using a Bayesian network inference method, whereby the Bayesian network calculates the causal influence weights between nodes based on the prior probability and conditional probability of each node, generating causal contribution data.
[0135] In this invention, a causal network is constructed based on key variables such as humidity variation, thermal radiation intensity, crack propagation amplitude, and surface subsidence in a high-resolution volumetric data model. The causal relationships in the causal graph are then quantified using Bayesian network inference methods to generate causal contribution data. This method can reveal the causal dependencies between variables and the quantitative relationships of their interactions, providing a scientific basis for dynamic monitoring and risk assessment.
[0136] The causal network represents the causal relationships between variables in the form of a directed acyclic graph (DAG), where nodes represent variables and edges represent causal dependencies. Humidity changes, thermal radiation intensity, crack propagation amplitude, and surface subsidence serve as the main nodes of the network, successively reflecting the different stages and effects of landslide induction.
[0137] Node Definition: The value of each node is calculated using a high-resolution volumetric data model. For example, the humidity change node value is derived from the near-infrared reflectance in multispectral frame data, and the crack propagation amplitude is obtained from high-frequency feature component analysis. Causal Graph Generation: An initial causal graph is constructed using a structure learning algorithm (such as the PC algorithm) and then manually adjusted using domain knowledge to ensure that the edges and directions of the causal graph conform to actual physical logic. For example, humidity changes affect crack propagation, and crack propagation further leads to surface subsidence.
[0138] Causal relationship quantification is achieved through Bayesian network inference, which calculates the weights (i.e., causal contributions) of edges in the causal graph based on the prior and conditional probabilities of each node. The formula for Bayesian inference is:
[0139]
[0140] Where P(A|B) represents the conditional probability of A when event B occurs; P(B|A) represents the probability of the influence of event A on B; and P(A) and P(B) represent the prior probabilities of nodes A and B.
[0141] The weight calculation formula is as follows:
[0142] w A→B =P(B|A)-P(B)
[0143] Among them, w A→B This represents the causal contribution of node A to node B, reflecting the portion of the change in the dependent variable caused by the change in the independent variable.
[0144] The weights of each edge are stored as causal contribution data to explain the interactions between variables in the network. For example, the weights of nodes with varying humidity can quantitatively illustrate their direct impact on crack propagation, as well as their indirect impact on surface subsidence through crack propagation. The generation of causal contribution data provides data support for the identification of landslide causes and the optimization of monitoring strategies.
[0145] The constructed causal network reveals the causal relationships between humidity changes, thermal radiation intensity, crack propagation, and surface subsidence, clarifying the dependency paths among variables and their relative impacts. Through Bayesian network inference, the interactions of landslide triggers are quantitatively described. For example, humidity changes contribute a weight of 0.6 to crack propagation, while crack propagation contributes a weight of 0.8 to surface subsidence. The causal contribution data provides a scientific basis for the dynamic adjustment of landslide monitoring and can be used to optimize the monitoring priority of drones, concentrating resources to cover areas with the greatest risk impact.
[0146] In an example, in a high-risk landslide area, a high-resolution volumetric data model was generated using multimodal data collected by a drone, and a causal network was constructed based on humidity changes, thermal radiation intensity, crack propagation amplitude, and surface subsidence.
[0147] The nodes include humidity change, thermal radiation intensity, crack propagation range, and surface subsidence. An initial causal graph is generated using the PC algorithm, and the direction is adjusted by combining domain knowledge so that humidity change directly affects crack propagation, and crack propagation further leads to surface subsidence.
[0148] The direct contribution of humidity variation to crack propagation was quantified as w using Bayesian inference. 湿度→裂缝 =0.6; the indirect contribution of crack propagation to surface subsidence is calculated as w 湿度→裂缝→沉降 =0.48;
[0149] The final causal contribution data stores the weight of each path and the interactions between nodes.
[0150] The causal network revealed that humidity variation was the primary trigger for crack propagation, while anomalies in thermal radiation intensity mainly influenced land subsidence. The causal contribution data provided a basis for UAV path planning, prioritizing coverage of areas with drastic humidity changes.
[0151] Preferably, the time series modeling uses a temporal convolutional network to model dynamic feature data. The convolutional kernel of the temporal convolutional network is dynamically adjusted according to the length of the time series and the rate of feature change to capture the time dependence between humidity changes, crack propagation, and thermal radiation intensity. The output of the temporal convolutional network is used to predict the future trend of the target area and generate extended prediction data.
[0152] In this invention, time series modeling uses a Temporal Convolutional Network (TCN) to analyze dynamic feature data, capturing the temporal dependencies between humidity changes, crack propagation, and thermal radiation intensity, and predicting future trends in the target area to generate extended prediction data. The core of the TCN is to capture long-distance dependent feature changes in the time series through convolutional operations, dynamically adjusting the size and stride of the convolutional kernel to adapt to the rate of feature change within different time windows.
[0153] Temporal convolutional networks (TCNs) are deep learning models based on one-dimensional convolutions that can efficiently process time-series data. Compared to traditional recurrent neural networks (such as LSTMs), TCNs achieve parallel computation through convolution operations and have a stronger ability to capture long-term dependencies.
[0154] Input layer: Dynamic feature data including humidity changes, crack propagation, and thermal radiation intensity are arranged in chronological order to form a multidimensional time series matrix.
[0155] X = [x1, x2, ..., x T ]
[0156] Where X is a time series matrix, x t This represents the feature vector at time t, where T is the length of the time series.
[0157] Convolutional layer: TCN extracts features from the input sequence through one-dimensional convolution operations.
[0158]
[0159] O t W represents the output at time t. i σ represents the kernel weights; k represents the kernel size; σ represents the activation function (e.g., ReLU); b represents the bias term. The kernel size is dynamically adjusted based on the time series length and the feature change rate to ensure sufficient time window coverage, thereby capturing the time dependence between humidity changes and crack propagation and thermal radiation intensity.
[0160] TCN captures feature correlations over long time spans through multi-layered stacked convolutional operations. For example, in dynamic sequences of humidity changes, local anomalies may trigger crack propagation, and the accumulation of crack propagation may lead to anomalous changes in thermal radiation intensity. The receptive field of TCN expands with the number of layers, enabling it to capture these dependencies across time periods.
[0161] The TCN output layer predicts future trends in the target area based on the captured temporal dependencies. Output data includes the distribution of future humidity changes, the rate and direction of crack propagation, and the spatial distribution of thermal radiation intensity, used to identify high-risk areas and the extent of crack expansion.
[0162] TCN (Transient Convolutional Neural Network) efficiently captures the nonlinear time dependencies between humidity changes, crack propagation, and thermal radiation intensity by dynamically adjusting the size and stride of the convolutional kernel, significantly improving the accuracy of landslide precursor feature analysis. The extended prediction data generated by the model clearly indicates future trends in the target area, providing precise guidance for dynamic monitoring and risk management by UAVs. Compared to recurrent neural networks, TCN's parallel computing capabilities significantly shorten training and inference time, making it suitable for the real-time processing needs of high-frequency UAV video stream analysis.
[0163] In an example, in a high-risk landslide area, based on time-series data of humidity changes, crack propagation, and thermal radiation intensity collected by drones, a temporal convolutional network is used to model and predict the trend of changes in the next 3 hours.
[0164] Input data: Time series matrix X, containing humidity change series, crack propagation series and thermal radiation intensity series, with a data length of 72 (sampling interval of 5 minutes).
[0165] Modeling process: The TCN convolution kernel size was set to 3, the stride to 1, the number of convolutional layers to 6, and the receptive field to cover a time step of 30 (approximately 150 minutes). The model captured the hysteretic effect of local anomalies in the humidity change sequence on crack propagation, as well as the cumulative effect of crack propagation on thermal radiation intensity anomalies.
[0166] Output results: Humidity change prediction: The humidity gradient will increase by 15% in high-risk areas over the next 3 hours. Crack propagation prediction: The average crack width will increase by 5 cm, and the length will increase by 20 cm. Thermal radiation intensity prediction: The number of thermal anomalies will increase by 50%, and the cracks will extend downstream by approximately 30 m.
[0167] The extended prediction data generated by the temporal convolutional network clearly marked the future trends of crack propagation and thermal anomalies, providing accurate reference for UAV monitoring mission planning and early warning of high-risk areas.
[0168] Input extended prediction data and generate dynamic risk heatmaps through feature fusion algorithms; use generative adversarial networks to optimize risk labeling of target areas, and combine reinforcement learning path planning algorithms with dynamic risk heatmaps to optimize UAV flight paths, prioritizing the monitoring of high-risk parts of target areas; feed back newly added multimodal acquisition data to dynamically update volumetric data models and risk heatmaps, achieving a closed-loop monitoring system.
[0169] This invention utilizes extended predictive data to generate dynamic risk heatmaps through feature fusion algorithms, optimizes the labeling of high-risk areas using generative adversarial networks (GANs), and adjusts the UAV's flight mission based on reinforcement learning path planning algorithms. Newly added multimodal acquisition data is fed back in real time to dynamically update the volumetric data model and risk heatmaps, forming a monitoring closed loop.
[0170] Dynamic risk heatmaps are generated by fusing and expanding key features such as humidity changes, crack propagation, and thermal radiation intensity from predictive data, and are used to visualize the risk distribution of labeled target areas. The feature fusion algorithm is based on the cross-attention mechanism, which calculates the weight relationships between multimodal features to dynamically adjust the contribution of each feature in the heatmap.
[0171] After the heatmap is generated, the annotation accuracy is optimized using a Generative Adversarial Network (GAN). A GAN consists of a generator and a discriminator. The generator produces a high-resolution map with risk region annotations, while the discriminator evaluates the realism of the generated map and guides the generator to improve. Generator: Based on a multi-layer convolutional neural network, it takes initial dynamic risk heatmap data as input and generates a high-resolution heatmap. Discriminator: Classifies the generated heatmap through convolutional layers, determines the accuracy of risk region annotations, and provides feedback to the generator for optimization.
[0172] Based on dynamic risk heatmaps, a reinforcement learning path planning algorithm calculates the optimal flight path for drones, prioritizing coverage of high-risk areas and dynamically adjusting monitoring tasks. The algorithm guides drone path planning through a reward function to achieve maximum coverage with minimal cost.
[0173] State space: The current location of the drone and the risk distribution in the dynamic risk heatmap;
[0174] Action space: the next flight path of the drone;
[0175] Reward function: Paths that cover high-risk areas receive higher rewards.
[0176] The newly added multimodal data acquisition is fed back in real time to dynamically update the volumetric data model and risk heatmap. The volumetric data model is updated by integrating the newly added humidity and thermal radiation features to optimize spatial distribution; the heatmap is updated by recalculating feature weights based on the new data to ensure the real-time nature and accuracy of risk labeling.
[0177] The dynamic risk heatmap generated by the feature fusion algorithm can accurately mark the range and changing trend of high-risk areas, providing a scientific basis for UAV mission planning; the generative adversarial network significantly improves the resolution and labeling accuracy of the heatmap, making the boundaries of risk areas clearer and the labeling of anomalies more accurate; the reinforcement learning path planning algorithm optimizes the flight path of the UAV based on the dynamic risk heatmap, prioritizing the coverage of high-risk areas and avoiding resource waste; and the dynamic linkage between UAV data collection, volume data updating and risk heatmap generation is realized, ensuring the real-time and accuracy of landslide monitoring.
[0178] Preferred, such as Figure 3 As shown, the dynamic risk heatmap integrates humidity features, thermal radiation features, and geometric features from a high-resolution volume data model through a multimodal interactive attention mechanism. The interactive attention mechanism dynamically adjusts the contribution of humidity features and thermal radiation features in the annotation of high-risk areas by calculating the weight matrix between each modal feature, thereby generating a dynamic risk heatmap with enhanced resolution and prominent key features.
[0179] This invention utilizes a multimodal interactive attention mechanism to fuse humidity, thermal radiation, and geometric features from a high-resolution volumetric data model, dynamically generating a dynamic risk heatmap with enhanced resolution and prominent key features. The core of the interactive attention mechanism lies in dynamically calculating the weight matrix of each modal feature, allowing the contribution weight of different features in the annotation of high-risk areas to adaptively adjust, ensuring that the risk heatmap accurately highlights the key risk points and their changing trends in the target area.
[0180] The high-resolution volumetric data model provides distribution data for humidity, thermal radiation, and geometric features. These features are fused into a unified risk heatmap through a multimodal interactive attention mechanism. The multimodal interactive attention mechanism calculates a weight matrix based on the spatial correlation of the features. Each feature T... i The weight w of (x,y) i (x, y) is calculated using the following formula:
[0181]
[0182] Among them, Q i and K i Represents the query vector and key vector for the i-th feature; d k Indicates the dimensionality of the feature; w i (x,y) represents the feature T i The weight of (x,y) at position (x,y) reflects the strength of the feature's contribution.
[0183] The calculated weight matrix w i(x,y) is applied to the feature values to generate the fused feature R(x,y):
[0184]
[0185] Where N is the number of modes, and R(x,y) represents the risk value at position (x,y) in the dynamic risk heatmap.
[0186] The interactive attention mechanism dynamically adjusts the distribution of the weight matrix by calculating the spatial and semantic relationships of each modality feature in real time. For example, when the local changes in humidity feature are abnormally significant, its weight will increase in the corresponding region, giving priority to labeling high-risk points; when the anomalies of thermal radiation feature and geometric feature overlap, the weights of these regions will be further superimposed, strengthening the labeling of high-risk areas.
[0187] The fused dynamic risk heatmap has the following characteristics: enhanced resolution, the fusion of multimodal features fully preserves the boundary and texture details of the target area; prominent risks, the annotation of high-risk areas is clearer, and the feature expression of key areas is more significant.
[0188] The multimodal interactive attention mechanism can accurately calculate the correlation between humidity, thermal radiation, and geometric features, ensuring a reasonable distribution of feature contributions in high-risk areas. For example, areas where humidity anomalies and crack propagation spatially overlap are assigned higher weights, improving the accuracy of heatmap annotation. The weight matrix is updated in real time with the dynamic changes of features, enabling the risk heatmap to quickly adapt to the risk evolution trend of the target area. The generated dynamic risk heatmap not only has higher resolution but also clearer boundaries of high-risk areas and richer detail.
[0189] In an example, a dynamic risk heat map is generated in a high-risk landslide area in a mountainous region using a multimodal interactive attention mechanism. This heat map is used to mark abnormal humidity changes, high-value areas of thermal radiation, and areas with significant changes in geometric features.
[0190] Input data: Humidity features, humidity gradient distribution data from the near-infrared band; thermal radiation features, thermal radiation intensity distribution data from the short-wave infrared band; geometric features, terrain variation information from point cloud depth data.
[0191] Feature fusion process: Calculate the weight matrix of humidity feature and thermal radiation feature, and dynamically adjust the contribution intensity of humidity feature through interactive attention mechanism to increase the weight of humidity abnormal area; in the overlapping area of humidity abnormality and thermal radiation abnormality, the weight is further superimposed to ensure that the marked high-risk area is significantly highlighted; apply the weight to feature value to generate fused risk heat map.
[0192] Output heat map: High-risk areas are clearly marked, and the crack edges in areas with significant humidity changes are clearly visible; areas with high thermal radiation values show the boundary range of local subsidence.
[0193] Dynamic risk heat maps accurately mark high-risk points and their changing trends in target areas, providing a scientific basis for UAV mission planning and technical support for regional risk warning and response strategies.
[0194] Preferred, such as Figures 4-5 As shown, the reinforcement learning path planning algorithm prioritizes high-risk areas based on the dynamic risk heatmap, and calculates the optimal flight path by combining the UAV's current position and flight parameters. The path planning algorithm optimizes the path covering high-risk areas through a reward function, adjusts the UAV's flight mission in real time, and generates mission feedback data to update the dynamic risk heatmap.
[0195] The reinforcement learning path planning algorithm in this invention uses a dynamic risk heatmap as input and combines the UAV's current location and flight parameters to calculate the optimal flight path. The algorithm guides the UAV to prioritize covering high-risk areas through a reward function, while simultaneously adjusting the flight mission in real time to adapt to the dynamically changing risk distribution. During path planning, the algorithm updates the dynamic risk heatmap by generating mission feedback data, achieving closed-loop optimization of the UAV monitoring mission.
[0196] The dynamic risk heatmap marks high-risk points and their priorities within the target area. Reinforcement learning algorithms use these priorities as input for path planning, optimizing the UAV's flight path by maximizing coverage of high-priority areas. The state space, consisting of the UAV's current position, flight direction, and the risk distribution in the dynamic risk heatmap, is defined as follows:
[0197] S t ={p t ,θ t ,R}
[0198] S t p represents the state at time t; t θ represents the position of the drone at time t; t R indicates the flight direction of the drone; R represents dynamic risk heatmap data.
[0199] Action space: defined as the set A of possible flight directions for the UAV's next move. t .
[0200] The core of the reward function lies in guiding the drone's path to maximize coverage of high-risk areas while minimizing flight costs (such as path length). The reward function can be expressed as:
[0201] R(S t At )=α·C high-risk -β·C cost
[0202] R(S t A t ) indicates that in state S t Take action A t Reward value; C high-risk Indicates the contribution value covering high-risk areas; C cost The cost of the flight path is represented by α and β, which are weighting parameters that adjust the balance between high-risk coverage and flight cost. The reward function calculates C based on the risk value in the dynamic risk heatmap. high-risk Drones that cover areas with higher risk levels will receive higher rewards.
[0203] The algorithm is based on reinforcement learning frameworks (such as Deep Q-Learning) and uses the value function Q(S) t A t Predicting in state S t Take action A t The maximum cumulative reward that can be obtained. During training, the value function is updated using the following formula:
[0204]
[0205] η represents the learning rate; γ represents the discount factor, which measures the importance of future rewards; and A′ represents the possible next action. Through repeated training, the algorithm can find the path that maximizes the cumulative reward, i.e., the optimal flight path for the drone.
[0206] During path planning, new multimodal data (such as humidity changes and thermal radiation intensity) collected in real time by the UAV are used to update the dynamic risk heat map. By recalculating the priority of risk areas, the algorithm can dynamically adjust the flight path, forming a closed-loop optimization.
[0207] The algorithm ensures that drones prioritize coverage of high-risk areas marked on dynamic risk heatmaps, achieving efficient resource utilization. Combined with real-time task feedback and dynamic heatmap updates, the algorithm can quickly adapt to dynamic changes in the risk distribution of the target area, ensuring the timeliness of monitoring tasks. The reward function balances high-risk coverage with flight costs, optimizes the drone's flight path, and significantly improves monitoring efficiency.
[0208] In an example, in a high-risk landslide area, a reinforcement learning path planning algorithm, combined with a dynamic risk heat map, is used to optimize the monitoring path of a drone.
[0209] Input data: Dynamic risk heatmap R, marking high-risk areas based on humidity changes and heat radiation intensity; current drone location p. t= (x, y) and flight direction θ t .
[0210] Path planning: The reward function is set as follows:
[0211] R(S t A t )=2·C high-risk -0.5°C cost
[0212] C high-risk The reward for covering high-risk areas is calculated based on the risk value in the heatmap; C cost This represents the penalty for path length. The algorithm predicts the optimal path using Deep Q-Learning, ensuring that the drone covers the area with the highest risk value in the heatmap.
[0213] The drones collect newly added humidity changes and thermal radiation intensity data to update the weight distribution of the dynamic risk heat map; the routes are replanned to include newly added high-risk areas in the monitoring task.
[0214] The optimized route enabled the drone to cover all high-risk points marked on the heatmap, reducing the route length by 20% and shortening the monitoring time by 15 minutes. With new data feedback, the algorithm replanned the route to cover dynamically changing risk areas.
[0215] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for analyzing UAV video streams based on high-frequency digital information analysis, characterized in that, Includes the following steps: Based on a dynamic partitioning strategy, UAVs are used to collect multimodal data within the target area. The multimodal data includes multispectral video frame data and point cloud depth data. Data denoising and spatial alignment are then performed to generate preliminary 3D model data. Based on the preliminary 3D model data, a low-resolution volumetric data model is generated using 3D modeling algorithms, and a high-resolution volumetric data model is generated by refining key areas. Using multimodal data fusion technology, the humidity and thermal radiation features of multispectral video frame data are mapped to the high-resolution volumetric data model using nonlinear interpolation methods. Based on the coordinate relationship between multispectral data and point cloud depth data in 3D space, material features are associated with geometric features to generate a multidimensional material 3D model with humidity gradient and thermal radiation intensity information. Finally, a feature extraction algorithm is used to extract the humidity and thermal radiation features from the multidimensional material 3D model to extract high-frequency feature data and low-frequency contour data of the target area. Based on high-resolution volumetric data models and high-frequency feature data, a causal network is constructed and causal graph nodes are defined. Causal inference methods are used to quantify the causal relationships between nodes and generate causal contribution data. Time series modeling is used to model dynamic feature data, predict the future change trend of the target area, and generate extended prediction data. Input extended prediction data and generate a dynamic risk heatmap through a feature fusion algorithm; Generative adversarial networks are used to optimize risk labeling in target areas, and a reinforcement learning path planning algorithm combined with dynamic risk heatmaps is used to optimize the UAV flight path, prioritizing the monitoring of high-risk parts of the target area. The newly added multimodal data is fed back to dynamically update the volume data model and risk heat map, thereby achieving a closed-loop monitoring system.
2. The method according to claim 1, characterized in that, The multi-spectral video frame data in the multimodal data includes the visible light band, used to capture surface texture features of the target area; the near-infrared band, used to monitor changes in soil and vegetation humidity; and the short-wave infrared band, used to detect local anomalies in surface thermal radiation intensity.
3. The method according to claim 2, characterized in that, The multispectral video frame data in the multimodal data is denoised using an adaptive filter. The adaptive filter dynamically adjusts the filtering parameters according to the band characteristics of the multispectral video frame data to eliminate noise caused by changes in light. The point cloud depth data is filtered to remove outliers. The statistical filter removes depth data points that deviate from the expected range based on the local distribution characteristics of the depth information.
4. The method according to claim 1, characterized in that, The three-dimensional modeling algorithm includes generating a low-resolution volume data model using a neural radiation field model. The neural radiation field model performs distributed encoding of the geometric and spectral information of the preliminary three-dimensional model data through ray sampling to generate a low-resolution three-dimensional representation. Based on the low-resolution volumetric data model, a deep super-resolution network is used to refine the cracks and subsidence areas in key regions. The deep super-resolution network reconstructs detailed information layer by layer through convolutional layers to generate a high-resolution volumetric data model.
5. The method according to claim 1, characterized in that, The feature extraction algorithm includes performing a two-dimensional fast Fourier transform on the humidity and thermal radiation features of the high-resolution volume data model to extract high-frequency components, which are used to capture details of crack propagation and local settlement in the target area. The extracted high-frequency components are further processed by continuous wavelet transform. The wavelet transform refines the spatial representation of crack edges and settlement features by decomposing the local changes of high-frequency components.
6. The method according to claim 1, characterized in that, The causal network constructs a causal graph based on humidity changes, thermal radiation intensity, crack propagation amplitude, and surface subsidence in a high-resolution volumetric data model. The causal relationships in the causal graph are quantified using a Bayesian network inference method. The Bayesian network calculates the causal influence weights between nodes based on the prior probability and conditional probability of each node, generating causal contribution data.
7. The method according to claim 6, characterized in that, The time series modeling uses a temporal convolutional network to model dynamic feature data. The convolutional kernel of the temporal convolutional network is dynamically adjusted according to the length of the time series and the rate of feature change to capture the time dependence between humidity changes, crack propagation, and thermal radiation intensity. The output of the temporal convolutional network is used to predict the future trend of the target area and generate extended prediction data.
8. The method according to claim 1, characterized in that, The dynamic risk heatmap integrates humidity features, thermal radiation features, and geometric features from a high-resolution volume data model through a multimodal interactive attention mechanism. The interactive attention mechanism dynamically adjusts the contributions of humidity and thermal radiation features in high-risk area labeling by calculating the weight matrix between various modal features, thereby generating a dynamic risk heatmap with enhanced resolution and prominent key features.
9. The method according to claim 8, characterized in that, The reinforcement learning path planning algorithm prioritizes high-risk areas based on the dynamic risk heatmap, and calculates the optimal flight path by combining the UAV's current position and flight parameters. The path planning algorithm optimizes the path covering high-risk areas through a reward function, adjusts the UAV's flight mission in real time, and generates mission feedback data to update the dynamic risk heatmap.