Unmanned aerial vehicle video stream analysis method based on high-frequency digital information analysis

By adopting high-frequency digital information analysis and multimodal data fusion technology in drone video stream analysis, combined with causal network and reinforced learning path planning, the problem of difficult to capture high-frequency changes in complex environments and fusion of multimodal features in the existing technology is solved, and efficient and accurate geological disaster monitoring and early warning are achieved.

CN120071198AActive Publication Date: 2025-05-30INNER MONGOLIA POLICE COLLEGE +1

Patent Information

Application Number
CN202510140196.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

Existing drone video stream analysis technology is difficult to capture high-frequency variation characteristics in complex environments, cannot effectively integrate other modal features such as humidity and thermal radiation, and lacks real-time path adjustment strategies based on risk priority, resulting in waste of resources or monitoring blind spots.

Method used

The drone video stream analysis method based on high-frequency digital information analysis is adopted to collect multimodal data through dynamic partitioning strategies, including multispectral video frame data and point cloud depth data, data denoising and spatial alignment processing are performed, and preliminary three-dimensional model data is generated. Then, through multimodal data fusion, causal network modeling, time series modeling and reinforcement learning path planning, a dynamic risk heat map is constructed and the drone flight path is optimized.

Benefits of technology

It significantly improves the real-time, flexibility and accuracy of drone monitoring, and can quickly label high-risk areas in complex environments, providing efficient solutions for early warning and response to geological disasters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071198A_ABST
    Figure CN120071198A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicle video stream analysis and geological disaster monitoring, in particular to an unmanned aerial vehicle video stream analysis method based on high-frequency digital information analysis, and the method comprises the following steps: generating a dynamic risk heat map through a multi-modal data fusion technology; performing time sequence modeling on the humidity change, the crack propagation and the thermal radiation intensity by using a time convolutional network, and predicting a future change trend of the target area; the flight path of the unmanned aerial vehicle is optimized based on a reinforcement learning path planning algorithm, and the high-risk area is preferentially covered; the unmanned aerial vehicle collects newly-added data in real time and is used for dynamically updating the heat map and the volume data model to form a monitoring closed loop. According to the method, high-risk region labeling is optimized through a multi-mode interactive attention mechanism, the thermal map resolution and precision are improved by utilizing the generative adversarial network, the efficiency and precision of unmanned aerial vehicle monitoring are improved by combining real-time path optimization, and the method is suitable for dynamic geological disaster monitoring in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of UAV video stream analysis and geological disaster monitoring, and particularly to a UAV video stream analysis method based on high-frequency digital information analysis. Background Art

[0002] UAV video stream analysis is of great significance in the field of UAV monitoring. It uses video streams collected at high frequencies and intelligent analysis technologies to perform real-time processing on dynamically changing scenes, especially in the fields of geological disaster monitoring, emergency response, environmental protection, etc., providing high-precision data support. Through high-frequency digital information analysis, abnormal changes can be identified faster, and an efficient early warning mechanism can be provided.

[0003] In the prior art (Chinese invention patent, publication number: CN117854256B, title: Geological disaster monitoring method based on UAV video stream analysis), a UAV is used to collect video streams and a geological disaster is monitored through an optical flow calculation method. Its main technical means include real-time collection of image frames, grayscale processing, optical flow calculation, and generation of a motion amplitude distribution map. However, its technical defects are mainly manifested as follows:

[0004] In the prior art, optical flow calculation is performed by processing grayscale images frame by frame, and it is difficult to capture high-frequency change features in complex environments. Especially in areas where crack propagation or humidity anomalies are relatively complex, there is an accuracy bottleneck in optical flow calculation; the prior art is mainly based on single video stream data and cannot fuse other modal features such as humidity and thermal radiation, resulting in insufficient description of complex geological disasters; the prior art lacks a real-time path adjustment strategy based on risk priority, making it difficult to optimize the monitoring task allocation of UAVs, resulting in resource waste or monitoring blind spots. Summary of the Invention

[0005] In view of the many problems existing in the above prior art, the present invention provides a UAV video stream analysis method based on high-frequency digital information analysis. Based on multi-modal data fusion and high-frequency digital information analysis, the present invention constructs a set of geological disaster monitoring methods integrating dynamic risk heat map generation, time series modeling, and reinforcement learning path planning. By fusing humidity, thermal radiation, and geometric features, the risk priority is dynamically calculated and the UAV path planning is optimized to achieve precise monitoring and risk early warning of the target area. The present invention significantly improves the real-time performance, flexibility, and accuracy of UAV monitoring, can quickly mark high-risk areas in complex environments, and provides an efficient solution for the early warning and response of geological disasters.

[0006] A UAV video stream analysis method based on high-frequency digital information analysis includes the following steps:

[0007] Based on the dynamic partitioning strategy, use drones to collect multimodal data within the target area. The multimodal data includes multispectral video frame data and point cloud depth data, and perform data denoising and spatial alignment processing to generate preliminary three-dimensional model data;

[0008] Based on the preliminary three-dimensional model data, use three-dimensional modeling algorithms to generate a low-resolution volume data model, and perform refinement processing on key areas to generate a high-resolution volume data model; Combine multispectral video frame data and point cloud depth data through multimodal data fusion technology, map the material characteristics of the target area to the three-dimensional model, and extract high-frequency feature data and low-frequency contour data of the target area through feature extraction algorithms;

[0009] Based on the high-resolution volume data model and high-frequency feature data, construct a causal network and define causal graph nodes; Use causal inference methods to quantify the causal relationships between nodes and generate causal contribution data; Use time series modeling to model dynamic feature data and predict the future change trend of the target area to generate extended prediction data;

[0010] Input the extended prediction data, generate a dynamic risk heat map through feature fusion algorithms; Use a generative adversarial network to optimize the risk annotation of the target area, and optimize the drone flight path by combining the dynamic risk heat map through a reinforcement learning path planning algorithm, giving priority to monitoring the high-risk parts of the target area; Feed back the newly added multimodal acquisition data for dynamically updating the volume data model and risk heat map to achieve a monitoring closed-loop.

[0011] Preferably, the multispectral video frame data in the multimodal data includes a visible light band for capturing the surface texture features of the target area; a near-infrared band for monitoring the humidity changes of soil and vegetation; a short-wave infrared band for detecting local anomalies in the surface thermal radiation intensity.

[0012] Preferably, the multispectral video frame data in the multimodal data is denoised through an adaptive filter, and the adaptive filter dynamically adjusts the filtering parameters according to the band characteristics of the multispectral video frame data to eliminate the noise caused by light changes; The point cloud depth data is cleared of outliers through statistical filtering, and the statistical filtering removes depth data points that deviate from the expected range based on the local distribution characteristics of the depth information.

[0013] Preferably, the three-dimensional modeling algorithm includes using a neural radiance field model to generate a low-resolution volume data model. The neural radiance field model performs distributed encoding on the geometric and spectral information of the preliminary three-dimensional model data through ray sampling to generate a low-resolution three-dimensional representation; On the basis of the low-resolution volume data model, use a depth super-resolution network to refine the cracks and settlement areas in key areas. The depth super-resolution network reconstructs the detailed information layer by layer through convolutional layers to generate a high-resolution volume data model.

[0014] Preferably, the multimodal data fusion technology includes using a non-linear interpolation method to map the humidity characteristics and thermal radiation characteristics of the multispectral video frame data into the three-dimensional model generated by the point cloud depth data. The non-linear interpolation method correlates the material characteristics with the geometric characteristics according to the coordinate relationship between the multispectral data and the point cloud depth data in the three-dimensional space, and generates a multi-dimensional material three-dimensional model with humidity gradient and thermal radiation intensity information.

[0015] Preferably, the feature extraction algorithm includes performing a two-dimensional fast Fourier transform on the humidity and thermal radiation characteristics of the high-resolution volume data model to extract high-frequency components for capturing details of crack propagation and local settlement in the target area; further processing the extracted high-frequency components using continuous wavelet transform. The wavelet transform refines the spatial representation of the crack edges and settlement characteristics by decomposing the local variations of the high-frequency components.

[0016] Preferably, the causal network constructs a causal graph based on the humidity change, thermal radiation intensity, crack propagation amplitude, and ground settlement amount in the high-resolution volume data model; the causal relationship of the causal graph is quantified by the Bayesian network inference method. The Bayesian network calculates the causal influence weights between nodes based on the prior probability and conditional probability of each node, and generates causal contribution data.

[0017] Preferably, the time series modeling models the dynamic feature data through a temporal convolutional network. The convolutional kernel of the temporal convolutional network is dynamically adjusted according to the length of the time series and the feature change rate to capture the temporal dependence relationship between humidity change, crack propagation, and thermal radiation intensity; the output result of the temporal convolutional network is used to predict the future change trend of the target area and generate extended prediction data.

[0018] Preferably, the dynamic risk heat map fuses the humidity characteristics, thermal radiation characteristics, and geometric characteristics from the high-resolution volume data model through a multimodal interactive attention mechanism; the interactive attention mechanism dynamically adjusts the contributions of the humidity characteristics and thermal radiation characteristics in the high-risk area annotation by calculating the weight matrix between the features of each modality, and generates a dynamic risk heat map with enhanced resolution and prominent focus.

[0019] Preferably, the reinforcement learning path planning algorithm calculates the optimal flight path based on the priority of the high-risk areas in the dynamic risk heat map, combined with the current position and flight parameters of the unmanned aerial vehicle; the path planning algorithm optimizes the path covering the high-risk areas through a reward function, adjusts the flight task of the unmanned aerial vehicle in real time, and generates task feedback data for updating the dynamic risk heat map.

[0020] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:

[0021] Through the multi-modal interaction attention mechanism, the present invention integrates humidity, thermal radiation, and geometric features to achieve accurate annotation of high-risk areas, overcoming the problem of insufficient description of complex geological disasters by single-modal data in the prior art;

[0022] Through the reinforcement learning path planning algorithm, the present invention dynamically adjusts the UAV flight mission to achieve path optimization for efficient coverage of high-risk areas, solving the defect of insufficient flexibility in path planning in the prior art;

[0023] Through the temporal convolutional network to model dynamic features, capture the temporal dependencies of humidity changes, crack propagation, and thermal radiation intensity, generate extended prediction data, and achieve accurate early warning of geological disasters, making up for the shortcoming of insufficient analysis of time series features in the prior art;

[0024] The present invention optimizes the dynamic risk heat map through the generative adversarial network, significantly improving the resolution and annotation accuracy of the heat map and enhancing the visualization effect of high-risk areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic flow chart of the method of the present invention;

[0026] Figure 2 is a schematic flow chart of multi-modal data fusion in the present invention;

[0027] Figure 3 is a schematic flow chart of generating a dynamic risk heat map in the present invention;

[0028] Figure 4 is a schematic flow chart of UAV path optimization in the present invention;

[0029] Figure 5 is a schematic flow chart of the monitoring closed-loop update in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure.

[0031] As Figure 1 shown, a UAV video stream analysis method based on high-frequency digital information analysis includes the following steps:

[0032] Based on the dynamic partitioning strategy, use the UAV to collect multi-modal data in the target area. The multi-modal data includes multi-spectral video frame data and point cloud depth data, and perform data denoising and spatial alignment processing to generate preliminary three-dimensional model data;

[0033] In the present invention, the core principle of the dynamic partitioning strategy is to divide the target area into multiple dynamically adjustable sub-areas to achieve efficient task allocation and coverage optimization of the unmanned aerial vehicle (UAV). This strategy uses the Voronoi partitioning method to adjust the partitioning result in real time according to the current position of the UAV and the terrain characteristics of the target area, ensuring the uniform distribution of UAV tasks in space and avoiding monitoring blind spots and data redundancy. Each sub-area corresponds to the task scope of a UAV, and the partitioning result is updated by a dynamic partitioning algorithm to adapt to real-time changing environmental conditions (such as changes in UAV battery power, obstacles, or the risk level of the target area).

[0034] The UAVs collect multi-modal data within their respective partition ranges, including multi-spectral video frame data and point cloud depth data. The multi-spectral video frame data includes the visible light band (for capturing surface textures), the near-infrared band (for monitoring humidity changes), and the short-wave infrared band (for detecting local anomalies in thermal radiation intensity); the point cloud depth data is collected by lidar and is used to provide three-dimensional geometric information of the terrain. Through a time synchronization mechanism, the multi-modal data is time-aligned frame by frame to ensure the spatial consistency of the multi-spectral frames and the point cloud depth frames.

[0035] The collected multi-modal data needs to be processed for data denoising and spatial alignment to improve the reliability and fusion accuracy of the data. Data denoising uses an adaptive filtering method, and by dynamically adjusting the filtering parameters, light interference and random noise are removed respectively for different band characteristics of the multi-spectral frame data. For the point cloud depth data, statistical filtering is used to remove outliers, and noise points deviating from the threshold range are removed according to the local distribution characteristics of the depth information. Spatial alignment uses a global registration method based on the Iterative Closest Point (ICP) algorithm to align the multi-spectral frames and the point cloud depth data in three-dimensional space, ensuring the fusion of multi-modal information in the same spatial coordinate system.

[0036] Through the dynamic partitioning strategy, efficient coverage of the target area by the UAVs can be achieved, maximizing the utilization of monitoring resources. Especially in the case of a large and complex terrain target area, this strategy effectively avoids duplication or omission of UAV tasks, improving the efficiency and integrity of data collection. The multi-modal data collection combines spectral information and three-dimensional geometric information, providing comprehensive data support for subsequent modeling and analysis. The data denoising and spatial alignment processing significantly improve the quality of the multi-modal data, ensuring the accuracy of subsequent volume data modeling and feature extraction.

[0037] Preferably, the multi-spectral video frame data in the multi-modal data includes the visible light band for capturing the surface texture features of the target area; the near-infrared band for monitoring the humidity changes of the soil and vegetation; and the short-wave infrared band for detecting local anomalies in the surface thermal radiation intensity.

[0038] The visible light band is mainly used to capture the surface texture features of the target area. The principle is to generate image frames by receiving the visible light signals reflected by the ground surface through a multispectral camera. Since the visible light reflection intensity is closely related to the surface materials (such as soil, rock, vegetation) and topographic features (such as cracks, settlements) of the ground surface, the frame data of this band can clearly display the geometric texture and details of the ground surface. In the multispectral camera carried by the drone, RGB (red, green, blue) three-channel sensors are used to collect the data of the visible light band. The resolution of each frame of image is set to be higher than 1 cm / pixel to extract high-precision texture features during subsequent processing.

[0039] The near-infrared band is used to monitor the humidity changes of soil and vegetation. The principle is based on the absorption characteristics of near-infrared light in water molecules and plant leaves. By analyzing the strength changes of the band signals, the humidity distribution of the target area is evaluated. The high penetration characteristic of the near-infrared band enables it to detect the humidity changes below the ground surface to a certain extent. The original data collected in the near-infrared band is used to calculate the humidity changes through spectral reflectance analysis. The formula for spectral reflectance R(λ) is:

[0040]

[0041] where R(λ) is the reflectance at wavelength λ, I r (λ) is the intensity of the reflected light received by the sensor, and I i (λ) is the intensity of the incident light. The decrease in reflectance usually corresponds to the increase in humidity. Combining with the humidity distribution change map, potential landslide risk areas can be marked.

[0042] The short-wave infrared band is used to detect local anomalies in the intensity of surface thermal radiation. The principle is to analyze the surface temperature distribution through thermal radiation characteristics. The short-wave infrared band is extremely sensitive to the changes in surface thermal signals and can be used to identify local temperature increases in soil and rocks caused by sliding and stress concentration.

[0043] Specific implementation: The data collection in the short-wave infrared band is carried out through an infrared sensor, and the data contains the temperature distribution of the target area. The thermal radiation intensity EEE of each pixel is calculated according to the simplified formula of Planck's law:

[0044] E = ∈·σ·T 4

[0045] where E is the radiation intensity, ∈ is the emissivity of the ground surface, σ is the Stefan-Boltzmann constant, and T is the absolute temperature. By comparing the temperature anomaly points in the area, high-risk precursor areas of landslides can be marked.

[0046] The multi - band combination of multi - spectral video frame data provides comprehensive support for the high - frequency digital information analysis of the target area. The visible light band makes the texture features of the ground surface (such as cracks and settlements) clearly visible. The near - infrared band enhances the monitoring ability of the ground surface humidity distribution, and the short - wave infrared band is used to identify the thermal radiation anomalies of landslide precursors. Through the fusion of the three - band data, potential risk points in the target area can be accurately located, providing multi - dimensional support for subsequent 3D modeling and causal network analysis.

[0047] In an embodiment, in a high - risk mountainous area prone to landslides, 3 drones equipped with multi - spectral cameras were deployed in the target area to collect data in the visible light, near - infrared, and short - wave infrared bands from different angles respectively. The flight altitude of the drones was set at 100 meters, and the cameras collected three - band data synchronously at a frame rate of 30Hz.

[0048] The collected RGB images clearly show the shape and distribution of surface cracks. By using an image enhancement algorithm to refine the crack edges, the changing trend of crack widths is quantified, providing a basis for the preliminary annotation of high - risk points.

[0049] In the soil and vegetation - covered areas of the target area, by analyzing the spectral reflectance, it is found that the humidity in some areas is abnormally high. These areas are further correlated with the crack distribution, indicating that the humidity anomaly may lead to local settlement.

[0050] In the landslide precursor area, through thermal radiation analysis, it is found that the temperature of some points is significantly higher than that of the surrounding areas. Combining the crack distribution and humidity data, these thermal anomaly points are confirmed as stress concentration areas induced by landslides.

[0051] A multi - modal risk distribution map is generated by integrating the three - band data. The cracks, humidity anomaly points, and high - temperature points of thermal radiation in the target area are accurately marked, providing accurate data input for subsequent dynamic modeling and path optimization.

[0052] Preferably, the multi - spectral video frame data in the multi - modal data is denoised by an adaptive filter. The adaptive filter dynamically adjusts the filtering parameters according to the band characteristics of the multi - spectral video frame data to eliminate the noise caused by light changes; the point - cloud depth data is cleared of abnormal points through statistical filtering, and the statistical filtering removes the depth data points that deviate from the expected range based on the local distribution characteristics of the depth information.

[0053] During the acquisition process, the multi - spectral video frame data is affected by light changes (such as natural light intensity fluctuations and local shadows) and sensor noise, which may lead to a decline in image quality. The adaptive filter adaptively enhances the effective signal and suppresses the invalid noise by dynamically adjusting the filtering parameters according to the band characteristics of the multi - spectral data and the local light environment. The adaptive filter calculates the statistical characteristics of the image based on a local window and dynamically adjusts the filtering weights. The filtering weight calculation formula is:

[0054]

[0055] Among them, W(x,y) represents the filter weight of the pixel point (x,y), is the local variance of the signal, is the estimated variance of the noise. This formula ensures that areas with drastic signal changes (such as edges or texture features) retain more details, while flat areas suppress noise more strongly. Spectral data in different bands (visible light, near infrared, short-wave infrared) have different signal characteristics, so the parameters of the filter need to be calculated independently according to the band characteristics. For example, the short-wave infrared band is sensitive to thermal radiation, and its noise mainly comes from photothermal interference; the filter will give priority to smoothing the noise signal in the low-frequency area.

[0056] During the acquisition process of point cloud depth data, outliers may be generated due to laser scattering, terrain occlusion or sensor error, which will affect the accuracy of terrain modeling. Statistical filtering analyzes the local distribution characteristics of depth data and removes outliers that deviate significantly from surrounding points, thereby improving the quality of point cloud data. In the neighborhood of each depth point, the mean μ and standard deviation σ of the depth value are calculated, and outliers that meet the following conditions are removed:

[0057] |D i -μ|>k·σ

[0058] Among them, D i is the depth value of the current depth point, and k is the filter threshold coefficient, which is used to control the strictness of removing abnormal points. The local distribution characteristics are quickly calculated by establishing a neighborhood search tree in the point cloud data to reduce processing time. In point cloud data of complex terrain (such as landslide areas), abnormal points may appear in the edge area or vegetation covered area. Statistical filtering can effectively remove these abnormal points to ensure the geometric accuracy of subsequent 3D modeling.

[0059] Adaptive filtering of multispectral video frame data significantly improves the signal-to-noise ratio of images in different bands, especially in environments with drastic light changes, retaining more detailed textures and regional features, making the annotation of humidity gradients and thermal radiation anomalies more accurate.

[0060] Statistical filtering of point cloud depth data eliminates depth deviation points caused by noise in lidar acquisition, which significantly improves the spatial consistency of point cloud data and lays a high-quality data foundation for 3D modeling and multimodal data fusion.

[0061] In an embodiment, in a high-risk area for landslides, multispectral video frame data and point cloud depth data are collected by a drone, and denoising is performed by applying an adaptive filter and a statistical filter, respectively.

[0062] Multispectral video frame data is collected by the onboard multispectral camera, including visible light, near-infrared, and short-wave infrared bands. The flight altitude is 120 meters, and the frame rate is 30Hz. Point cloud depth data is collected by lidar, and the generated point cloud density is 1000 points per square meter.

[0063] In the short-wave infrared band, due to obvious signal fluctuations caused by surface thermal radiation and light interference, the adaptive filter dynamically adjusts the weights according to the local changes of the signal, filters out the noise signals in the low-frequency region, and enhances the detailed texture of the crack and humidity anomaly regions. In the vegetated areas, there are many depth anomaly points caused by laser reflection. Statistical filtering removes the points that deviate from the local mean by 3 standard deviations. These points are usually inconsistent with the surrounding terrain.

[0064] After the adaptive filtering process of the multispectral video frame data, the texture details of the thermal radiation anomaly region are clearer, and the boundaries of humidity changes are more obvious. The point cloud depth data after statistical filtering is more coherent in the edge region, and the geometric shape of the terrain is more accurate, providing high-quality input for subsequent modeling.

[0065] Based on the preliminary three-dimensional model data, a low-resolution volume data model is generated using a three-dimensional modeling algorithm, and the key areas are refined to generate a high-resolution volume data model; through multi-modal data fusion technology, the multispectral video frame data and the point cloud depth data are combined, and the material characteristics of the target area are mapped into the three-dimensional model, and the high-frequency feature data and low-frequency contour data of the target area are extracted through feature extraction algorithms;

[0066] The preliminary three-dimensional model data is the product of the denoising and registration processing of the multi-modal data collected by the drone, and contains the geometric structure and basic material characteristics of the target area. The three-dimensional modeling algorithm uses the Neural Radiance Field model (NeRF), and by sampling the volume distribution of each ray in space, a low-resolution volume data model of the reconstructed area is generated. The core of the Neural Radiance Field model is to integrate the color and density of each ray r(t):

[0067]

[0068] Among them, C(r) represents the ray color, σ(r(t)) represents the volume density, c(r(t)) is the color, T(t) is the ray cumulative transmittance, t n and t f are the starting and ending points of the ray. This model can quickly generate a low-resolution three-dimensional representation and provide the basic structure of the global area.

[0069] Based on the low-resolution volume data model, the marked key areas (such as cracks and settlements) are refined. The refinement uses a deep super-resolution network (SR-Net), and the features of the area are enhanced layer by layer through multi-layer convolution to restore details. Each convolution operation is based on the following formula:

[0070] f l =ReLU(W l ·f l-1 +b l )

[0071] where f l represents the feature output of the l-th layer, W l and b l are the weight and bias respectively, and ReLU is the activation function. Through multiple feature extractions and refinements, a three-dimensional model with high spatial resolution is generated.

[0072] To enhance the material characteristics of the three-dimensional model, the multi-modal data fusion technology is used to map the humidity features and thermal radiation features in the multi-spectral video frame data to the three-dimensional model. The non-linear interpolation technology fuses the multi-spectral frame data and the point cloud depth data through the three-dimensional coordinate relationship to achieve the precise association of geometric information and material information.

[0073] The high-resolution volume data model is subjected to feature decomposition to extract the high-frequency feature data (such as the tiny changes at the crack edges and settlement areas) and low-frequency contour data (such as the overall terrain changes) of the target area. The high-frequency features are extracted by two-dimensional fast Fourier transform (2D-FFT), and its formula is:

[0074]

[0075] where F(u, v) represents the frequency-domain signal, f(x, y) is the spatial signal, and M and N are the spatial dimensions. The texture details are extracted through the high-frequency components in the frequency domain, and the edge features are refined by combining with the continuous wavelet transform. The low-frequency contour retains the overall shape and change trend of the target area for global modeling.

[0076] Through the three-dimensional modeling algorithm and multi-modal data fusion technology, the geometric and material information of the target area is represented with high precision, and the high-resolution volume data model provides a solid foundation for subsequent analysis. The high-frequency feature extraction enhances the detail expression of the target area. Especially in landslide monitoring, cracks, settlements, and local anomalies are accurately marked, while the low-frequency contour data provides the overall trend information for subsequent dynamic expansion prediction.

[0077] Example: In a certain mountainous area, the initial three-dimensional model data of the target area is generated from the multi-modal data collected by drones. Cracks and settlements in the area are the key monitoring objects and require high-resolution modeling.

[0078] Use a neural radiance field model to globally model the preliminary 3D model data and generate a low-resolution volume data model. The model shows the basic terrain of the area, including the main ridges, slopes, and depressions.

[0079] For the crack and settlement areas, use a deep super-resolution network to enhance features layer by layer and generate a high-resolution volume data model. After processing, the width changes of the cracks and the boundaries of the settlement areas are clearly visible.

[0080] Map the humidity gradient and thermal radiation intensity of the multispectral frame data into the 3D model through non-linear interpolation technology. Humidity anomaly points are shown in the crack areas, while the thermal radiation intensity in the settlement areas is higher than that of the surroundings.

[0081] Perform a two-dimensional fast Fourier transform on the high-resolution volume data model, extract the high-frequency features of the crack edges and settlement areas, and further refine the details of the edges through continuous wavelet transform. Extract low-frequency contour data globally to show the overall trend of the area.

[0082] The results of 3D modeling and feature extraction clearly show the distribution and expansion direction of the cracks, and the humidity and thermal anomaly points in the settlement areas are accurately marked, providing high-precision input data for dynamic risk assessment and expansion prediction.

[0083] Preferably, the 3D modeling algorithm includes using a neural radiance field model to generate a low-resolution volume data model. The neural radiance field model performs distributed encoding on the geometric and spectral information of the preliminary 3D model data through ray sampling to generate a low-resolution 3D representation; based on the low-resolution volume data model, use a deep super-resolution network to refine the cracks and settlement areas in the key areas. The deep super-resolution network reconstructs the detailed information layer by layer through convolutional layers to generate a high-resolution volume data model.

[0084] The present invention uses a neural radiance field model (NeRF) to generate a low-resolution volume data model, and on this basis, uses a deep super-resolution network to refine the key parts (such as cracks and settlements) in the target area, and finally generates a high-resolution volume data model, laying a foundation for subsequent high-frequency digital information analysis.

[0085] The neural radiance field model is an algorithm that performs distributed encoding on the spectral and geometric information in 3D space through ray sampling. In the present invention, the preliminary 3D model data is mapped into a continuous voxel distribution in space, and the neural radiance field generates a low-resolution 3D representation of the target area by sampling the color and density of the ray r(t). The light integral formula is as follows:

[0086]

[0087] C(r) represents the light color, which represents the color accumulation from t n (the starting point of the light) to t f (the ending point of the light); σ(r(t)) represents the voxel density, which describes the geometric information of points in space; c(r(t)) represents the color value of the position r(t)r(t)r(t) where the light passes through; T(t) represents the cumulative transmittance of the light, which represents the transmittance from t n to t.

[0088] The neural radiance field model fits σ(r(t)) and c(r(t)) through a multi-layer perceptron (MLP) to generate a low-resolution three-dimensional volume data model, providing a global representation of the target area.

[0089] Based on the low-resolution volume data model, key areas such as cracks and settlements are refined, and a deep super-resolution network is used to gradually enhance the detailed features layer by layer. This network extracts, refines, and reconstructs the three-dimensional features of the input through convolutional operations. Each convolutional layer calculates the output features through the following formula:

[0090] f l = ReLU(W l ·f l-1 + b l )

[0091] f l represents the output features of the l-th layer; W l represents the weight matrix; f l-1 represents the features of the previous layer; b l represents the bias term. The last layer of the network outputs a high-resolution three-dimensional model, where the texture at the crack edges and the geometric details in the settlement area are significantly enhanced.

[0092] The neural radiance field model quickly generates a low-resolution volume data model, fully presenting the basic geometric features of the target area, providing a reliable global framework for subsequent refinement processing. The deep super-resolution network refines the details in the crack and settlement areas, improving the spatial resolution of the model, enabling the clear expression of high-frequency features (such as the discontinuous changes at the crack edges and the local protrusions in the settlement areas). The high-resolution volume data model provides high-precision input data for subsequent high-frequency feature extraction and dynamic expansion prediction, significantly improving the accuracy and reliability of the entire solution.

[0093] Example, in a high-risk landslide area, using the preliminary three-dimensional model data collected by drones, a high-resolution volume data model is generated by combining the neural radiance field model and the deep super-resolution network.

[0094] The initial three-dimensional model data contains multimodal information of the target area, including point cloud depth data and multispectral video frame data. After denoising and registration processing, it is input into the neural radiance field model.

[0095] The neural radiance field model encodes the geometric and spectral information of the target area in a distributed manner by ray sampling, generating a low-resolution volume data model. This model shows the global terrain features of the target area, such as the inclination of the hillside and the distribution of depressions.

[0096] Based on the low-resolution model, the crack and settlement areas are selected as the key areas for refinement. The depth super-resolution network enhances the features layer by layer for these areas, and the texture changes of the cracks and the boundary details of the settlements are clearly expressed in the high-resolution model.

[0097] The generated high-resolution volume data model shows the width, depth, and extension direction of the cracks, and at the same time, it finely describes the boundary and internal morphology of the settlement area, providing high-precision input data for landslide risk assessment.

[0098] Preferably, as Figure 2 shown, the multimodal data fusion technology includes using a non-linear interpolation method to map the humidity feature and thermal radiation feature of the multispectral video frame data into the three-dimensional model generated by the point cloud depth data. The non-linear interpolation method correlates the material feature with the geometric feature according to the coordinate relationship between the multispectral data and the point cloud depth data in the three-dimensional space, generating a multi-dimensional material three-dimensional model with humidity gradient and thermal radiation intensity information.

[0099] The multimodal data fusion technology in the present invention maps the humidity feature and thermal radiation feature of the multispectral video frame data into the three-dimensional model generated by the point cloud depth data through a non-linear interpolation method to generate a multi-dimensional material three-dimensional model with tightly combined material features and geometric features. This model not only provides the geometric shape of the target area but also covers the spatial distribution information of the humidity gradient and thermal radiation intensity, thus providing comprehensive data support for high-frequency digital information analysis.

[0100] The non-linear interpolation method projects the material features (humidity feature and thermal radiation feature) accurately into the three-dimensional geometric model represented by the point cloud data by modeling the coordinate relationship between the multispectral video frame data and the point cloud depth data in the three-dimensional space. Compared with the linear interpolation method, the non-linear interpolation can handle irregularly distributed data points more accurately and retain the detailed features. The non-linear interpolation uses an algorithm of inverse weighted distance to map the material features, and the interpolation function can be expressed as:

[0101]

[0102] F(p) represents the interpolation result of point p, indicating the material feature value of this point; f(p i ) represents the material feature value of the input data point p i ; w i represents the weight, which is determined by the distance d(p, p i ) from point p to p i ; k represents the parameter controlling the distance attenuation, and usually k = 2 is taken to balance the weight attenuation rate. Through the above interpolation calculation, the material features in the multispectral data are mapped point by point into the three-dimensional coordinates generated by the point cloud data.

[0103] The point cloud depth data provides the geometric shape information of the three-dimensional space, while the multispectral video frame data contains humidity and thermal radiation features. Through the non-linear interpolation technique, these material features are assigned to each three-dimensional coordinate point of the point cloud data, thus generating a multi-dimensional material three-dimensional model that contains both geometric structure and material information. In the specific implementation, the interpolated material features will be bound to the point cloud data to form three-dimensional points with the following attributes:

[0104] Geometric attributes: the three-dimensional coordinates (x, y, z) of the point;

[0105] Material attributes: humidity feature value, thermal radiation feature value.

[0106] The finally output multi-dimensional material three-dimensional model retains the spatial structure of the point cloud data, and at the same time superimposes the material feature distribution information on each point, providing a complete input for subsequent modeling and high-frequency feature analysis.

[0107] The non-linear interpolation technique enables the precise combination of the spatial distributions of humidity gradient and thermal radiation intensity with the three-dimensional geometric model. The generated multi-dimensional material three-dimensional model can simultaneously display the morphological features and material changes of the target area. Through the mapping and fusion of material features, it is possible to clearly mark the humidity anomalies and thermal radiation anomaly points in scenarios such as landslide monitoring, providing a scientific basis for identifying the precursor areas of landslides. The non-linear characteristic of the interpolation method ensures that when the data distribution is uneven or there are local details, the mapping of material features can maintain high precision and high details, avoiding feature blurring caused by linear interpolation.

[0108] Example: In a high-risk landslide area, an unmanned aerial vehicle is used to collect multispectral video frame data and point cloud depth data, and the non-linear interpolation technique is adopted for multi-modal data fusion.

[0109] Input data: Point cloud depth data, collected by lidar, with a point cloud density of 1000 points per square meter; multispectral video frame data, where the near-infrared band reflects the humidity gradient distribution and the short-wave infrared band shows the change of thermal radiation intensity.

[0110] For each point cloud data point p i, find the corresponding area in the multi-spectral frame data through three-dimensional coordinates, and calculate the humidity eigenvalue and the thermal radiation eigenvalue; according to the principle of inverse weighted distance, the interpolation function assigns the material eigenvalue to each point cloud coordinate point.

[0111] Output result: A multi-dimensional material three-dimensional model, where each point contains three-dimensional coordinates and material characteristics (humidity gradient value and thermal radiation intensity value). The model shows the result of the superposition of the geometric shape and material characteristics of the high-risk landslide area.

[0112] In the generated model, the humidity gradient around the crack area shows significant anomalies, and the thermal radiation intensity in the short-wave infrared forms high-value points in the local settlement area. Through the superposition of material and geometric features, the model clearly marks the high-risk points of landslide precursors, providing complete data for subsequent dynamic expansion prediction and risk assessment.

[0113] Preferably, the feature extraction algorithm includes performing two-dimensional fast Fourier transform on the humidity and thermal radiation characteristics of the high-resolution volume data model to extract high-frequency components for capturing details of crack propagation and local settlement in the target area; further processing the extracted high-frequency components using continuous wavelet transform, and the wavelet transform refines the spatial representation of crack edges and settlement characteristics by decomposing the local variations of the high-frequency components.

[0114] The feature extraction algorithm in the present invention takes the high-resolution volume data model as input, extracts high-frequency components of humidity and thermal radiation characteristics through two-dimensional fast Fourier transform (2D-FFT) for capturing detailed features (such as crack propagation and local settlement) of the target area. Subsequently, the extracted high-frequency components are further processed using continuous wavelet transform (CWT) to decompose the local variations of the high-frequency components, thereby refining the spatial representation of crack edges and settlement characteristics.

[0115] 2D-FFT is used to transform the spatial domain data into the frequency domain and extract high-frequency components corresponding to the detailed changes in humidity and thermal radiation characteristics. The changes in humidity gradient and thermal radiation intensity are manifested as high-frequency signals in the frequency domain, so these detailed features can be significantly amplified through Fourier transform. The humidity characteristics and thermal radiation characteristics of the high-resolution volume data model are stored in the form of a two-dimensional matrix, and the Fourier transform formula is as follows:

[0116]

[0117] F(u, v) represents the frequency domain signal; f(x, y) represents the input spatial signal; M, N represent the row and column sizes of the matrix; (u, v) represents the frequency domain coordinates. By performing high-pass filtering on the frequency domain signal, only the high-frequency components are retained to enhance the detailed features of crack propagation and local settlement.

[0118] Based on the extracted high-frequency components, continuous wavelet transform is used to decompose and refine local variations. Wavelet transform decomposes high-frequency signals by constructing a mother wavelet function with localization characteristics, making the features of cracks and settlement edges clearer. The wavelet transform formula is as follows:

[0119]

[0120] W(a,b) represents the wavelet transform coefficient, indicating the component intensity of the signal under scaling a and translation b; f(t) represents the input signal; ψ(t) represents the mother wavelet function; a represents the scaling parameter; b represents the translation parameter. The magnitude of the wavelet coefficient represents the degree of change of the signal at different scales. By selecting an appropriate mother wavelet function (such as the Morlet wavelet), the details of the crack edges and settlement areas can be refined more precisely.

[0121] The high-frequency features output by wavelet transform are used to relabel the spatial features of crack and settlement areas. Combining with the geometric information of the high-resolution volume data model, a refined feature representation is finally formed.

[0122] Extracting high-frequency components through 2D-FFT can significantly enhance the crack and settlement details corresponding to humidity gradients and thermal radiation changes. These details may be masked by low-frequency information in the spatial domain and are clearly expressed through frequency-domain analysis. Continuous wavelet transform further amplifies the local features of high-frequency components through multi-scale decomposition, significantly refining the discontinuous changes at crack edges and the boundary features of settlement areas, especially suitable for the fine analysis of landslide precursor areas. The combination of high-frequency features and the geometric information of the high-resolution volume data model generates a feature representation that not only includes the geometric structure of the target area but also refines local details, providing high-precision input for subsequent dynamic expansion prediction.

[0123] Example: In a high-risk landslide area, an unmanned aerial vehicle is used to collect multi-modal data to generate a high-resolution volume data model. Through the feature extraction algorithm, high-frequency analysis is performed on humidity gradient and thermal radiation features to capture crack expansion and settlement details.

[0124] Apply 2D-FFT to the humidity gradient and thermal radiation feature matrices respectively, retain the high-frequency components in the frequency domain, and enhance the texture details of crack expansion and settlement areas. Through a high-pass filter, low-frequency signals are removed to highlight the edges of humidity changes and local anomalies of thermal radiation.

[0125] The extracted high-frequency components are decomposed by continuous wavelet transform at multiple scales. Appropriate wavelet scaling parameter a and translation parameter b are selected to finely label the width, direction, and edge protrusions of cracks. For the boundary changes in the settlement area, further optimization is performed through the spatial distribution of wavelet coefficients.

[0126] Fracture characteristics: After wavelet refinement, the width and depth of the fractures are clearly visible, showing the expansion trend and local details of the fractures. Settlement characteristics: The high-frequency components of the thermal radiation characteristics form a concentrated distribution in the settlement area, and the boundary details are clearly marked through wavelet analysis.

[0127] Based on the high-resolution volume data model and high-frequency feature data, construct a causal network and define the nodes of the causal graph; use the causal inference method to quantify the causal relationship between the nodes and generate causal contribution data; use time series modeling to model the dynamic feature data and predict the future change trend of the target area to generate extended prediction data.

[0128] The present invention constructs a causal network using a high-resolution volume data model and high-frequency feature data, quantifies the causal relationship between the nodes through a causal inference method, and predicts the dynamic change trend of the target area based on time series modeling to generate extended prediction data.

[0129] A causal network is a directed acyclic graph (DAG) that can reflect the causal relationship between variables. In the present invention, the nodes of the causal network include humidity change, thermal radiation intensity, fracture expansion amplitude, ground settlement amount, and the future change trend of the target area. These nodes represent the causal relationship through edges, and the direction of the edges reflects the direction of the causal relationship. The values of the nodes are calculated through the high-resolution volume data model and high-frequency feature data. For example, the humidity change is calculated from the reflectance of the near-infrared band of the multispectral video frame data, and the fracture expansion amplitude is obtained by refining the high-frequency feature components.

[0130] Causal relationship modeling uses a structure learning algorithm (such as the PC algorithm or the GRA-GCN model) to generate the causal graph structure, and combines domain knowledge and data correlation analysis to adjust the edges and directions of the graph.

[0131] The core of causal inference is to quantify the contribution value of each causal relationship, that is, to analyze the direct or indirect impact of the change of one node on the change of other nodes. Through Bayesian network inference, calculate the conditional probability distribution and quantify the causal contribution. Through inference, generate the weight (contribution value) of each causal path and store it as causal contribution data for explaining the interaction between variables.

[0132] The dynamic feature data reflects the changes of humidity, thermal radiation, fracture expansion, etc. over time. Time series modeling predicts the future change trend of the target area by capturing the time dependence of these changes. A time convolutional network (TCN) is used to model the dynamic feature data, and its core lies in using one-dimensional convolution to capture long-term dependence relationships. The time series prediction results output by the TCN are combined with the causal contribution data to generate extended prediction data, marking the high-risk points and expansion ranges of the target area in the future.

[0133] The causal network clearly defines the causal relationships between variables. Especially in landslide risk monitoring, it can quantify the impacts of factors such as humidity changes, thermal radiation intensity, crack propagation, and ground surface settlement on landslide triggering, providing a scientific basis for risk prediction. Causal contribution data provides the influence weights between variables, helping to identify the dominant role of key inducing factors (such as abnormal humidity or abnormal thermal radiation) in regional changes. Dynamic feature data predicts the future expansion trend of the target area through time series modeling, providing a forward-looking reference for the dynamic monitoring of high-risk areas.

[0134] Preferably, the causal network constructs a causal graph based on humidity changes, thermal radiation intensity, crack propagation amplitude, and ground surface settlement amount in the high-resolution volume data model; the causal relationships of the causal graph are quantified through the Bayesian network inference method. The Bayesian network calculates the causal influence weights between nodes based on the prior probability and conditional probability of each node, generating causal contribution data.

[0135] In the present invention, based on key variables such as humidity changes, thermal radiation intensity, crack propagation amplitude, and ground surface settlement amount in the high-resolution volume data model, a causal network is constructed, and the causal relationships in the causal graph are quantified through the Bayesian network inference method, generating causal contribution data. This method can reveal the causal dependencies between variables and the quantified relationships of their interactions, providing a scientific basis for dynamic monitoring and risk assessment.

[0136] The causal network represents the causal relationships between variables in the form of a directed acyclic graph (DAG), where nodes represent variables and edges represent causal dependencies. Humidity changes, thermal radiation intensity, crack propagation amplitude, and ground surface settlement amount are used as the main nodes of the network, successively reflecting different stages and effects of landslide inducing factors.

[0137] Node definition: The value of each node is calculated through the high-resolution volume data model. For example, the value of the humidity change node is derived from the reflectance of the near-infrared band in the multi-spectral frame data, and the crack propagation amplitude is obtained through high-frequency feature component analysis. Causal graph generation: An initial causal graph is constructed through a structure learning algorithm (such as the PC algorithm) and manually adjusted in combination with domain knowledge to ensure that the edges and directions of the causal graph conform to the actual physical logic. For example, humidity changes affect crack propagation, and crack propagation further leads to ground surface settlement.

[0138] The quantification of causal relationships is achieved through Bayesian network inference. Based on the prior probability and conditional probability of each node, the weight of the edge (i.e., causal contribution) in the causal graph is calculated. The formula for Bayesian inference is:

[0139]

[0140] Among them, P(A|B) represents the conditional probability of A when event B occurs; P(B|A) represents the influence probability of event A on B; P(A) and P(B) represent the prior probabilities of nodes A and B.

[0141] The weight calculation formula is expressed as:

[0142] w A→B = P(B|A) - P(B)

[0143] Among them, w A→B represents the causal contribution of node A to node B, reflecting the part of the change amplitude of the dependent variable caused by the change of the independent variable.

[0144] The weight of each edge is stored as causal contribution data, which is used to explain the interaction between variables in the network. For example, the weight of the humidity change node can quantitatively describe its direct impact on crack propagation and its indirect impact on ground settlement through crack propagation. The generation of causal contribution data provides data support for the identification of landslide inducements and the optimization of monitoring strategies.

[0145] The constructed causal network reveals the causal relationship between humidity change, heat radiation intensity, crack propagation, and ground settlement amount, clarifies the dependence path between variables and their relative impacts; through Bayesian network reasoning, the interaction of landslide inducements is quantitatively described. For example, the contribution weight of humidity change to crack propagation is 0.6, and the contribution weight of crack propagation to ground settlement is 0.8; the causal contribution data provides a scientific basis for the dynamic adjustment of landslide monitoring, which can be used to optimize the monitoring priority of drones and concentrate resources to cover the areas with the greatest risk impact.

[0146] In an embodiment, in a high-risk landslide area, multi-modal data collected by drones is used to generate a high-resolution volume data model, and a causal network is constructed based on humidity change, heat radiation intensity, crack propagation amplitude, and ground settlement amount.

[0147] The nodes include humidity change, heat radiation intensity, crack propagation amplitude, and ground settlement amount; an initial causal graph is generated through the PC algorithm, and the direction is adjusted in combination with domain knowledge to make the humidity change directly affect crack propagation, and crack propagation further causes ground settlement.

[0148] The direct contribution of humidity change to crack propagation is quantified as w 湿度→裂缝 = 0.6 through Bayesian inference; the indirect contribution of crack propagation to ground settlement is calculated as w 湿度→裂缝→沉降 = 0.48;

[0149] The finally generated causal contribution data stores the weights of each path and the interaction between nodes.

[0150] The causal network reveals that humidity changes are the main inducement for crack propagation, while abnormal thermal radiation intensity mainly affects ground settlement. The causal contribution data provides a basis for UAV path planning, preferentially covering areas with drastic humidity changes.

[0151] Preferably, the time series modeling models the dynamic feature data through a temporal convolutional network, and the convolutional kernel of the temporal convolutional network is dynamically adjusted according to the length of the time series and the feature change rate to capture the temporal dependence relationship among humidity changes, crack propagation, and thermal radiation intensity; the output result of the temporal convolutional network is used to predict the future change trend of the target area and generate extended prediction data.

[0152] In the present invention, the time series modeling analyzes the dynamic feature data through a temporal convolutional network (TCN) to capture the temporal dependence relationship among humidity changes, crack propagation, and thermal radiation intensity, and predict the future change trend of the target area to generate extended prediction data. The core of the TCN is to capture the feature changes of long-distance dependencies in the time series through convolutional operations, and dynamically adjust the size and stride of the convolutional kernel to adapt to the feature change rate within different time windows.

[0153] The temporal convolutional network is a deep learning model based on one-dimensional convolution, which can efficiently process time series data. Compared with traditional recurrent neural networks (such as LSTM), the TCN achieves parallel computing through convolutional operations and has stronger long-term dependence capture ability.

[0154] Input layer: The dynamic feature data includes humidity changes, crack propagation, and thermal radiation intensity, which are arranged in chronological order to form a multi-dimensional time series matrix:

[0155] X = [x 1 , x 2 , …, x T

[0156] where X is the time series matrix, and x t represents the feature vector at time t, and T is the length of the time series.

[0157] Convolutional layer: The TCN extracts features from the input sequence through one-dimensional convolutional operations:

[0158]

[0159] O t represents the output at time t; W i ​represents the convolutional kernel weight; k represents the size of the convolutional kernel; σ represents the activation function (such as ReLU); b represents the bias term. The size of the convolutional kernel is dynamically adjusted according to the time series length and the feature change rate to ensure that enough time windows are covered, so as to capture the time-dependent relationship between humidity changes, crack propagation, and thermal radiation intensity.

[0160] TCN captures feature associations over long time spans through multi-layer stacked convolutional operations. For example, in the dynamic sequence of humidity changes, local anomalies may trigger crack propagation, and the accumulation of crack propagation may lead to abnormal changes in thermal radiation intensity. The receptive field of TCN expands as the number of layers increases, enabling it to capture these cross-time dependencies.

[0161] The output layer of TCN predicts the future change trend of the target area based on the captured time-dependent relationship. The output data includes the distribution of future humidity changes, the speed and direction of crack propagation, and the spatial distribution of thermal radiation intensity, which are used to label high-risk areas and the scope of expansion.

[0162] TCN efficiently captures the non-linear time-dependent relationship between humidity changes, crack propagation, and thermal radiation intensity by dynamically adjusting the size and stride of the convolutional kernel, significantly improving the accuracy of landslide precursor feature analysis. The extended prediction data generated by the model can clearly label the future change trend of the target area, providing accurate guidance for the dynamic monitoring and risk management of drones. Compared with recurrent neural networks, the parallel computing characteristics of TCN greatly shorten the training and inference time, making it suitable for real-time processing requirements in high-frequency drone video stream analysis.

[0163] Example: In a high-risk landslide area, based on the time series data of humidity changes, crack propagation, and thermal radiation intensity collected by drones, a temporal convolutional network is used for modeling to predict the change trend within the next 3 hours.

[0164] Input data: The time series matrix X, which includes the humidity change sequence, the crack propagation sequence, and the thermal radiation intensity sequence, with a data length of 72 (sampling interval of 5 minutes).

[0165] Modeling process: The size of the convolutional kernel of TCN is set to 3, the stride is 1, the number of convolutional layers is 6, and the receptive field covers 30 time steps (about 150 minutes). The model captures the lag effect of local anomalies in the humidity change sequence on crack propagation, and the cumulative effect of crack propagation on thermal radiation intensity anomalies.

[0166] Output results: Humidity change prediction, the humidity gradient will increase by 15% in the high-risk area within the next 3 hours. Crack propagation prediction, the average crack width will expand by 5 cm and the length will increase by 20 cm. Thermal radiation intensity prediction, the number of thermal anomaly points will increase by 50% and expand approximately 30 m downstream.

[0167] The extended prediction data generated by the temporal convolutional network clearly marks the future trends of crack propagation and thermal anomalies, providing accurate references for the monitoring task planning of drones and the early warning of high-risk areas.

[0168] Input the extended prediction data, and generate a dynamic risk heat map through the feature fusion algorithm; use the generative adversarial network to optimize the risk annotation of the target area, and optimize the flight path of the drone by combining the dynamic risk heat map with the reinforcement learning path planning algorithm, giving priority to monitoring the high-risk parts of the target area; feedback the newly added multi-modal acquisition data for dynamically updating the volume data model and the risk heat map to achieve a monitoring closed-loop.

[0169] The present invention utilizes the extended prediction data to generate a dynamic risk heat map through the feature fusion algorithm, combines the generative adversarial network (GAN) to optimize the annotation of high-risk areas, and at the same time adjusts the flight mission of the drone based on the reinforcement learning path planning algorithm. The newly added multi-modal acquisition data is fed back in real time for dynamically updating the volume data model and the risk heat map to form a monitoring closed-loop.

[0170] The dynamic risk heat map is generated by fusing key features such as humidity change, crack propagation, and thermal radiation intensity in the extended prediction data, and is used to visually annotate the risk distribution of the target area. The feature fusion algorithm is based on the cross-attention mechanism, calculating the weight relationship between multi-modal features to dynamically adjust the contribution degree of each feature in the heat map.

[0171] After the heat map is generated, the annotation accuracy is optimized through the generative adversarial network (GAN). The GAN includes a generator and a discriminator. The generator generates a high-resolution map of the risk area annotation, and the discriminator evaluates the authenticity of the generated map and guides the generator to improve. Generator: Based on a multi-layer convolutional neural network, input the initial data of the dynamic risk heat map to generate a high-resolution heat map. Discriminator: Classify the generated heat map through the convolutional layer, judge whether the risk area annotation is accurate, and feedback it to the generator for optimization.

[0172] Based on the dynamic risk heat map, the reinforcement learning path planning algorithm calculates the optimal flight path of the drone, giving priority to covering high-risk areas and dynamically adjusting the monitoring task. The algorithm guides the drone path planning through the reward function to achieve the maximum coverage at the minimum cost.

[0173] State space: The current position of the drone and the risk distribution in the dynamic risk heat map;

[0174] Action space: The next flight direction of the drone;

[0175] Reward function: The path covering high-risk areas obtains a higher reward.

[0176] The newly added real-time feedback of multi-modal acquisition data is used to dynamically update the volume data model and the risk heat map. The update of the volume data model optimizes the spatial distribution by fusing the newly added humidity feature and thermal radiation feature; the heat map update recomputes the feature weights in combination with the newly added data to ensure the real-time and accuracy of risk annotation.

[0177] The dynamic risk heat map generated by the feature fusion algorithm can accurately label the scope and its changing trend of the high-risk area, providing a scientific basis for the mission planning of the UAV; the generative adversarial network significantly improves the resolution and annotation accuracy of the heat map, making the boundary of the risk area clearer and the annotation of abnormal points more accurate; the reinforcement learning path planning algorithm optimizes the flight path of the UAV according to the dynamic risk heat map, preferentially covering the high-risk area to avoid resource waste; realizing the dynamic linkage of UAV acquisition, volume data update and risk heat map generation, ensuring the real-time and accuracy of landslide monitoring.

[0178] Preferably, as Figure 3 shown, the dynamic risk heat map fuses the humidity feature, thermal radiation feature and geometric feature from the high-resolution volume data model through the multi-modal interactive attention mechanism; the interactive attention mechanism dynamically adjusts the contributions of the humidity feature and thermal radiation feature in the annotation of the high-risk area by calculating the weight matrix between the features of each modality, generating a dynamic risk heat map with enhanced resolution and prominent focus.

[0179] The present invention dynamically generates a dynamic risk heat map with enhanced resolution and prominent focus by fusing the humidity feature, thermal radiation feature and geometric feature in the high-resolution volume data model through the multi-modal interactive attention mechanism. The core of the interactive attention mechanism lies in dynamically calculating the weight matrix of the features of each modality, enabling the contribution weights of different features to be adaptively adjusted in the annotation of the high-risk area, ensuring that the risk heat map can accurately highlight the key risk points and their changing trends in the target area.

[0180] The high-resolution volume data model provides the distribution data of the humidity feature, thermal radiation feature and geometric feature. Through the multi-modal interactive attention mechanism, these features are fused into a unified risk heat map. The multi-modal interactive attention mechanism calculates the weight matrix according to the spatial correlation of the features. The weight w i (x,y) of each feature T i is calculated by the following formula:

[0181]

[0182] where, Q i and K i represent the query vector and key vector of the i-th feature; d k represents the dimension size of the feature; w i (x,y) represents the feature T i(x, y) The weight at position (x, y), which reflects the contribution intensity of the feature.

[0183] Apply the calculated weight matrix w i (x, y) to the eigenvalue to generate the fused feature R(x, y):

[0184]

[0185] where N is the number of modalities, and R(x, y) represents the risk value at position (x, y) in the dynamic risk heat map.

[0186] The interactive attention mechanism dynamically adjusts the distribution of the weight matrix by calculating the spatial and semantic relationships of the features of each modality in real time. For example, when the local change of the humidity feature is extremely obvious, its weight will increase in the corresponding area and is preferentially used to label high-risk points; when the abnormal points of the thermal radiation feature and the geometric feature overlap, the weights in these areas will be further superimposed to strengthen the labeling of high-risk areas.

[0187] The fused dynamic risk heat map has the following characteristics: enhanced resolution, and the fusion of multi-modal features fully preserves the boundaries and texture details of the target area; prominent risk, the labeling of high-risk areas is clearer, and the feature expression of key areas is more significant.

[0188] The multi-modal interactive attention mechanism can accurately calculate the correlation between the humidity, thermal radiation, and geometric features, ensuring a reasonable distribution of the contributions of the features in high-risk areas. For example, the spatial overlap area between humidity anomalies and crack propagation is given a higher weight, improving the accuracy of heat map labeling. The weight matrix is updated in real time with the dynamic changes of the features, enabling the risk heat map to quickly adapt to the risk evolution trend of the target area. The generated dynamic risk heat map not only has a higher resolution, but also the boundaries of high-risk areas are clearer and the detail expression is more abundant.

[0189] Example: In a high-risk area of mountain landslides, a dynamic risk heat map is generated through the multi-modal interactive attention mechanism to label abnormal points of humidity change, high-value areas of thermal radiation, and areas with significant changes in geometric features.

[0190] Input data: Humidity feature, humidity gradient distribution data from the near-infrared band; thermal radiation feature, thermal radiation intensity distribution data from the short-wave infrared band; geometric feature, terrain change information from point cloud depth data.

[0191] Feature fusion process: Calculate the weight matrices of the humidity feature and the thermal radiation feature, dynamically adjust the contribution intensity of the humidity feature through the interactive attention mechanism to increase the weight in the humidity anomaly area; in the overlapping area of humidity anomaly and thermal radiation anomaly, the weights are further superimposed to ensure that the labeled high-risk areas are significantly prominent; apply the weights to the eigenvalues to generate the fused risk heat map.

[0192] Output heat map: High-risk areas are clearly marked, and the edges of cracks in areas with significant humidity changes are clearly visible; the high-value areas of thermal radiation show the boundary range of local settlement.

[0193] The dynamic risk heat map accurately marks the high-risk points in the target area and their changing trends, providing a scientific basis for the mission planning of drones and technical support for the risk early warning and response strategies of the area.

[0194] Preferably, as Figure 4 - Figure 5 shown, the reinforcement learning path planning algorithm calculates the optimal flight path based on the priority of high-risk areas in the dynamic risk heat map, combining the current position and flight parameters of the drone; the path planning algorithm optimizes the path covering high-risk areas through a reward function, adjusts the flight mission of the drone in real time, and generates task feedback data for updating the dynamic risk heat map.

[0195] The reinforcement learning path planning algorithm in the present invention takes the dynamic risk heat map as input, combines the current position and flight parameters of the drone, and calculates the optimal flight path. The algorithm guides the drone to preferentially cover high-risk areas through a reward function, and at the same time adjusts the flight mission in real time to adapt to the dynamically changing risk distribution. During the path planning process, the algorithm updates the dynamic risk heat map by generating task feedback data, realizing the closed-loop optimization of the drone monitoring task.

[0196] The dynamic risk heat map marks the high-risk points and their priorities in the target area. The reinforcement learning algorithm uses these priorities as the input for path planning and optimizes the flight path of the drone by maximizing the coverage of high-priority areas. State space: Composed of the current position of the drone, the flight direction, and the risk distribution in the dynamic risk heat map, defined as:

[0197] S t ={p t ,θ t ,R}

[0198] S t represents the state at time t; p t represents the position of the drone at time t; θ t represents the flight direction of the drone; R represents the dynamic risk heat map data.

[0199] Action space: Defined as the set A t .

[0200] The core of the reward function is to guide the drone path to maximize the coverage of high-risk areas while minimizing the flight cost (such as path length). The reward function can be expressed as:

[0201] R(S t ,A t ) = α·C high-risk -β·C cost

[0202] R(S t ,A t ) represents the reward value for taking action A t in state S t ; C high-risk represents the contribution value of covering high-risk areas; C cost represents the cost of the flight path; α and β represent weight parameters that adjust the balance between high-risk coverage and flight cost. The reward function calculates C high-risk through the risk values in the dynamic risk heat map. The drone will receive a higher reward for covering areas with higher risk values.

[0203] The algorithm is based on a reinforcement learning framework (such as Deep Q-Learning) and predicts the maximum cumulative reward that can be obtained by taking action A t ,A t in state S t through the value function Q(S t ). During the training process, the value function is updated by the following formula:

[0204]

[0205] η represents the learning rate; γ represents the discount factor that measures the importance of future rewards; A′ represents the possible action in the next step. Through repeated training, the algorithm can find the path that maximizes the cumulative reward, that is, the optimal flight path of the drone.

[0206] During the execution of path planning, the newly acquired multi-modal data (such as humidity changes, thermal radiation intensity) collected by the drone in real time is used to update the dynamic risk heat map. By recalculating the priorities of the risk areas, the algorithm can dynamically adjust the flight path to form a closed-loop optimization.

[0207] The algorithm can ensure that the drone preferentially covers the high-risk areas marked in the dynamic risk heat map, achieving efficient utilization of resources. Combining real-time task feedback and dynamic heat map updates, the algorithm can quickly adapt to the dynamic changes in the risk distribution of the target area, ensuring the timeliness of the monitoring task. The reward function balances high-risk coverage and flight cost, optimizes the flight path of the drone, and significantly improves the monitoring efficiency.

[0208] Example: In a high-risk landslide area, use the reinforcement learning path planning algorithm in combination with the dynamic risk heat map to optimize the monitoring path of the drone.

[0209] Input data: Dynamic risk heatmap R, marking high-risk areas of humidity change and heat radiation intensity; current position p of the drone t =(x,y) and flight direction θ t .

[0210] Path planning: The reward function is set as:

[0211] R(S t ,A t ) = 2·C high-risk -0.5·C cost

[0212] C high-risk represents the reward obtained for covering high-risk areas, calculated from the risk values in the heatmap; C cost represents the penalty for the path length. The algorithm predicts the optimal path through Deep Q-Learning to ensure that the drone covers the area with the highest risk value in the heatmap.

[0213] The drone collects new humidity change and heat radiation intensity data, updates the weight distribution of the dynamic risk heatmap; re-plans the path to include the newly added high-risk areas in the monitoring task.

[0214] The optimized path enables the drone to cover all the high-risk points marked in the heatmap, reducing the path length by 20% and shortening the monitoring time by 15 minutes. After the new data is fed back, the algorithm re-plans the path to cover the dynamically changing risk area.

[0215] The above are only examples of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. A method for analyzing drone video streams based on high-frequency digital information analysis, characterized in that: The following steps are involved: Based on the dynamic partitioning strategy, a drone is used to collect multimodal data in the target area, wherein the multimodal data includes multispectral video frame data and point cloud depth data, and data denoising and spatial alignment are performed to generate preliminary three-dimensional model data; Based on the preliminary 3D model data, a 3D modeling algorithm is used to generate a low-resolution volume data model, and a high-resolution volume data model is generated by refining the key areas. Multi-spectral video frame data is combined with point cloud depth data through multimodal data fusion technology to map the material features of the target area to the 3D model, and the high-frequency feature data and low-frequency contour data of the target area are extracted through feature extraction algorithms. Based on high-resolution volume data models and high-frequency feature data, a causal network is constructed and causal graph nodes are defined. Causal reasoning methods are used to quantify the causal relationship between nodes and generate causal contribution data. Time series modeling is used to model dynamic feature data, predict future change trends in the target area, and generate extended prediction data. Input the extended prediction data and generate a dynamic risk heat map through the feature fusion algorithm; Generative adversarial networks are used to optimize the risk annotation of target areas, and reinforcement learning path planning algorithms combined with dynamic risk heat maps are used to optimize the flight paths of drones, giving priority to monitoring high-risk parts of target areas. The newly added multimodal acquisition data is fed back to dynamically update the volume data model and risk heat map to achieve a closed-loop monitoring system.

2. The method according to claim 1, characterized in that The multispectral video frame data in the multimodal data includes a visible light band for capturing surface texture features of the target area; a near infrared band for monitoring moisture changes in soil and vegetation; and a short-wave infrared band for detecting local anomalies in the intensity of thermal radiation on the surface.

3. The method according to claim 2, characterized in that The multispectral video frame data in the multimodal data is denoised by an adaptive filter, which dynamically adjusts the filter parameters according to the band characteristics of the multispectral video frame data to eliminate the noise caused by light changes; the point cloud depth data is cleared of abnormal points by statistical filtering, and the statistical filtering removes depth data points that deviate from the expected range based on the local distribution characteristics of the depth information.

4. The method according to claim 1, characterized in that: The three-dimensional modeling algorithm includes generating a low-resolution volume data model using a neural radiation field model, wherein the neural radiation field model performs distributed encoding of geometric and spectral information of preliminary three-dimensional model data by ray sampling to generate a low-resolution three-dimensional representation; On the basis of the low-resolution volume data model, the cracks and settlement areas in the key areas are refined using a deep super-resolution network, which reconstructs detail information layer by layer through convolutional layers to generate a high-resolution volume data model.

5. The method according to claim 4, characterized in that The multimodal data fusion technology includes using a nonlinear interpolation method to map the humidity characteristics and thermal radiation characteristics of multispectral video frame data to a three-dimensional model generated by point cloud depth data. The nonlinear interpolation method associates material characteristics with geometric characteristics based on the coordinate relationship between multispectral data and point cloud depth data in three-dimensional space to generate a multidimensional material three-dimensional model with humidity gradient and thermal radiation intensity information.

6. The method according to claim 1, characterized in that The feature extraction algorithm includes performing a two-dimensional fast Fourier transform on the humidity and thermal radiation characteristics of the high-resolution volume data model to extract high-frequency components for capturing the details of crack extension and local settlement in the target area; further processing the extracted high-frequency components using a continuous wavelet transform, which refines the spatial representation of crack edges and settlement features by decomposing local changes in the high-frequency components.

7. The method according to claim 1, characterized in that The causal network constructs a causal graph based on humidity changes, thermal radiation intensity, crack expansion amplitude and surface settlement in a high-resolution volume data model; the causal relationship of the causal graph is quantified through a Bayesian network reasoning method. The Bayesian network calculates the causal influence weights between nodes based on the prior probability and conditional probability of each node to generate causal contribution data.

8. The method according to claim 7, characterized in that The time series modeling models dynamic feature data through a time convolution network. The convolution kernel of the time convolution network is dynamically adjusted according to the length of the time series and the feature change rate to capture the time dependency between humidity change, crack expansion and thermal radiation intensity. The output result of the time convolution network is used to predict the future change trend of the target area and generate extended prediction data.

9. The method according to claim 1, characterized in that: The dynamic risk heat map fuses humidity features, thermal radiation features and geometric features from a high-resolution volume data model through a multimodal interactive attention mechanism; The interactive attention mechanism dynamically adjusts the contribution of humidity features and thermal radiation features in the labeling of high-risk areas by calculating the weight matrix between the features of each modality, thereby generating a dynamic risk heat map with enhanced resolution and highlighted key points.

10. The method according to claim 9, characterized in that The reinforcement learning path planning algorithm calculates the optimal flight path based on the high-risk area priority of the dynamic risk heat map and the current position and flight parameters of the UAV; the path planning algorithm optimizes the path covering the high-risk area through the reward function, adjusts the flight mission of the UAV in real time, and generates mission feedback data for updating the dynamic risk heat map.

Citation Information

Patent Citations

  • Geological disaster monitoring method based on UAV video stream analysis

    CN117854256B

  • Unmanned aerial vehicle routing inspection fault diagnosis system based on integrated pod and AI technology

    CN107992067A

  • Flight situation accurate perception method based on multiple views

    CN111144290A

  • Vehicle path planning method, device and equipment and storage medium

    CN118408564A

  • Small unmanned aerial vehicle inspection data management method

    CN119322941A

Cited By

  • Intelligent image signal processing method and system based on multi-modal fusion

    CN120318603A

  • Terrain surveying and mapping method and system based on unmanned aerial vehicle, equipment and medium

    CN120947592A

  • Unmanned aerial vehicle-based topographic mapping method and system, device, medium

    CN120947592B