Unmanned aerial vehicle surveying and mapping information processing method and system of prefabricated model
By constructing a multimodal prefabricated model and dynamically adjusting the confidence weight, combined with the spatiotemporal coding change detection of U-Net and Transformer networks, the problems of low efficiency, poor adaptability and low accuracy in drone mapping technology are solved, and efficient and accurate mapping results are achieved.
Patent Information
- Application Number
- CN202510427658.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing drone surveying and mapping technology has shortcomings in data processing efficiency, model adaptability, data fusion accuracy, and positioning stability, and it is difficult to effectively deal with changes in complex terrain and dynamic environments, resulting in waste of computing resources, high false alarm rates and unstable positioning.
Build a multimodal prefabricated model, dynamically adjust the confidence weight, combine the U-Net network, Transformer network and cross-attention mechanism, build a spatiotemporal coding change detection model, realize the feature extraction and fusion of real-time data and historical data, and dynamically update the surveying and mapping model.
It significantly improves surveying and mapping efficiency, reduces false alarm rate, improves the accuracy and reliability of surveying and mapping results, enhances the detection ability of slight changes, and provides high-precision and adaptive surveying and mapping support.
Smart Images

Figure CN120352886A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of UAV mapping, and in particular to a method and system for processing UAV mapping information of prefabricated models. Background Art
[0002] In the technical field of UAV mapping, although certain progress has been made in the existing technology, there are still several significant defects and industry pain points, which restrict the further improvement of mapping efficiency and accuracy.
[0003] Traditional UAV mapping methods mainly rely on real-time full-scale data collection. When dealing with complex terrains or large areas, this method often performs redundant calculations on repeatedly mapped areas, resulting in waste of computing resources and low mapping efficiency. In addition, traditional static prefabricated models are unable to cope with dynamic changes in the natural environment, such as vegetation growth and seasonal changes, because these models cannot reflect these dynamic changes in a timely manner, leading to a high false alarm rate in change detection and affecting the accuracy and reliability of mapping results.
[0004] In terms of data fusion, existing multi-source data (including visible light, LiDAR, multispectral, etc.) fusion technologies are not yet perfect and it is difficult to achieve precise identification of subtle changes. For example, minor changes in key infrastructure such as ditches and pipelines are often overlooked or misjudged. This problem of insufficient fusion limits the application potential of UAV mapping in refined management. In addition, fixed markers (such as infrastructure like iron towers) as positioning reference points in mapping are prone to being affected by natural factors such as weather and light, resulting in unstable positioning results and further affecting the accuracy and reliability of mapping data.
[0005] In summary, the existing UAV mapping technology has many deficiencies in aspects such as data processing efficiency, model adaptability, data fusion accuracy, and positioning stability, and urgently needs to be improved and enhanced through technological innovation. Summary of the Invention
[0006] Embodiments of the present invention provide a method and system for processing UAV mapping information of prefabricated models to solve the following technical problems: The existing UAV mapping technology has many deficiencies in aspects such as data processing efficiency, model adaptability, data fusion accuracy, and positioning stability, and urgently needs to be improved and enhanced through technological innovation.
[0007] Embodiments of the present invention adopt the following technical solutions:
[0008] On the one hand, embodiments of the present invention provide a method for processing UAV mapping information of prefabricated models, the method comprising: constructing a multi-modal prefabricated model based on UAV historical data collected in a historical period;
[0009] Dynamically adjust the confidence weights of each element in the multimodal prefabricated model according to real-time environmental parameters to obtain a dynamic prefabricated model;
[0010] Based on the U-Net network, Transformer network, and cross-attention mechanism, construct a spatio-temporal encoding change detection model;
[0011] Extract features from the real-time data of the drone in the target area and the multimodal historical data loaded in the dynamic prefabricated model respectively to obtain real-time multimodal features and historical multimodal feature sequences;
[0012] Input the real-time multimodal features and historical multimodal feature sequences into the spatio-temporal encoding change detection model to obtain the mapping change information of the target area.
[0013] In a feasible implementation manner, construct a multimodal prefabricated model based on the historical data of the drone collected in the historical period, specifically including:
[0014] Obtain the historical data of the drone collected in multiple historical periods of the target area; wherein, the historical data of the drone at least includes: RGB image data, lidar point cloud data, and multispectral image data;
[0015] Input the RGB image data and the multispectral image data into the DeepLabV3+ network for target recognition and classification to obtain the fixed elements and dynamic elements in the target area; wherein, the dynamic elements at least include vegetation and water bodies;
[0016] Extract the normalized difference vegetation index NDVI and the normalized difference water index NDWI in each frame of multispectral image data;
[0017] According to the shooting dates of the current frame RGB image data and the multispectral image data, add spatio-temporal tags to the dynamic elements; wherein, the spatio-temporal tags at least include timestamps, season encodings, and meteorological codes;
[0018] Based on the nearest point search algorithm, register the lidar point cloud data collected in multiple historical periods, and construct a three-dimensional reference model of the target area in multiple historical periods according to the registered lidar point cloud data;
[0019] Label all the fixed elements and dynamic elements in the three-dimensional reference model, and embed the normalized difference vegetation index NDVI, the normalized difference water index NDWI, and the spatio-temporal tags into the dynamic elements of the three-dimensional reference model to obtain the multimodal prefabricated model.
[0020] In a feasible implementation, according to the real-time environmental parameters, the confidence weights of each element in the multi-modal prefabricated model are dynamically adjusted to obtain a dynamic prefabricated model, which specifically includes:
[0021] Cascade several layers of MLP networks and connect them to a Softmax network to construct a weight generation network and train it;
[0022] Collect the real-time environmental parameters of the target area through an environmental parameter sensor; wherein, the real-time environmental parameters at least include the current light intensity, the current meteorological code, the current season code, the current average temperature, and the current average humidity;
[0023] Construct a multi-dimensional input feature based on the real-time environmental parameters and the historical environmental parameters in the multi-modal prefabricated model and perform preprocessing;
[0024] Input the preprocessed multi-dimensional input feature into the weight generation network, and output the confidence weights of each element in the multi-modal prefabricated model at the current moment;
[0025] Based on the confidence weights, construct a weighted fusion model of the real-time data of the unmanned aerial vehicle and the multi-modal prefabricated model to obtain the dynamic prefabricated model.
[0026] In a feasible implementation, according to the real-time environmental parameters and the historical environmental parameters in the multi-modal prefabricated model, construct a multi-dimensional input feature and perform preprocessing, which specifically includes:
[0027] Construct a multi-dimensional input feature based on the real-time environmental parameters and the historical environmental parameters in the multi-modal prefabricated model: X = [L, W, ||S now -S hist ||2, ΔT avg , ΔH avg ; where L is the current light intensity, W is the current meteorological code, ||S now -S hist ||2 is the Euclidean distance between the current season code S now and the historical season code S in the multi-modal prefabricated model hist ; ΔT avg is the difference between the current average temperature and the historical average temperature in the same period in the multi-modal prefabricated model; ΔH avg is the difference between the current average humidity and the historical average humidity in the same period in the multi-modal prefabricated model;
[0028] Perform standardization processing on each feature channel in the multi-dimensional input feature to obtain the final input feature.
[0029] In a feasible implementation manner, based on the confidence weights, a weighted fusion model of the real-time data of the drone and the multi-modal prefabricated model is constructed to obtain the dynamic prefabricated model, which specifically includes:
[0030] According to Perform weighted fusion on the real-time data of the drone and the multi-modal prefabricated model to obtain the dynamic prefabricated model M r ;
[0031] where ω i represents the confidence weight of the i-th element, N represents the total number of fixed elements and dynamic elements in the multi-modal prefabricated model, Align is the data alignment operation, is the real-time data of the drone, is the historical data in the multi-modal prefabricated model;
[0032] Calculate the real-time confidence weights of each element according to the real-time environmental parameters collected each time, and substitute them into the dynamic prefabricated model to update the dynamic prefabricated model in the prefabricated model library in real time.
[0033] In a feasible implementation manner, based on the U-Net network, Transformer network and cross-attention mechanism, a spatio-temporal coding change detection model is constructed, which specifically includes:
[0034] Based on the U-Net convolutional neural network, a basic spatial network is constructed; among them, the basic spatial network includes an input layer, multiple encoder layers, a bridging layer, multiple decoder layers and an output layer;
[0035] Introduce an alternating convolutional layer with a preset dilation rate in the bridging layer to expand the receptive field; add channel attention at each skip connection to suppress invalid band interference, and obtain a spatial change detection branch network;
[0036] Based on the Transformer neural network architecture, a time change detection branch network is constructed; the time change detection branch network includes a patch embedding layer for dividing each frame of the input temporal image sequence into 16×16 patches;
[0037] Construct a cross-attention module for projecting the multi-scale feature map output by the spatial change detection branch network into a feature matrix, and fusing the feature matrix with the temporal features output by the time change detection branch network through a spatio-temporal attention calculation formula to output fused features;
[0038] Construct a change detection head for identifying the probability of change of each pixel according to the fused features and outputting a change probability map;
[0039] The spatial change detection branch network, the temporal change detection branch network, the cross-attention module, and the change detection head constitute the spatio-temporal encoded change detection model.
[0040] In a feasible implementation, feature extraction is respectively performed on the real-time data of the drone in the target area and the multi-modal historical data loaded in the dynamic prefabricated model to obtain real-time multi-modal features and historical multi-modal feature sequences, specifically including:
[0041] During actual drone mapping, real-time data of the drone in the target area is collected in real time by the drone; wherein, the real-time data of the drone at least includes: RGB image data, lidar point cloud data, and multi-spectral image data;
[0042] In the prefabricated model library, the dynamic prefabricated model corresponding to the target area is extracted, the multi-modal historical data therein is loaded, and the multi-modal historical data of each element is multiplied by the corresponding confidence weight to obtain multi-modal weighted historical data; wherein, the multi-modal historical data at least includes geometric data, spectral data, and spatio-temporal labels;
[0043] Feature encoding and alignment are respectively performed on the real-time data of the drone and the multi-modal weighted historical data to obtain real-time multi-modal features and historical multi-modal feature sequences; wherein, the feature encoding at least includes at least one or more encoding methods such as point cloud feature extraction, spectral time series encoding, and semantic embedding.
[0044] In a feasible implementation, the real-time multi-modal features and historical multi-modal feature sequences are input into the spatio-temporal encoded change detection model to obtain the mapping change information of the target area, specifically including:
[0045] The real-time multi-modal features and historical multi-modal feature sequences are simultaneously input into the spatio-temporal encoded change detection model for spatial multi-scale feature extraction and temporal feature extraction;
[0046] The multi-scale features are projected into a Query matrix through convolution, and the temporal features are extended into a Key-Value matrix through a fully connected layer;
[0047] Based on the spatio-temporal attention calculation formula in the spatio-temporal encoded change detection model, the fused features of the Query matrix and the Key-Value matrix are output;
[0048] The fused features are input into the change detection head to determine the probability of change for each pixel in the real-time RGB image data and the real-time multi-spectral image data, and a corresponding change probability map is generated;
[0049] In the change probability map, extract the mapping change information of the target area; wherein, the mapping change information at least includes the positions of important change areas in the target area and the corresponding change types.
[0050] In a feasible implementation manner, after obtaining the mapping change information of the target area, the method further includes:
[0051] When there are important change areas in the target area, insert the real-time data of the drone in the current frame into the dynamic prefabricated model, and replace the oldest historical frame in the dynamic prefabricated model;
[0052] For the unchanged areas in the target area, only update the timestamps of each element in the unchanged areas in the dynamic prefabricated model, and retain their geometric features and spectral features to achieve incremental update of the dynamic prefabricated model.
[0053] On the other hand, an embodiment of the present invention also provides a UAV mapping information processing system for a prefabricated model, characterized in that the system includes:
[0054] A prefabricated model construction module, configured to construct a multi-modal prefabricated model based on the historical UAV data collected in the historical period; dynamically adjust the confidence weights of each element in the multi-modal prefabricated model according to real-time environmental parameters to obtain a dynamic prefabricated model;
[0055] A real-time mapping module, configured to construct a spatio-temporal encoded change detection model based on the U-Net network, the Transformer network, and the cross-attention mechanism; respectively extract features from the real-time UAV data of the target area and the multi-modal historical data loaded in the dynamic prefabricated model to obtain real-time multi-modal features and historical multi-modal feature sequences; input the real-time multi-modal features and historical multi-modal feature sequences into the spatio-temporal encoded change detection model to obtain the mapping change information of the target area.
[0056] Compared with the prior art, a UAV mapping information processing method and system for a prefabricated model provided by an embodiment of the present invention have the following beneficial effects:
[0057] 1. Significantly improved efficiency: By introducing a prefabricated model, the present invention effectively reduces most of the repeated mapping areas in the UAV mapping process, thereby greatly reducing the unnecessary flights of the UAV during the mapping task, and also reducing the time consumption of a single task. This improvement not only significantly improves the efficiency of the mapping work, but also reduces the mapping cost, providing strong support for large-scale and high-frequency mapping operations.
[0058] 2. Powerful adaptive capability: This invention innovatively introduces a dynamic weight mechanism, breaking through the limitations of traditional static prefabricated models and realizing multi-source data fusion with environmental adaptation. This mechanism enables the drone to dynamically adjust the confidence weight of each element data in the prefabricated model according to environmental factors such as seasonal changes, thereby effectively reducing the false alarm rate caused by seasonal changes and improving the accuracy and reliability of surveying and mapping results.
[0059] 3. Innovation of spatiotemporal joint coding: The present invention also uses spatiotemporal joint coding technology, which can capture spatial details and temporal evolution rules at the same time, thus significantly improving the ability to detect small changes. This innovation provides more comprehensive and detailed data support for UAV mapping, which helps to discover potential environmental problems and changing trends.
[0060] In summary, the present invention provides a new solution for UAV mapping with high precision, low redundancy and strong adaptability through the deep coupling of prefabricated models and real-time perception. It not only performs well in efficiency, precision and adaptability, but also further improves the comprehensiveness and accuracy of mapping results through the innovative application of spatiotemporal joint coding technology. Compared with the existing technology, the present invention has significant advantages and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0062] Figure 1 A flow chart of a method for processing surveying and mapping information of a prefabricated model of an unmanned aerial vehicle provided in an embodiment of the present invention;
[0063] Figure 2 A schematic structural diagram of a prefabricated model UAV surveying and mapping information processing system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0064] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0065] An embodiment of the present invention provides a method for processing UAV mapping information of a prefabricated model, as follows Figure 1 As shown, the method for processing UAV mapping information of the prefabricated model specifically includes steps S101 - S105:
[0066] S101. Based on the UAV historical data collected in the historical period, construct a multi-modal prefabricated model.
[0067] Specifically, obtain the UAV historical data collected in multiple historical periods of the target area; wherein, the UAV historical data at least includes: RGB image data, lidar point cloud data, and multi-spectral image data.
[0068] Further, input the RGB image data and multi-spectral image data into the DeepLabV3+ network for target recognition and classification to obtain the fixed elements and dynamic elements in the target area; wherein, the dynamic elements at least include vegetation and water bodies. Extract the Normalized Difference Vegetation Index (NDVI) and Normalized Difference Water Index (NDWI) in each frame of multi-spectral image data.
[0069] Further, according to the shooting dates of the current frame RGB image data and multi-spectral image data, add spatio-temporal tags to the dynamic elements. Wherein, the spatio-temporal tags at least include time stamps, season codes, and meteorological codes.
[0070] In one embodiment, obtain 12 periods of UAV inspection data for a certain park in the past year. The UAV is equipped with a high-definition camera, a multi-spectral camera, and a lidar, and can simultaneously obtain the RGB image data, lidar point cloud data, and multi-spectral image data of the inspection area. Input the obtained historical data into the DeepLabV3+ network to extract the fixed markers (such as street lights, public facilities, buildings, etc.) and dynamic elements (such as vegetation, water bodies, etc.) in the park. Then, according to the inspection time of each period of data, add a time stamp t and a season code S ∈ R 4 (spring / summer / autumn / winter) to each element.
[0071] Further, based on the nearest point search algorithm, register the lidar point cloud data collected in multiple historical periods, and construct a three-dimensional reference model of the target area in multiple historical periods according to the registered lidar point cloud data.
[0072] Finally, label all the fixed elements and dynamic elements in the three-dimensional reference model, and embed the Normalized Difference Vegetation Index (NDVI), Normalized Difference Water Index (NDWI), and spatio-temporal tags into the dynamic elements of the three-dimensional reference model to obtain a multi-modal prefabricated model.
[0073] S102. Dynamically adjust the confidence weights of each element in the multi-modal prefabricated model according to the real-time environmental parameters to obtain a dynamic prefabricated model.
[0074] Specifically, a number of MLP networks are cascaded and then connected to a Softmax network to construct a weight generation network and train it. The weight generation network can be trained using the environmental parameters in the historical inspection data.
[0075] Furthermore, real-time environmental parameters of the target area are collected by an environmental parameter sensor; the real-time environmental parameters at least include the current light intensity, the current meteorological code, the current season code, the current average temperature, and the current average humidity.
[0076] Furthermore, based on the real-time environmental parameters and the historical environmental parameters in the multi-modal prefabricated model, multi-dimensional input features are constructed and preprocessed. The specific implementation method is as follows:
[0077] Based on the real-time environmental parameters and the historical environmental parameters in the multi-modal prefabricated model, multi-dimensional input features are constructed: X = [L, W, ||S now -S hist ||2, ΔT avg , ΔH avg . Where L is the current light intensity, W is the current meteorological code, ||S now -S hist ||2 is the Euclidean distance between the current season code S now and the historical season code S hist in the multi-modal prefabricated model; ΔT avg is the difference between the current average temperature and the historical average temperature in the same period in the multi-modal prefabricated model; ΔH avg is the difference between the current average humidity and the historical average humidity in the same period in the multi-modal prefabricated model.
[0078] Then, each feature channel in the multi-dimensional input features is separately normalized to obtain the final input features.
[0079] Furthermore, the preprocessed multi-dimensional input features are input into the trained weight generation network, and the confidence weights ω i = f(L, W, ||S now -S hist ||2) of each element in the multi-modal prefabricated model at the current moment are output. Where f is the objective function in the weight generation network.
[0080] Furthermore, based on the confidence weights, a weighted fusion model of the UAV real-time data and the multi-modal prefabricated model is constructed to obtain a dynamic prefabricated model. The specific implementation method is as follows:
[0081] According to the UAV real-time data and the multi-modal prefabricated model are weighted and fused to obtain a dynamic prefabricated model Mr 。
[0082] Among them, ω i represents the confidence weight of the i-th element, N represents the total number of fixed elements and dynamic elements in the multi-modal prefabricated model, Align is the data alignment operation, is the real-time data of the drone, is the historical data in the multi-modal prefabricated model.
[0083] As a feasible implementation method, the dynamic prefabricated models of each different target area are stored in the prefabricated model library and regularly maintained. In the daily maintenance of the prefabricated model library, the real-time confidence weights of each element are calculated according to the real-time environmental parameters collected each time and substituted into the dynamic prefabricated model to update the dynamic prefabricated model in the prefabricated model library in real time, so that the dynamic elements in the dynamic prefabricated model can adjust the confidence according to the changes of seasons, time and climate in a timely manner, avoiding false alarms caused by the system misinterpreting the huge changes in detected vegetation as abnormal changes due to seasonal changes during the actual mapping process.
[0084] S103. Based on the U-Net network, Transformer network and cross-attention mechanism, construct a spatio-temporal coding change detection model.
[0085] Specifically, based on the U-Net convolutional neural network, construct a basic spatial network; among them, the basic spatial network includes an input layer, multiple encoder layers, a bridging layer, multiple decoder layers and an output layer.
[0086] Introduce an alternating convolutional layer with a preset dilation rate in the bridging layer to expand the receptive field; add channel attention at the skip connection of each layer to suppress the interference of invalid bands, and improve the basic spatial network to obtain a spatial change detection branch network. The U-Net convolutional neural network can extract multi-level spatial features from single-temporal data and capture detailed textures and global semantics.
[0087] Furthermore, based on the Transformer neural network architecture, construct a time change detection branch network; the time change detection branch network includes a patch embedding layer for dividing each frame of the input temporal image sequence into 16×16 patches. The Transformer neural network can encode the temporal dependencies between multi-temporal data and model the change evolution law.
[0088] Furthermore, construct a cross-attention module for projecting the multi-scale feature map output by the spatial change detection branch network into a feature matrix, and fusing the feature matrix with the temporal features output by the time change detection branch network through the spatio-temporal attention calculation formula to output the fused features.
[0089] Further, a change detection head is constructed to identify the probability of change for each pixel based on the fused features and output a change probability map.
[0090] Finally, the spatial change detection branch network, the temporal change detection branch network, the cross-attention module, and the change detection head are combined to form a spatio-temporal encoded change detection model.
[0091] S104: Respectively extract features from the real-time UAV data of the target area and the multi-modal historical data loaded in the dynamic pre-built model to obtain real-time multi-modal features and historical multi-modal feature sequences.
[0092] Specifically, during actual UAV mapping, the real-time UAV data of the target area is collected in real time by the UAV; among them, the real-time UAV data at least includes: RGB image data, lidar point cloud data, and multi-spectral image data.
[0093] Further, extract the dynamic pre-built model corresponding to the target area from the pre-built model library, load the multi-modal historical data in it, and multiply the multi-modal historical data of each element by the corresponding confidence weight to obtain multi-modal weighted historical data. Among them, the multi-modal historical data at least includes geometric data, spectral data, and spatio-temporal labels.
[0094] Further, respectively perform feature encoding and alignment on the real-time UAV data and the multi-modal weighted historical data to obtain real-time multi-modal features and historical multi-modal feature sequences; among them, the feature encoding at least includes at least one or more encoding methods such as point cloud feature extraction, spectral time series encoding, and semantic embedding.
[0095] As a feasible implementation method, loading the multi-modal historical data of the target area from the pre-built model library includes:
[0096] Geometric data: Registered lidar point cloud data, containing information such as three-dimensional coordinates and reflection intensity;
[0097] Spectral data: Information such as multi-temporal NDVI, NDWI, and raster data;
[0098] Spatio-temporal labels: Information such as timestamps, season encodings, and meteorological codes.
[0099] The feature encoding process includes:
[0100] Point cloud feature extraction: Use PointNet++ to generate a point cloud global feature vector;
[0101] Spectral time series encoding: Input the NDVI / NDWI sequence into LSTM and output time series features;
[0102] Semantic embedding: Encode the spatio-temporal tags of the features through BERT to obtain the semantic encoding features of the spatio-temporal tags.
[0103] S105. Input the real-time multimodal features and the historical multimodal feature sequence into the spatio-temporal encoding change detection model to obtain the mapping change information of the target area.
[0104] Specifically, input the real-time multimodal features and the historical multimodal feature sequence into the spatio-temporal encoding change detection model at the same time to perform spatial multi-scale feature extraction and temporal feature extraction. Project the multi-scale features into a Query matrix through convolution, and expand the temporal features into a Key-Value matrix through a fully connected layer.
[0105] Furthermore, based on the spatio-temporal attention calculation formula in the spatio-temporal encoding change detection model, output the fused feature Attention(Q, K, V) of the Query matrix and the Key-Value matrix.
[0106] Furthermore, input the fused feature into the change detection head to determine the probability P change ∈[0, 1] that each pixel in the real-time RGB image data and the real-time multispectral image data has changed, and generate a corresponding change probability map.
[0107] Furthermore, in the change probability map, extract the mapping change information of the target area; wherein, the mapping change information at least includes the position of the important change area in the target area and the corresponding change type.
[0108] As a feasible implementation manner, in the change probability map, by comparing with the change probability threshold, the pixels exceeding the threshold can be determined, and then the change area composed of the pixels can be determined. According to the element type in this change area and the difference between its change parameter and the historical change parameter, the change type of this change area can be determined, such as types like abnormal vegetation growth and abnormal increase in river area.
[0109] In one embodiment, in one embodiment, if a certain park is inspected in summer, first call the spring dynamic prefabricated model of the park, and then input the real-time data features collected by the drone and the historical multimodal feature sequence extracted from the spring dynamic prefabricated model into the spatio-temporal encoding change detection model at the same time. The drone detects in real time that the NDVI of a certain area has increased by 0.3 (suspected rapid vegetation growth); the result obtained after spatio-temporal encoding analysis is: compared with the historical data of the same period, it is found that the NDVI increase in this area in previous years was only 0.1. The newly added shrub area is identified through the cross-attention mechanism. Insert the lidar point cloud data and multispectral data of this area into the dynamic prefabricated model for updating, and mark the change type as "high-risk vegetation encroachment area", triggering a pruning work order.
[0110] Furthermore, when there are important change regions in the target area, the real-time data of the drone in the current frame is inserted into the dynamic prefabricated model to replace the oldest historical frame in the dynamic prefabricated model. For the unchanged regions in the target area, only the timestamps of the elements in the unchanged regions in the dynamic prefabricated model are updated, and their geometric and spectral features are retained to achieve incremental updates of the dynamic prefabricated model.
[0111] In addition, an embodiment of the present invention also provides a UAV mapping information processing system for a prefabricated model, as Figure 2 shown. The UAV mapping information processing system 200 for the prefabricated model specifically includes:
[0112] A prefabricated model construction module 210, configured to construct a multi-modal prefabricated model based on the UAV historical data collected in the historical period; dynamically adjust the confidence weights of the elements in the multi-modal prefabricated model according to real-time environmental parameters to obtain a dynamic prefabricated model;
[0113] A real-time mapping module 220, configured to construct a spatio-temporal encoded change detection model based on a U-Net network, a Transformer network, and a cross-attention mechanism; respectively extract features from the real-time data of the UAV in the target area and the multi-modal historical data loaded in the dynamic prefabricated model to obtain real-time multi-modal features and a historical multi-modal feature sequence; input the real-time multi-modal features and the historical multi-modal feature sequence into the spatio-temporal encoded change detection model to obtain the mapping change information of the target area.
[0114] Each embodiment in the present invention is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0115] The above specifically describes certain embodiments of the present invention. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0116] The above are only the embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiments of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for processing unmanned aerial vehicle mapping information of a prefabricated model, characterized in that, The method includes: Constructing a multi-modal prefabricated model based on the historical data of drones collected in historical periods; Dynamically adjusting the confidence weights of each element in the multi-modal prefabricated model according to real-time environmental parameters to obtain a dynamic prefabricated model; Constructing a spatio-temporal coding change detection model based on the U-Net network, the Transformer network, and the cross-attention mechanism; Performing feature extraction on the real-time data of the drone in the target area and the multi-modal historical data loaded in the dynamic prefabricated model respectively to obtain real-time multi-modal features and a historical multi-modal feature sequence; Inputting the real-time multi-modal features and the historical multi-modal feature sequence into the spatio-temporal coding change detection model to obtain the mapping change information of the target area.
2. The method for processing unmanned aerial vehicle surveying and mapping information of a prefabricated model according to claim 1, wherein, Constructing a multi-modal prefabricated model based on the historical data of drones collected in historical periods, specifically including: Obtaining the historical data of drones collected in multiple historical periods of the target area; wherein, the historical data of the drones at least includes: RGB image data, lidar point cloud data, and multi-spectral image data; Inputting the RGB image data and the multi-spectral image data into the DeepLabV3+ network for target recognition and classification to obtain fixed elements and dynamic elements in the target area; wherein, the dynamic elements at least include vegetation and water bodies; Extracting the normalized difference vegetation index NDVI and the normalized difference water index NDWI in each frame of multi-spectral image data; Adding spatio-temporal labels to the dynamic elements according to the shooting dates of the current frame RGB influence data and the multi-spectral image data; wherein, the spatio-temporal labels at least include timestamps, season codes, and meteorological codes; Registering the lidar point cloud data collected in multiple historical periods based on the nearest point search algorithm, and constructing a three-dimensional reference model of the target area in multiple historical periods according to the registered lidar point cloud data; Labeling all the fixed elements and dynamic elements in the three-dimensional reference model, and embedding the normalized difference vegetation index NDVI, the normalized difference water index NDWI, and the spatio-temporal labels into the dynamic elements of the three-dimensional reference model to obtain the multi-modal prefabricated model.
3. A method for processing unmanned aerial vehicle mapping information of a prefabricated model according to claim 1, characterized in that, Dynamically adjusting the confidence weights of each element in the multi-modal prefabricated model according to real-time environmental parameters to obtain a dynamic prefabricated model, specifically including: Cascading several layers of MLP networks and connecting them to a Softmax network to construct a weight generation network and train it; Collecting real-time environmental parameters of the target area through an environmental parameter sensor; wherein, the real-time environmental parameters at least include the current light intensity, the current meteorological code, the current season code, the current average temperature, and the current average humidity; Constructing a multi-dimensional input feature according to the real-time environmental parameters and the historical environmental parameters in the multi-modal prefabricated model and performing preprocessing; Inputting the preprocessed multi-dimensional input feature into the weight generation network to output the confidence weights of each element in the multi-modal prefabricated model at the current moment; Constructing a weighted fusion model of the real-time data of the drone and the multi-modal prefabricated model based on the confidence weights to obtain the dynamic prefabricated model.
4. The method for processing unmanned aerial vehicle mapping information of a prefabricated model according to claim 3, wherein, Construct multi-dimensional input features based on the real-time environmental parameters and the historical environmental parameters in the multi-modal prefabricated model and perform preprocessing, specifically including: Construct a multi-dimensional input feature based on the real-time environmental parameters and the historical environmental parameters in the multi-modal prefabricated model: X = [L, W, ||S now -S hist ||2, ΔT avg , ΔH avg ; where L is the current light intensity, W is the current meteorological code, ||S now -S hist ||2 is the Euclidean distance between the current season code S now and the historical season code S in the multi-modal prefabricated model hist ; ΔT avg is the difference between the current average temperature and the historical average temperature in the same period in the multi-modal prefabricated model; ΔH avg is the difference between the current average humidity and the historical average humidity in the same period in the multi-modal prefabricated model; Perform standardization processing on each feature channel in the multi-dimensional input features to obtain the final input features.
5. The method for processing unmanned aerial vehicle mapping information of a prefabricated model according to claim 3, wherein, Based on the confidence weights, construct a weighted fusion model of the UAV real-time data and the multi-modal prefabricated model to obtain the dynamic prefabricated model, specifically including: According to weightedly fuse the real-time data of the drone and the multi-modal prefabricated model to obtain the dynamic prefabricated model M r ; Among them, ω i represents the confidence weight of the i-th element, N represents the total number of fixed elements and dynamic elements in the multi-modal prefabricated model, Align is the data alignment operation, is the real-time data of the drone, is the historical data in the multi-modal prefabricated model; Calculate the real-time confidence weights of each element according to the real-time environmental parameters collected each time, and substitute them into the dynamic prefabricated model to update the dynamic prefabricated model in the prefabricated model library in real time.
6. A method for processing unmanned aerial vehicle mapping information of a prefabricated model according to claim 1, characterized in that, Based on the U-Net network, Transformer network and cross-attention mechanism, construct a spatio-temporal coding change detection model, specifically including: Based on the U-Net convolutional neural network, construct a basic spatial network; wherein, the basic spatial network includes an input layer, multiple encoder layers, a bridging layer, multiple decoder layers and an output layer; Introduce an alternating convolutional layer with a preset dilation rate in the bridging layer to expand the receptive field; add channel attention at each skip connection to suppress ineffective band interference, and obtain a spatial change detection branch network; Based on the Transformer neural network architecture, construct a temporal change detection branch network; the temporal change detection branch network includes a patch embedding layer for dividing each frame of the input temporal image sequence into 16×16 patches; Construct a cross-attention module for projecting the multi-scale feature map output by the spatial change detection branch network into a feature matrix, and fusing the feature matrix with the temporal features output by the temporal change detection branch network through the spatio-temporal attention calculation formula to output fused features; Construct a change detection head for identifying the probability of change for each pixel based on the fused features and outputting a change probability map; The spatial change detection branch network, the temporal change detection branch network, the cross-attention module and the change detection head constitute the spatio-temporal coding change detection model.
7. A method for processing unmanned aerial vehicle mapping information of a prefabricated model according to claim 1, characterized in that, Extract features from the UAV real-time data of the target area and the multi-modal historical data loaded in the dynamic prefabricated model respectively to obtain real-time multi-modal features and historical multi-modal feature sequences, specifically including: During actual UAV surveying and mapping, the UAV real-time data of the target area is collected in real time by the UAV; wherein, the UAV real-time data at least includes: RGB image data, lidar point cloud data and multi-spectral image data; Extract the dynamic prefabricated model corresponding to the target area in the prefabricated model library, load the multi-modal historical data therein, and multiply the multi-modal historical data of each element by the corresponding confidence weight to obtain multi-modal weighted historical data; wherein, the multi-modal historical data at least includes geometric data, spectral data and spatio-temporal labels; Perform feature encoding and alignment on the UAV real-time data and the multi-modal weighted historical data respectively to obtain real-time multi-modal features and historical multi-modal feature sequences; wherein, the feature encoding at least includes at least one or more encoding methods such as point cloud feature extraction, spectral temporal coding and semantic embedding.
8. A method for processing unmanned aerial vehicle mapping information of a prefabricated model according to claim 1, characterized in that, Input the real-time multimodal features and the historical multimodal feature sequence into the spatio-temporal encoded change detection model to obtain the mapping change information of the target area, specifically including: Input the real-time multimodal features and the historical multimodal feature sequence into the spatio-temporal encoded change detection model simultaneously to perform spatial multi-scale feature extraction and temporal feature extraction; Project the multi-scale features into a Query matrix through convolution, and expand the temporal features into a Key-Value matrix through a fully connected layer; Based on the spatio-temporal attention calculation formula in the spatio-temporal encoded change detection model, output the fused features of the Query matrix and the Key-Value matrix; Input the fused features into a change detection head to determine the probability of change for each pixel in the real-time RGB image data and the real-time multispectral image data, and generate a corresponding change probability map; Extract the mapping change information of the target area from the change probability map; wherein, the mapping change information at least includes the position of the important change area in the target area and the corresponding change type.
9. The method for processing unmanned aerial vehicle mapping information of a prefabricated model according to claim 8, characterized in that, After obtaining the mapping change information of the target area, the method further includes: When there is an important change area in the target area, insert the real-time data of the drone in the current frame into the dynamic prefabricated model to replace the oldest historical frame in the dynamic prefabricated model; For the unchanged areas in the target area, only update the timestamps of each element in the unchanged areas in the dynamic prefabricated model, and retain their geometric features and spectral features to achieve incremental update of the dynamic prefabricated model.
10. An unmanned aerial vehicle mapping information processing system for prefabricated models, characterized in that, The system includes: A prefabricated model construction module, which is used to construct a multimodal prefabricated model based on the historical drone data collected in the historical period; dynamically adjust the confidence weight of each element in the multimodal prefabricated model according to the real-time environmental parameters to obtain a dynamic prefabricated model; A real-time mapping module, which is used to construct a spatio-temporal encoded change detection model based on the U-Net network, the Transformer network and the cross-attention mechanism; perform feature extraction on the real-time drone data of the target area and the multimodal historical data loaded in the dynamic prefabricated model respectively to obtain real-time multimodal features and a historical multimodal feature sequence; input the real-time multimodal features and the historical multimodal feature sequence into the spatio-temporal encoded change detection model to obtain the mapping change information of the target area.
Citation Information
Patent Citations
Dynamic remote sensing monitoring surveying and mapping method and system
CN116778104A
Deformation monitoring method and device based on unmanned aerial vehicle remote sensing technology
CN118424194A
Barley crop planting recommendation method
CN119066270A
Multi-modal space-time fusion target detection method and device and medium
CN119107638A
Fire prevention method in battery discharge prevention mode of eco-friendly car
KR1020250018246A
Cited By
Available cultivated land area evaluation method and system based on topographic mapping
CN120894414A
A method and system for assessing available arable land area based on topographic mapping
CN120894414B
Multi-mode full-autonomous inspection method and system for electric unmanned aerial vehicle
CN120909339A