A natural resource investigation and monitoring method, device and electronic equipment

Through the collaborative work of cloud-based devices and master control vehicles, the accuracy of natural resource survey and monitoring data in complex terrain has been improved, the problem of large errors in traditional methods has been solved, and detailed reports have been generated.

CN120528948BActive Publication Date: 2025-10-10CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511028421.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-10-10
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

In complex and dangerous areas such as mountains, hills, and plateaus, natural resource survey and monitoring work is limited by rugged terrain and a lack of ground communications and power facilities. This leads to large errors in traditional manual surveys and makes it difficult to achieve high-accuracy data collection.

Method used

A method based on the natural resource survey and monitoring system is adopted. The cloud device is used to obtain the mission planning plan through the decision support system. The main control vehicle collects raw multimodal data and generates bit stream data through spatiotemporal calibration processing and target detection. Finally, the data is decoded and processed on the cloud device to generate a detailed report.

Benefits of technology

It improves the temporal and spatial consistency of data, reduces false alarm and missed alarm rates, reduces transmission bandwidth requirements, ensures data integrity, and generates detailed natural resource survey and monitoring reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120528948B_ABST
    Figure CN120528948B_ABST
Patent Text Reader

Abstract

The application provides a natural resource investigation and monitoring method, device and electronic equipment, relates to the technical field of resource investigation and monitoring, and is based on a natural resource investigation and monitoring system. The natural resource investigation and monitoring system comprises a master control vehicle and a cloud device. The natural resource investigation and monitoring method comprises the following steps: obtaining a task planning scheme by using the cloud device through a decision support system; controlling the master control vehicle to collect original multi-modal data through the task planning scheme; performing space-time calibration processing on the original multi-modal data to obtain target multi-modal data; obtaining target detection image data by performing target detection on the target multi-modal data; performing data compression on the target detection image data to obtain bit stream data, and sending the bit stream data to the cloud device; and performing decoding processing on the bit stream data by using the cloud device to obtain a natural resource investigation and monitoring report. The application improves the accuracy of natural resource investigation and monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of resource survey and monitoring, and in particular to a natural resource survey and monitoring method, device and electronic equipment. Background Art

[0002] Natural resource surveys and monitoring facilitate the development of scientifically sound resource protection measures and development and utilization plans, promoting sustainable resource use. They are crucial for identifying environmental trends, assessing ecological risks, and implementing effective ecological protection measures. However, due to factors such as rugged terrain and a lack of ground communications and power infrastructure, conducting natural resource surveys and monitoring in complex and difficult areas such as mountains, hills, and plateaus has long been an industry challenge. Currently, this work primarily relies on traditional manual field surveys, which can lead to errors in the data collected. Summary of the Invention

[0003] The problem solved by the present invention is how to improve the accuracy of natural resource survey and monitoring.

[0004] To solve the above problems, the present invention provides a natural resource survey and monitoring method, device and electronic equipment.

[0005] In a first aspect, the present invention provides a natural resource survey and monitoring method based on a natural resource survey and monitoring system, wherein the natural resource survey and monitoring system includes a master control vehicle and a cloud device; the natural resource survey and monitoring method includes:

[0006] Obtaining a mission planning solution through a decision support system using the cloud device;

[0007] Controlling the master vehicle to collect original multimodal data through the mission planning scheme;

[0008] Performing spatiotemporal calibration processing on the original multimodal data to obtain target multimodal data, wherein the target multimodal data is used to represent multimodal data with spatial position information and time information;

[0009] Obtaining target detection image data by performing target detection on the target multimodal data;

[0010] Compressing the target detection image data to obtain bit stream data, and sending the bit stream data to the cloud device;

[0011] The cloud device is used to decode the bit stream data to obtain a natural resources survey and monitoring report.

[0012] Optionally, performing spatiotemporal calibration processing on the original multimodal data to obtain target multimodal data includes:

[0013] Obtaining spatially calibrated multimodal data by performing spatial calibration processing on the original multimodal data;

[0014] Based on a time synchronization protocol, time calibration processing is performed on the original multimodal data to obtain time-calibrated multimodal data, wherein the time synchronization protocol includes a PPS pulse per second signal, a GPRMC data transmission format, and a master-slave PTP protocol;

[0015] The target multimodal data is obtained according to the spatially calibrated multimodal data and the temporally calibrated multimodal data.

[0016] Optionally, the original multimodal data includes original surround camera data, original lidar data, and original hyperspectral data, and the spatially calibrated multimodal data is obtained by performing spatial calibration on the original multimodal data, including:

[0017] Performing unified coordinate system processing on the original surround-view camera data, the original lidar data, and the original hyperspectral data to obtain surround-view camera data in a unified coordinate system, lidar data in a unified coordinate system, and hyperspectral data in a unified coordinate system, wherein the original surround-view camera data includes first surround-view camera data, second surround-view camera data, third surround-view camera data, and fourth surround-view camera data;

[0018] The process of processing the original laser radar data into a unified coordinate system is as follows:

[0019] ,

[0020] Among them, P camera#1 is the coordinate of point P in the coordinate system of the first surround-view camera data, P LiDAR is the coordinate of point P in the coordinate system of the original lidar data, R is the lidar data rotation matrix, and T is the lidar data translation vector;

[0021] The process of processing the original hyperspectral data into a unified coordinate system is as follows:

[0022] ,

[0023] Among them, P H is the coordinate of point P in the coordinate system of the original hyperspectral data, R H is the hyperspectral data rotation matrix, T H is the hyperspectral data translation vector;

[0024] The process of processing the original surround view camera data into a unified coordinate system is as follows:

[0025] ,

[0026] ,

[0027] ,

[0028] Among them, P camera#2 is the coordinate of point P in the coordinate system of the second surround-view camera data, P camera#3 is the coordinate of point P in the coordinate system of the third surround-view camera data, P camera#4 is the coordinate of point P in the coordinate system of the fourth surround-view camera data, R camera#2 is the second surround camera rotation matrix, R camera#3 is the third surround camera rotation matrix, R camera#4 is the fourth surround camera rotation matrix, T camera#2 is the translation vector of the second surround camera, T camera#3 is the translation vector of the third surround camera, T camera#4 is the translation vector of the fourth surround camera;

[0029] Inputting GNSS positioning data, surround view camera data in the unified coordinate system, and lidar data in the unified coordinate system into a spatial precision positioning model for feature extraction and fusion to obtain predicted values ​​of the position, speed, and attitude of the master vehicle;

[0030] Wherein, the spatial precise positioning model is:

[0031] ,

[0032] ,

[0033] ,

[0034] ,

[0035] Among them, B 时空感知 B is the spatiotemporal perception feature processing branch of the spatial precise positioning model, 视觉感知 B is the visual perception feature processing branch of the spatial precise positioning model, 几何感知 is the geometric perception feature processing branch of the spatial precise positioning model, F 时空感知 is the feature vector obtained by the spatiotemporal perception feature processing branch, F 视觉感知 is the feature vector obtained by the visual perception feature processing branch, F 几何感知 is the feature vector obtained by the geometric perception feature processing branch, the GNSSRTK+IMU time series data is the GNSS positioning data, and F 联合 is to fuse feature vectors, and Concat is the feature fusion module;

[0036] The spatial calibration multimodal data is obtained according to the surround view camera data in the unified coordinate system, the lidar data in the unified coordinate system, the hyperspectral data in the unified coordinate system, and the predicted values ​​of the position, speed, and posture of the master vehicle.

[0037] Optionally, the target multimodal data includes target surround camera data, target lidar data, and target hyperspectral data, and the target detection image data is obtained by performing target detection on the target multimodal data, including:

[0038] Performing RGB image enhancement on the target surround view camera data according to the target surround view camera data, the target lidar data, and the target hyperspectral data to obtain enhanced RGB image data;

[0039] The enhanced RGB image data, the target lidar data and the target hyperspectral data are respectively subjected to feature extraction by a feature extraction backbone network to obtain low-level features, mid-level features and high-level features;

[0040] Inputting the low-level features, the mid-level features and the high-level features into a low-level feature fusion module, a mid-level feature fusion module and a high-level feature fusion module respectively to obtain low-level feature fusion data, mid-level feature fusion data and high-level feature fusion data;

[0041] The low-level feature fusion data, the mid-level feature fusion data and the high-level feature fusion data are input into a PAN feature pyramid module to obtain the target detection image data.

[0042] Optionally, performing RGB image enhancement on the target surround view camera data according to the target surround view camera data, the target lidar data, and the target hyperspectral data to obtain enhanced RGB image data includes:

[0043] Extracting features from the target surround camera data through a convolutional layer to obtain RGB image features;

[0044] Decomposing the RGB image features using a discrete wavelet transform module to obtain target high-frequency features and target low-frequency features;

[0045] Among them, the target high-frequency features are:

[0046] ,

[0047] Among them, F RGB is the RGB image feature, F High ' is the target high frequency feature, F Highis the high-frequency feature, DWT()_High is the high-frequency feature obtained by decomposing the discrete wavelet transform module, and Conv() is the convolution operation;

[0048] Wherein, the target low-frequency feature is:

[0049] ,

[0050] Among them, F Low ' is the target low-frequency feature, F Low It is a low-frequency feature, and DWT()_Low is a low-frequency feature obtained by decomposing it using the discrete wavelet transform module;

[0051] Extracting features of the target lidar data through a convolutional layer to obtain lidar image features;

[0052] Inputting the target high-frequency feature and the lidar image feature into a first adaptive attention mechanism fusion module to obtain a first fusion feature;

[0053] The first fusion feature is:

[0054] ,

[0055] Among them, F Fusion1 is the first fusion feature, fusion module #1 is the first adaptive attention mechanism fusion module, F LiDAR is the laser radar image feature;

[0056] Extracting features of the target hyperspectral data through a convolutional layer to obtain hyperspectral image features;

[0057] Inputting the target low-frequency feature and the hyperspectral image feature into a second adaptive attention mechanism fusion module to obtain a second fused feature;

[0058] The second fusion feature is:

[0059] ,

[0060] Among them, F Fusion2 is the second fusion feature, fusion module #2 is the second adaptive attention mechanism fusion module, F Hyperspectral is the hyperspectral image feature;

[0061] The enhanced RGB image data is obtained by fusing the first fusion feature and the second fusion feature and processing them through a convolution layer.

[0062] Optionally, compressing the target detection image data to obtain bit stream data includes:

[0063] Performing mask processing on the target detection image data to obtain filtered target detection image data;

[0064] The bit stream data is obtained by compressing the filtered target detection image data through an entropy coding module.

[0065] Optionally, before compressing the filtered target detection image data by the entropy coding module to obtain the bit stream data, the method further includes:

[0066] Obtain meteorological data, cloud images, and sky images;

[0067] Inputting the meteorological data, the cloud image and the sky image into a satellite communication quality prediction model to obtain an effective communication window;

[0068] The data compression rate of the entropy coding module is adjusted according to the effective communication window.

[0069] Optionally, the natural resource survey and monitoring system further includes a drone and a quadruped robot, wherein the drone and the quadruped robot are provided with a sensor module including an RGB camera, a hyperspectral camera, and a laser radar, and the quadruped robot is further provided with a sampling manipulator arm; the natural resource survey and monitoring method further includes:

[0070] Obtaining multimodal data of the drone through the RGB camera, the hyperspectral camera, and the lidar of the drone;

[0071] Obtaining multimodal data of the quadruped robot through the sampling manipulator arm of the quadruped robot, the RGB camera, the hyperspectral camera and the laser radar;

[0072] The natural resources survey and monitoring report is obtained based on the original multimodal data, the drone multimodal data and the quadruped robot multimodal data.

[0073] In a second aspect, the present invention provides a natural resource survey and monitoring device, based on a natural resource survey and monitoring system, wherein the natural resource survey and monitoring system includes a master control vehicle and a cloud device; the natural resource survey and monitoring device includes:

[0074] A mission planning scheme acquisition module is used to obtain a mission planning scheme through a decision support system using the cloud device;

[0075] An original multimodal data acquisition module, configured to control the master vehicle to collect original multimodal data through the mission planning scheme;

[0076] The spatio-temporal calibration processing module is configured to perform spatio-temporal calibration processing on the original multi-modal data to obtain target multi-modal data, wherein the target multi-modal data is used to represent multi-modal data with spatial position information and time information.

[0077] The target detection module is configured to perform target detection on the target multi-modal data to obtain target detection image data.

[0078] The data compression module is configured to perform data compression on the target detection image data to obtain bit stream data, and send the bit stream data to the cloud device.

[0079] The natural resource investigation and monitoring report acquisition module is configured to decode the bit stream data by using the cloud device to obtain a natural resource investigation and monitoring report.

[0080] In a third aspect, the present application provides an electronic device comprising a memory and a processor.

[0081] The memory is configured to store a computer program.

[0082] The processor is configured to implement the natural resource investigation and monitoring method according to the first aspect when executing the computer program.

[0083] The natural resource investigation and monitoring method, device and electronic device have the following advantages: the cloud device is used to obtain a task planning scheme through a decision support system, and a master control vehicle is controlled to collect original multi-modal data. The spatio-temporal calibration processing on the original multi-modal data can ensure the consistency of data from different sources in time and space, which helps to improve the accuracy of subsequent data analysis. The target detection on the target multi-modal data can more accurately identify objects or features of interest, which not only improves the detection accuracy, but also reduces the false positive rate and the false negative rate. The data compression on the target detection image data to obtain bit stream data and the sending of the bit stream data to the cloud device reduce the data volume, which reduces the demand for transmission bandwidth and ensures the integrity of the data in the case of poor network conditions in field work. The powerful computing capacity and storage resources of the cloud device are used to decode the compressed bit stream data and generate a detailed natural resource investigation and monitoring report, which effectively improves the accuracy of natural resource investigation and monitoring. BRIEF DESCRIPTION OF DRAWINGS

[0084] Figure 1 FIG. 1 is a flowchart of a natural resource investigation and monitoring method according to an embodiment of the present application;

[0085] Figure 2 FIG. 3 is a schematic diagram of a spatio-temporal calibration processing module according to an embodiment of the present application;

[0086] Figure 3 A schematic diagram of a spatial calibration process according to an embodiment of the present invention;

[0087] Figure 4 A schematic diagram of an RGB image enhancement process according to an embodiment of the present invention;

[0088] Figure 5 A schematic diagram of a process for obtaining a valid communication window according to an embodiment of the present invention;

[0089] Figure 6 This is a structural diagram of a natural resource survey and monitoring device according to an embodiment of the present invention;

[0090] Figure 7 The figure is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0091] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0092] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0093] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to"; the term "based on" means "based at least in part on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0094] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0095] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0096] In response to the problems existing in the above-mentioned related technologies, this embodiment provides a natural resource survey and monitoring method, device and electronic equipment.

[0097] like Figure 1 As shown, an embodiment of the present invention provides a natural resource survey and monitoring method, characterized in that it is based on a natural resource survey and monitoring system, the natural resource survey and monitoring system includes a master control vehicle and a cloud device; the natural resource survey and monitoring method includes:

[0098] Step 110: Utilize the cloud device to obtain a mission planning solution through a decision support system.

[0099] In some more specific embodiments, the decision support system builds an external knowledge base for natural resource survey and monitoring based on relevant materials. This knowledge base is used by the large model to reference relevant background knowledge and provide appropriate responses. Input includes user questions and relevant processing results transmitted by the vehicle system. For user questions, the system connects to the external knowledge base using RAG technology, injects the retrieved knowledge into a prompt template, and then encodes and inputs it into the large model. RAG (Retrieval-Augmented Generation) is a method that combines retrieval and generation techniques. It retrieves relevant information from an external knowledge base and injects it into the input (prompt template) of the generation model, thereby improving the relevance and accuracy of the generated content. Real-time sensor data is encoded using an encoder and also input into the large model for computational processing. The output is the natural resource survey and monitoring report required by the user. Furthermore, a multimodal corpus dedicated to natural resource survey and monitoring is constructed and used, and a large model dedicated to natural resource survey and monitoring is trained using LoRA (Low-Rank Adaptation) fine-tuning technology. LoRA is an efficient parameter fine-tuning method, particularly suitable for adapting large pre-trained models in specific domains. The multimodal corpus dedicated to natural resource surveys and monitoring includes natural resource-related textual materials (guides, local chronicles), maps, historical satellite images, on-site photos collected by drones, real-time data from vehicle-mounted sensors, data processing module results, and historical mission records. All data is linked by labeling them with "time + location."

[0100] Step 120 : Control the master vehicle to collect original multimodal data through the mission planning solution.

[0101] Specifically, a cloud-based decision support system generates a mission plan based on the user's question request. There are two main types of mission plans. The first is for automated, multi-dimensional on-site verification of surface change patterns in natural resource surveys and monitoring. The mission plan primarily includes the survey sequence and route for the master vehicle, the location and angle of data collection by the master vehicle, the flight path and frequency of drone photography, and the type of ground object detection model. The second is for the natural resource survey and monitoring needs of both on-site sampling (point-based) and large-scale remote sensing monitoring (area-based). The mission plan primarily includes the types of sensors that should be equipped on drones and quadruped robots, as well as the locations of sampling points.

[0102] Step 130 : Performing spatiotemporal calibration processing on the original multimodal data to obtain target multimodal data, wherein the target multimodal data is used to represent multimodal data with spatial position information and time information.

[0103] Specifically, users receive a mission plan based on the survey and monitoring needs of the area to be inspected, and simultaneously plan and formulate the master vehicle's work path on the in-vehicle command screen. The master vehicle's rooftop sensor module is equipped with numerous sensor units, including four surround-view cameras for the front, rear, left, and right sides, a 32-beam LiDAR, and a lightweight hyperspectral camera. A deep neural network-based, end-to-end spatial positioning model with adaptive fusion of multivariate data was designed to meet the master vehicle's high-precision spatial positioning requirements in these complex work scenarios. It is crucial for sensors to coordinate with each other to achieve simultaneous data acquisition based on highly accurate time calibration information. Time calibration enables synchronization of time information among multiple sensors on the master vehicle.

[0104] Step 140 : Obtain target detection image data by performing target detection on the target multimodal data.

[0105] Specifically, an image enhancement method was designed to address the problem that surround-view cameras are easily affected by lighting conditions, resulting in blurred and unusable images. Furthermore, based on this image enhancement, a deep learning-based multimodal data fusion method for detecting ground objects based on surface change patterns was designed, leveraging the multimodal data captured by the roof-mounted sensor. This method enhances the system's fine-grained detection capabilities on the ground and outputs the object's category and location.

[0106] Step 150: compress the target detection image data to obtain bit stream data, and send the bit stream data to the cloud device.

[0107] Specifically, to address the issue of insufficient storage space and inefficient transmission due to the large amount of data collected by the master vehicle's roof sensor module, a timestamp- and space-stamp-driven multimodal data collaborative compression method was designed. This method first performs the first half of the processing in the vehicle's onboard computer, processing the image data to produce a compressed bitstream. This bitstream is then transmitted to the cloud via a multimodal communication module, where it is decompressed to recover the data.

[0108] Step 160: Utilize the cloud device to decode the bit stream data to obtain a natural resources survey and monitoring report.

[0109] Specifically, data related to natural resource surveys and monitoring is compressed and transmitted on the vehicle side. After further processing and analysis in a large cloud-based model, the final results are fed back to the vehicle-side user. There are two types of natural resource survey and monitoring reports: one that includes the type, location, and area of ​​surface changes observed during natural resource surveys and monitoring; and one that includes large-scale remote sensing monitoring results, such as water quality distribution.

[0110] In this embodiment, a cloud device is used to obtain a mission planning plan through a decision support system, and the main control vehicle is controlled to collect raw multimodal data. The spatiotemporal calibration processing of the raw multimodal data can ensure the temporal and spatial consistency of data from different sources, which helps to improve the accuracy of subsequent data analysis. By performing target detection on the target multimodal data, the object or feature of interest can be identified more accurately, which not only improves the accuracy of detection, but also may reduce the false alarm rate and missed alarm rate. The target detection image data is compressed to obtain bit stream data, and sent to the cloud device, which reduces the amount of data. In the case of poor network conditions in field operations, it reduces the demand for transmission bandwidth while ensuring the integrity of the data. The powerful computing power and storage resources of the cloud device are used to decode the compressed bit stream data and generate a detailed natural resource survey and monitoring report, effectively improving the accuracy of natural resource survey and monitoring.

[0111] Optionally, performing spatiotemporal calibration processing on the original multimodal data to obtain target multimodal data includes:

[0112] Obtaining spatially calibrated multimodal data by performing spatial calibration processing on the original multimodal data;

[0113] Based on a time synchronization protocol, time calibration processing is performed on the original multimodal data to obtain time-calibrated multimodal data, wherein the time synchronization protocol includes a PPS pulse per second signal, a GPRMC data transmission format, and a master-slave PTP protocol;

[0114] The target multimodal data is obtained according to the spatially calibrated multimodal data and the temporally calibrated multimodal data.

[0115] In some more specific embodiments, Figure 2 As shown, the master vehicle includes a positioning module, an onboard computer, and a sensor module. The positioning module includes an IMU and a GNSS positioning module. The master vehicle's rooftop sensor module is equipped with numerous sensor units, including four surround-view cameras for the front, back, and left and right, a 32-beam LiDAR, and a lightweight hyperspectral camera. The positioning module receives and resolves clock signals, and uses PPS+GPRMC to enable the onboard computer to obtain ultra-high-precision clock synchronization information. Other devices, such as the surround-view cameras, hyperspectral cameras, and LiDAR, use the master-slave PTP protocol to synchronize time with the onboard computer. Master-slave PTP (Precision Time Protocol) is a protocol for synchronizing clocks within a network. Defined by the IEEE 1588 standard, it aims to provide sub-microsecond clock synchronization accuracy and is suitable for time-sensitive applications such as measurement and control systems. PPS (Pulse Per Second) is a standard for providing pulse-per-second signals. These signals are typically generated by devices such as GPS receivers, emitting a precise pulse every second to synchronize clocks or other devices. GPRMC (Recommended Minimum Specific GPS / Transit Data) is a sentence format in the NMEA 0183 protocol used to transmit the most basic GPS data, including time, date, latitude, longitude, speed, and heading information.

[0116] Specifically, triggered time soft synchronization is used between the surround-view camera, lidar, and hyperspectral camera to achieve synchronized data acquisition between the different sensors, ensuring highly temporally aligned data. When the onboard computer captures the first frame of imagery from the first surround-view camera, it triggers the issuance of a PPS signal. When the other surround-view cameras, lidar, and hyperspectral cameras receive the PPS signal, data acquisition is triggered on the rising edge of the signal. First, the surround-view camera, lidar, and hyperspectral camera deployed on the roof of the master vehicle subsystem use this triggered soft synchronization method to synchronize data acquisition and store the corresponding timestamp information. Timestamp information is stored in the format of "year_month_day_hour_minute_second_microsecond"; images captured by the four cameras on the roof (front, rear, left, and right) are named in the format of "surround view camera 1_id.jpg," "surround view camera 2_id.jpg," "surround view camera 3_id.jpg," and "surround view camera 4_id.jpg," respectively, where the id represents the number of the captured image; lidar data is named "lidar_id.pcap"; and hyperspectral images are named "hyperspectral camera_id.hdr" and "hyperspectral camera_id.spe." The storage and naming format for data collected by the drone subsystem and quadruped robot subsystem is similar to the above. The original multimodal data is updated using the predicted values ​​of position, velocity, and attitude obtained from the spatial calibration multimodal data, and device time information is synchronized using the temporal calibration multimodal data to obtain the target multimodal data.

[0117] In this optional embodiment, in terms of time calibration of the master control vehicle, the PPS+GPRMC+master-slave PTP protocol is adopted to achieve time information synchronization between devices, reduce data analysis errors caused by time asynchrony, and establish a time basis for data synchronization collection.

[0118] Optionally, the original multimodal data includes original surround camera data, original lidar data, and original hyperspectral data, and the spatially calibrated multimodal data is obtained by performing spatial calibration on the original multimodal data, including:

[0119] Performing unified coordinate system processing on the original surround-view camera data, the original lidar data, and the original hyperspectral data to obtain surround-view camera data in a unified coordinate system, lidar data in a unified coordinate system, and hyperspectral data in a unified coordinate system, wherein the original surround-view camera data includes first surround-view camera data, second surround-view camera data, third surround-view camera data, and fourth surround-view camera data;

[0120] The process of processing the original laser radar data into a unified coordinate system is as follows:

[0121] ,

[0122] Among them, P camera#1 is the coordinate of point P in the coordinate system of the first surround-view camera data, P LiDAR is the coordinate of point P in the coordinate system of the original lidar data, R is the lidar data rotation matrix, and T is the lidar data translation vector;

[0123] The process of processing the original hyperspectral data into a unified coordinate system is as follows:

[0124] ,

[0125] Among them, P H is the coordinate of point P in the coordinate system of the original hyperspectral data, R H is the hyperspectral data rotation matrix, T H is the hyperspectral data translation vector;

[0126] The process of processing the original surround view camera data into a unified coordinate system is as follows:

[0127] ,

[0128] ,

[0129] ,

[0130] Among them, P camera#2 is the coordinate of point P in the coordinate system of the second surround-view camera data, P camera#3 is the coordinate of point P in the coordinate system of the third surround-view camera data, P camera#4 is the coordinate of point P in the coordinate system of the fourth surround-view camera data, R camera#2 is the second surround camera rotation matrix, R camera#3 is the third surround camera rotation matrix, R camera#4 is the fourth surround camera rotation matrix, T camera#2 is the translation vector of the second surround camera, T camera#3 is the translation vector of the third surround camera, T camera#4 is the translation vector of the fourth surround camera;

[0131] Inputting GNSS positioning data, surround view camera data in the unified coordinate system, and lidar data in the unified coordinate system into a spatial precision positioning model for feature extraction and fusion to obtain predicted values ​​of the position, speed, and attitude of the master vehicle;

[0132] Wherein, the spatial precise positioning model is:

[0133] ,

[0134] ,

[0135] ,

[0136] ,

[0137] Among them, B 时空感知 B is the spatiotemporal perception feature processing branch of the spatial precise positioning model, 视觉感知 B is the visual perception feature processing branch of the spatial precise positioning model, 几何感知 is the geometric perception feature processing branch of the spatial precise positioning model, F 时空感知 is the feature vector obtained by the spatiotemporal perception feature processing branch, F 视觉感知 is the feature vector obtained by the visual perception feature processing branch, F 几何感知 is the feature vector obtained by the geometric perception feature processing branch, GNSSRTK+IMU time series data is the GNSS positioning data, F 联合 is to fuse feature vectors, and Concat is the feature fusion module;

[0138] The spatial calibration multimodal data is obtained according to the surround view camera data in the unified coordinate system, the lidar data in the unified coordinate system, the hyperspectral data in the unified coordinate system, and the predicted values ​​of the position, speed, and posture of the master vehicle.

[0139] Specifically, if Figure 3 As shown in the figure, the spatial position relationship between the surround-view camera, lidar, and hyperspectral camera is determined by solving the rotation matrix R and translation matrix T in the following formula, so that the data collected by different sensors are also highly aligned in space. In terms of the precise spatial positioning of the master vehicle, based on GNSS RTK+IMU positioning, GNSS RTK, IMU, surround-view camera images, and lidar multivariate data are used as input to establish an end-to-end spatial precision positioning model based on deep neural networks with adaptive fusion of multivariate data, further improving the positioning accuracy of the master vehicle in complex environments. GNSS RTK is combined with IMU. GNSS RTK can provide long-term stable absolute position information, but it may be affected by environmental factors and temporarily lose lock. IMU is not affected by these environmental factors, but it will have large cumulative errors during long-term operation. The combination of the two can make up for each other's shortcomings.

[0140] The model first extracts features from GNSS RTK+IMU time series data, surround-view camera image data, and laser radar point cloud projection image respectively, and is divided into three independent processing branches, namely space-time perception branch, visual perception branch and geometric perception branch. In the space-time perception branch, the GNSS RTK+IMU time series data is organized into a size of (T x 7). Among them, T represents the time window size, which is 10 by default; 7 represents the longitude, latitude, height obtained by GNSS RTK and the three-axis acceleration and angular velocity of IMU. Multi-layer BiLSTM is used to effectively capture forward historical dependence and backward future dependence, and multi-head attention mechanism is used to dynamically weight the importance of different time steps. Multi-layer BiLSTM (Bidirectional Long Short-Term Memory Network) is a deep learning architecture that enhances the model's feature extraction capability for sequence data by stacking multiple BiLSTM layers. When the input data size is 10 x 7, after two bidirectional LSTM and a multi-head attention layer, a 1 x 256 feature vector is obtained, wherein:

[0141] ,

[0142] Among them, B 时空感知 is the space-time perception feature processing branch of the spatial precise positioning model, F 时空感知 is the feature vector obtained by processing the space-time perception feature processing branch, GNSS RTK+IMU time series data is the GNSS positioning data, and Multiheadattention (multi-head attention mechanism) is a very important technology in the field of modern deep learning.

[0143] In the visual perception branch and the geometric perception branch, for the surround-view camera image and the laser radar point cloud projection image, a feature extraction branch with RepVGGBlock+RepBlock as the backbone network is used to effectively extract features. The multi-branch structure in it enables the model to more comprehensively capture the features of the input data. RepVGGBlock is a simple convolution block based on VGG style, but an additional branch (such as 1x1 convolution and identity mapping) is introduced during training to enhance the model's expression ability. RepBlock is a more general reparameterization module, usually composed of multiple sub-modules. Among them, the original size of the surround-view camera image and the laser radar point cloud projection image is scaled to 3 x 640 x 640, wherein:

[0144] ,

[0145] ,

[0146] Among them, B 视觉感知 is the visual perception feature processing branch of the spatial precise positioning model, and B几何感知 is the geometric perception feature processing branch of the spatial precise positioning model, F 视觉感知 is the feature vector obtained by the visual perception feature processing branch, F 几何感知 This is the feature vector generated by the geometric-aware feature processing branch. Next, after extracting from three independent processing branches, the feature vectors output by each branch are concatenated to form a joint feature with a size of 1×768. This feature vector F_joint is then dimensionalized by a fully connected layer (Dense + LayerNorm) to unify it into a common feature space. The Dense layer (fully connected layer) is one of the most basic components in a neural network, typically consisting of a linear transformation (weight matrix multiplication) and a nonlinear activation function. LayerNorm is a normalization technique designed to accelerate training and improve model stability by standardizing the distribution of activation values ​​within each layer of a neural network. This eliminates inter-modal scale differences and enhances feature nonlinearity, resulting in an output feature size of 1×512. Next, a Transformer encoder is used to perform cross-modal feature attention computation, fully exploiting implicit correlations between multivariate data. The output feature size is 1×512. The MLP ReLU activation function performs a nonlinear transformation on the input before passing it to the next layer. Finally, three decoupled MLP regression heads are used for regression prediction, generating predicted values ​​for position, velocity, and attitude, with sizes of 1×3, 1×3, and 1×4, respectively. An MLP (Multilayer Perceptron) regression head is a neural network structure consisting of one or more fully connected layers, used for regression tasks. It is typically located at the end of the model and is responsible for mapping features extracted by previous layers to specific predicted values.

[0147] In some more specific embodiments, in terms of spatial calibration, due to factors such as manufacturing and assembly deviations, it is necessary to perform distortion correction on the surround view camera to eliminate radial distortion and tangential distortion, as shown in the following equation:

[0148] ,

[0149] ,

[0150] Among them, (x, y) is the undistorted normal coordinate, (x0, y0) is the distorted coordinate, k1, k2, k3 are the radial distortion coefficients, p1, p2 are the tangential distortion coefficients, and r is the distance from the pixel to the center of the image.

[0151] In this optional embodiment, spatial calibration of the raw multimodal data eliminates data inconsistencies between sensors due to factors such as position and orientation. This allows data from different sources to accurately represent real-world information within the same coordinate system, thereby improving overall data accuracy. When both space and time are precisely calibrated, fusion analysis is much easier.

[0152] Optionally, the target multimodal data includes target surround camera data, target lidar data, and target hyperspectral data, and the target detection image data is obtained by performing target detection on the target multimodal data, including:

[0153] Performing RGB image enhancement on the target surround view camera data according to the target surround view camera data, the target lidar data, and the target hyperspectral data to obtain enhanced RGB image data;

[0154] The enhanced RGB image data, the target lidar data and the target hyperspectral data are respectively subjected to feature extraction by a feature extraction backbone network to obtain low-level features, mid-level features and high-level features;

[0155] Inputting the low-level features, the mid-level features and the high-level features into a low-level feature fusion module, a mid-level feature fusion module and a high-level feature fusion module respectively to obtain low-level feature fusion data, mid-level feature fusion data and high-level feature fusion data;

[0156] The low-level feature fusion data, the mid-level feature fusion data and the high-level feature fusion data are input into a PAN feature pyramid module to obtain the target detection image data.

[0157] Specifically, the deep learning-based multimodal data fusion target detection model deployed in the onboard computer fuses the RGB image data, LiDAR data, and hyperspectral image data simultaneously acquired by the roof sensor module. Three deep neural network-based feature extraction backbone networks, Backbone_RGB, Backbone_Hyperspectral, and Backbone_LiDAR, are used to extract features from RGB images, hyperspectral images, and LiDAR point cloud mapping images, respectively, obtaining low-, medium-, and high-level data features. Backbone_RGB, the RGB image feature extraction backbone network, is used to extract rich spatial information from RGB images. Backbone_Hyperspectral, the hyperspectral image feature extraction backbone network, is used to extract spectral information and its combination with spatial information from hyperspectral images. Backbone_LiDAR, the point cloud data feature extraction backbone network, is used to extract 3D geometric structure information from LiDAR point cloud data. Then, three multimodal data feature fusion modules, namely the low-level feature fusion module "fusion module_1", the medium-level feature fusion module "fusion module_2" and the high-level feature fusion module "fusion module_3", are used to adaptively fuse low, medium and high data features respectively.

[0158] In this optional embodiment, based on image enhancement, for the detection of ground objects in surface change patterns, full use is made of the multimodal data obtained by the roof sensor, and a deep learning-based multimodal data fusion change pattern ground object detection method is designed to enhance the system's fine-grained detection capability on the ground and output the ground object category and location.

[0159] Optionally, performing RGB image enhancement on the target surround view camera data according to the target surround view camera data, the target lidar data, and the target hyperspectral data to obtain enhanced RGB image data includes:

[0160] Extracting features from the target surround camera data through a convolutional layer to obtain RGB image features;

[0161] Decomposing the RGB image features using a discrete wavelet transform module to obtain target high-frequency features and target low-frequency features;

[0162] Among them, the target high-frequency features are:

[0163] ,

[0164] Among them, F RGB is the RGB image feature, F High ' is the target high frequency feature, F Highis the high-frequency feature, DWT()_High is the high-frequency feature obtained by decomposing the discrete wavelet transform module, and Conv() is the convolution operation;

[0165] Wherein, the target low-frequency feature is:

[0166] ,

[0167] Among them, F Low ' is the target low-frequency feature, F Low It is a low-frequency feature, and DWT()_Low is a low-frequency feature obtained by decomposing it using the discrete wavelet transform module;

[0168] Extracting features of the target lidar data through a convolutional layer to obtain lidar image features;

[0169] Inputting the target high-frequency feature and the lidar image feature into a first adaptive attention mechanism fusion module to obtain a first fusion feature;

[0170] The first fusion feature is:

[0171] ,

[0172] Among them, F Fusion1 is the first fusion feature, fusion module #1 is the first adaptive attention mechanism fusion module, F LiDAR is the laser radar image feature;

[0173] Extracting features of the target hyperspectral data through a convolutional layer to obtain hyperspectral image features;

[0174] Inputting the target low-frequency feature and the hyperspectral image feature into a second adaptive attention mechanism fusion module to obtain a second fused feature;

[0175] The second fusion feature is:

[0176] ,

[0177] Among them, F Fusion2 is the second fusion feature, fusion module #2 is the second adaptive attention mechanism fusion module, F Hyperspectral is the hyperspectral image feature;

[0178] The enhanced RGB image data is obtained by fusing the first fusion feature and the second fusion feature and processing them through a convolution layer.

[0179] Specifically, if Figure 4As shown in FIG, target surround view camera data, target lidar data and target hyperspectral data are used to enhance the RGB image of target surround view camera data. The target surround view camera data is used to represent the RGB image data, and the complementary characteristics of RGB image data, lidar data and hyperspectral data are used. Among them, point cloud data has strong high-frequency features such as shape and distance, while hyperspectral data contains more low-frequency features. This method includes a main branch B for processing RGB image data. RGB and two lidar processing branches B that process lidar images and hyperspectral images respectively LiDAR , Hyperspectral image processing branch B Hyperspectral . Main branch B RGB Extract features F through convolutional layers RGB Then, the discrete wavelet transform module (DWT) is used to divide the features into high-frequency features F High and low-frequency features F Low After processing separately, the two branches simultaneously use superimposed convolutional layers to extract the target high-frequency features and target low-frequency features, where the target high-frequency features are:

[0180] ,

[0181] Among them, F RGB is the RGB image feature, F High ' is the target high frequency feature, F High is the high-frequency feature, DWT()_High is the high-frequency feature obtained by decomposing the discrete wavelet transform module, and Conv() is the convolution operation;

[0182] Wherein, the target low-frequency feature is:

[0183] ,

[0184] Among them, F Low ' is the target low-frequency feature, F Low It is a low-frequency feature, and DWT()_Low is a low-frequency feature obtained by decomposing it using the discrete wavelet transform module.

[0185] Two fusion modules based on channel adaptive attention mechanism are used to extract the high-frequency features F from the RGB image. High ' and the feature F extracted from the lidar image LiDAR Fusion to get F Fusion1 , the low-frequency features F extracted from the RGB image Low ' and the feature F extracted from the hyperspectral image Hyperspectral Fusion to get F Fusion2 ;

[0186] The target high-frequency features and the lidar image features are input into the first adaptive attention mechanism fusion module to obtain the first fused feature, which is:

[0187] ,

[0188] Among them, F Fusion1 is the first fusion feature, fusion module #1 is the first adaptive attention mechanism fusion module, F LiDAR is the laser radar image feature;

[0189] The target low-frequency features and the hyperspectral image features are input into the second adaptive attention mechanism fusion module to obtain the second fused features, which are:

[0190] ,

[0191] Among them, F Fusion2 is the second fusion feature, fusion module #2 is the second adaptive attention mechanism fusion module, F Hyperspectral is the hyperspectral image feature.

[0192] For the first adaptive attention mechanism fusion module, the following calculations are performed:

[0193] ,

[0194] ,

[0195] ,

[0196] ,

[0197] ,

[0198] ,

[0199] in, is the high-frequency feature F High ' and the feature F extracted from the lidar image LiDAR The mixed features obtained after addition, H, W, c represent the mixed features respectively The height, width and band index of For Take the eigenvalue after the average operation in each band, is the intermediate eigenvector after compression calculation, is the high frequency feature F High 'Calculation weight, F LiDAR The calculation weight of is the high frequency feature F High 'With the laser radar feature F LiDAR The weight of the c-th band after fusion;

[0200] For the second adaptive attention mechanism fusion module, the following calculations are performed:

[0201] ,

[0202] ,

[0203] ,

[0204] ,

[0205] ,

[0206] ,

[0207] in, is the low-frequency feature F low 'With features extracted from hyperspectral images The mixed features obtained after addition, H, W, c represent the mixed features respectively The height, width and band index of For Take the eigenvalue after the average operation in each band, is the intermediate eigenvector after compression calculation, is the low-frequency feature F low 'Calculation weight, for The calculation weight of is the low-frequency feature F low 'With hyperspectral image features The weight of the c-th band after fusion.

[0208] The fused high-frequency features are further fused with the low-frequency features to obtain the final features. Finally, two convolutional layers are used to process the restored RGB image data. During model training, the loss function is defined as the mean square error (MSE) between the restored RGB image and the well-illuminated RGB image, and the global SSIM (structural similarity index).

[0209] In some more specific embodiments, the fusion module calculates the RGB image features U obtained by the feature extraction backbone network. RGB , hyperspectral image features U Hyper and lidar feature U LiDAR Add element by element to get the mixed feature U:

[0210] ,

[0211] Then, the mixed feature U is calculated by the following formula to obtain S c 、Z:

[0212] ,

[0213] ,

[0214] Where H, W, and c represent the mixed features U and c The height, width and band index of S c is the average value of all pixel values ​​in the channel, Z is the result after batch normalization, i and j are used to traverse each pixel position in the feature map, δ represents the ReLU function, and S is the final S c , Bn represents batch normalization, W∈R d×C d is calculated by the following formula:

[0215] ,

[0216] Where r represents the reduction ratio, L represents the minimum value of d, and max represents the maximum value.

[0217] Next, the fusion weights of the features of the three modal data are calculated by the following formula:

[0218] ,

[0219] ,

[0220] ,

[0221] Among them, RGB c 、Hyper c 、LiDAR c Represent the calculation weights of RGB image, hyperspectral image and lidar in the Cth band respectively, O, P, Q are the calculation weights obtained through model training; O c 、P c , Q c They are the corresponding calculation weights on the cth band respectively. The three are multiplied element by element through the following formula to finally obtain the fused feature map U Fusion :

[0222] ,

[0223] Where, the RGB image feature U on the cth band is RGB_c , hyperspectral image features U Hyper_cand lidar feature U LiDAR_c , U Fusion_c Represents the feature map obtained by fusing the three modal data on the c-th band.

[0224] After fusing these multimodal data features, we generate three levels of fused data. This data is then fed into the PAN feature pyramid module, where it is further fused using a top-down and then bottom-up approach. Finally, three decoupled and efficient detection heads enable detection of objects of varying sizes.

[0225] In this optional embodiment, a multimodal RGB image enhancement method based on multimodal data is designed to address the problem that surround view cameras are susceptible to lighting conditions, resulting in blurred and unusable images. This method leverages the geometric characteristics of lidar data and the spectral detail of hyperspectral image data to output RGB images with enhanced visual quality.

[0226] Optionally, compressing the target detection image data to obtain bit stream data includes:

[0227] Performing mask processing on the target detection image data to obtain filtered target detection image data;

[0228] The bit stream data is obtained by compressing the filtered target detection image data through an entropy coding module.

[0229] Specifically, the collected target surround camera data, target lidar data, and target hyperspectral data are highly aligned in terms of timestamps and spatial locations. In addition to the target of interest, these data often contain a large amount of redundant information. Therefore, a region-of-interest (ROI)-guided compression operation is required to maximize the preservation of valuable information in the region of interest and reduce useless information in other areas. First, the lidar data is converted from point cloud information such as distance and reflection intensity into raster data for subsequent processing by a deep neural network. Furthermore, the results obtained by the aforementioned target detection model are used to guide the compression framework using the ROI. Specifically, encoding is performed first. The encoder utilizes a convolutional structure based on a wavelet transform, which achieves a large receptive field while maintaining locality. The detection box generated by the target detection model is then masked with a 0-1 mask, with the target box region set to 1 and all other regions set to 0. This 0-1 mask image is then used to process the features generated by the encoder, maximizing the preservation of valuable information and reducing redundancy. The resulting features are then encoded using an entropy coding module, resulting in a bitstream that is then transmitted via a transmission module. Deployed in the cloud, a decoder symmetrical to the encoder processes the input features and recovers restored data that approximates the original data. The error between the restored data and the input data is measured using the mean squared error (MSE) of the region of interest and the global structural similarity index (SSIM).

[0230] In this optional embodiment, a timestamp- and space-stamp-driven multimodal data collaborative compression method is designed to address the issues of insufficient storage space and efficient transmission caused by the large amount of data collected by the master vehicle's roof sensor module. This method first generates a compressed bitstream, then transmits the bitstream to the cloud via a multimodal communication module. The recovered data is then decompressed and backed up in the cloud.

[0231] Optionally, before compressing the filtered target detection image data by the entropy coding module to obtain the bit stream data, the method further includes:

[0232] Obtain meteorological data, cloud images, and sky images;

[0233] Inputting the meteorological data, the cloud image and the sky image into a satellite communication quality prediction model to obtain an effective communication window;

[0234] The data compression rate of the entropy coding module is adjusted according to the effective communication window.

[0235] In some more specific embodiments, adjusting the data compression rate of the entropy coding module according to the effective communication window includes:

[0236] adjusting the data compression rate of the entropy coding module according to the effective communication window;

[0237] The objective function of the entropy coding module is:

[0238] ,

[0239] ,

[0240] Among them, R is the bit rate (the amount of data after image compression), D is the degree of distortion between the decompressed and reconstructed image and the original image, and N is the number of image pixels. are the original image and the reconstructed image respectively, 、 are the pixel values ​​of the original image and the reconstructed image at the i-th position, is the 0-1 mask data of the i-th position, is the structural similarity index (SSIM) between the original image and the reconstructed image, and WindowTime is the data compression rate, which is used to control the coefficient of the trade-off between R and D.

[0241] WindowTime is represented by the effective communication window level obtained by the satellite communication quality prediction model, with values ​​ranging from 8192, 4096, and 2048. These values ​​represent excellent (greater than 90 seconds), good (60-90 seconds), and poor (less than 60 seconds) effective communication windows, respectively. When the predicted effective communication window is excellent, WindowTime is set to 8192, shifting the optimization objective toward reducing distortion, resulting in a larger compressed file size but better image reconstruction quality. When the predicted effective communication window is poor, WindowTime is set to 2048, shifting the optimization objective toward reducing the bitrate, resulting in a smaller compressed file size but lowering reconstruction quality. This dynamic adjustment and selection of the data compression ratio based on the effective communication window allows for data transmission under varying communication conditions.

[0242] In some more specific embodiments, to ensure stable data transmission under varying communication conditions, the natural resources 3D survey and monitoring vehicle system incorporates a multimode communication module that supports both local and wide-area communications (4G, Tiantong high-orbit narrowband satellites, and Xingyun low-orbit narrowband satellites). This communication module consists of two submodules: a local communication module and a wide-area communication module. This module meets the system's communication requirements under varying communication conditions. The local communication module supports protocols such as LoRa and WiFi, providing the vehicle system with robust local communication capabilities. The wide-area communication module supports 4G, Tiantong high-orbit narrowband satellites, and Xingyun low-orbit narrowband satellites. This allows the vehicle system to utilize terrestrial 4G networks for high-speed data transmission in areas with terrestrial network coverage. In areas without terrestrial network coverage, the vehicle system can flexibly select either Tiantong high-orbit narrowband satellites or Xingyun low-orbit narrowband satellites for data transmission.

[0243] Specifically, a satellite communication quality prediction method based on multimodal data is embedded in the system, such as Figure 5 As shown. The input data of this method include meteorological data (temperature, air pressure, humidity, etc.), cloud image, sky image taken by the camera, and the output data is the effective communication window size. The overall structure of this method includes several feature extraction branches, which are used to extract high-level features of different modal data respectively. For image data, a two-dimensional convolution feature extraction backbone network is used; for meteorological data, a one-dimensional convolution feature extraction backbone network is used. For the high-level features obtained using the one-dimensional convolution feature extraction backbone network, a matrix broadcast operation is first performed on it. For example, the feature vector obtained by the temperature branch is processed so that the 1*N high-level feature vector is expanded into an M*N feature matrix. Then, several M*N feature matrices obtained by different branches are spliced ​​on the channel to obtain a C*M*N feature matrix, where C represents the number of categories of meteorological data used. Then, the feature Feature extracted from the meteorological data is Meteorology , Feature extracted from cloud image Cloud , and features extracted from the sky image taken by the camera Image Input into the Transformer fusion module for fusion. The fused feature Fusion The fully connected layer then processes the data to obtain the output data, which is the effective communication window size.

[0244] The model is then trained using a two-stage training strategy. To address the limited number of training samples, a data augmentation method based on a generative adversarial network is used to expand the sample size. The generator generates data on a large scale using random noise, producing a large number of generated samples (fake samples), which are then fed into the discriminator. Simultaneously, real samples from a small amount of real data are fed into the discriminator for discrimination. The generator and the discriminator engage in a game of training. As training progresses, the fake samples generated by the generator become difficult for the discriminator to detect as fake, and the discriminator's discrimination ability gradually improves. Ultimately, the samples generated by the generator are able to achieve an effect that is "indistinguishable from the real." The model is pre-trained using a large number of samples generated by the generator, and then fine-tuned using a small amount of real data.

[0245] For natural resource survey and monitoring work, which often involves harsh environments and lacks ground network coverage, a multimodal communication module was designed that supports both local-area communications (LoRa, WiFi) and wide-area communications (4G, Tiantong high-orbit narrowband satellites, and Xingyun low-orbit narrowband satellites). Where ground networks are available, 4G networks are used for data transmission; in areas without ground network coverage, Tiantong high-orbit narrowband satellites or Xingyun low-orbit narrowband satellites are used for data transmission. Furthermore, a satellite communication quality prediction method based on multimodal data was designed to predict the effective satellite communication window in advance to guide the execution of the multimodal communication module. The inputs include sky images acquired by the master vehicle, meteorological data collected by the master vehicle's built-in meteorological monitoring equipment, and historical remote sensing cloud image data. The output is the effective communication window size. The predicted effective satellite communication window guides the aforementioned data compression model to dynamically adjust the data compression ratio in advance to maximize data transmission within the communication window. If the predicted communication quality is poor and communication is unavailable, the data is temporarily stored in the NAS cache queue, awaiting subsequent transmission. The cache queue is updated dynamically, that is, when storage space is insufficient, the earliest historical data can be replaced by the latest data.

[0246] Optionally, the natural resource survey and monitoring system further includes a drone and a quadruped robot, wherein the drone and the quadruped robot are provided with a sensor module including an RGB camera, a hyperspectral camera, and a laser radar, and the quadruped robot is further provided with a sampling manipulator arm; the natural resource survey and monitoring method further includes:

[0247] Obtaining multimodal data of the drone through the RGB camera, the hyperspectral camera, and the lidar of the drone;

[0248] Obtaining multimodal data of the quadruped robot through the sampling manipulator arm of the quadruped robot, the RGB camera, the hyperspectral camera and the laser radar;

[0249] The natural resources survey and monitoring report is obtained based on the original multimodal data, the drone multimodal data and the quadruped robot multimodal data.

[0250] In some more specific embodiments, the system primarily consists of three subsystems: a master vehicle subsystem, an unmanned aerial vehicle (UAV) subsystem, and a quadruped robot subsystem. The master vehicle serves as the core, collaborating with the UAV and quadruped robot subsystems. Based on the survey and monitoring needs of the area to be surveyed, users can plan and configure the work paths, sensor acquisition frequencies, and other parameters for the master vehicle, UAV, and quadruped robot on the in-vehicle command screen. The UAV can be equipped with sensors such as RGB cameras, hyperspectral cameras, and LiDAR, as needed; the quadruped robot can also be equipped with RGB cameras, LiDAR, and sampling manipulators, as needed. Before data collection officially begins, the system performs spatiotemporal calibration on the master vehicle. This includes distortion correction for the surround-view camera in the master vehicle's roof sensor module, spatial calibration between sensors, and precise spatial positioning and temporal calibration of the master vehicle. The multimodal data from the UAV and quadruped robot are processed identically to the original multimodal data. The UAV and quadruped robot can detect locations inaccessible to the master vehicle. The collected multimodal data is ultimately transmitted to a cloud device, which generates a natural resource survey and monitoring report.

[0251] Due to the influence of factors such as the manufacturing or assembly process of the surround view camera, the natural resource survey monitoring ground objects in the area around the camera's field of view will be severely distorted, affecting the subsequent data processing. Therefore, distortion correction is performed to eliminate radial distortion and tangential distortion. In addition, the RGB camera carried by the unmanned aerial vehicle and the quadruped robot also performs distortion correction operation to eliminate the information deviation caused by distortion. In addition, the positions between the various sensor units equipped in the roof sensor module are determined by matrix calibration method to accurately obtain the translation and rotation matrix between the four surround view cameras, one 32-line laser radar and one light and small hyperspectral camera, which are used for subsequent data alignment operation. In the natural resource survey and monitoring work, the working environment of the host vehicle is complex, including mountainous areas, forest areas, tunnels and so on. Only relying on single GNSS RTK or using GNSS+IMU scheme cannot better meet the spatial positioning demand. Therefore, on the basis of the soft synchronization scheme of PPS signal triggering data synchronization collection, a multi-element data adaptive fusion end-to-end spatial accurate positioning model based on deep neural network is designed to meet the high-precision spatial positioning demand of the host vehicle in the above complex working scene. The input is the multi-element data of GNSS RTK, IMU, surround view camera image and laser radar point cloud; the output is the position, speed and attitude information of the host vehicle. Secondly, the roof sensor module on the host vehicle is configured with a large number of sensor units, including four surround view cameras, one 32-line laser radar and one light and small hyperspectral camera. It is essential to realize the simultaneous collection of data through the cooperation between the sensors on the basis of obtaining high-precision time calibration information. Therefore, in the time calibration of the host vehicle, PPS+GPRMC+master-slave PTP protocol is used to realize the time information synchronization between the vehicle-mounted computer, four surround view cameras, one 32-line laser radar and one light and small hyperspectral camera, which establishes the time basis for the next step of data synchronization collection. After this processing, the multi-modal data with accurate spatial position information and time information are output.

[0252] After the task planning is completed and the time and space calibration related processing is completed, the host vehicle, the unmanned aerial vehicle and the quadruped robot begin to perform the survey and monitoring task. First of all, as mentioned above, the host vehicle mainly performs fine ground data collection in the area with relatively flat ground and paved road. Then, for the areas that the host vehicle cannot reach, the unmanned aerial vehicle takes off from the host vehicle to perform the survey and monitoring task, and the sensor carried by the unmanned aerial vehicle rushes to the upper air of such area to collect the corresponding data. Secondly, the quadruped robot enters the dangerous area with steep terrain, serious shelter and unknown environment (the host vehicle and the unmanned aerial vehicle are difficult to collect data) to work, and uses the sensor carried to obtain the corresponding data, or uses the sampling mechanical arm carried to collect rock, water and soil samples.

[0253] Data collected by the master vehicle, drone, and quadruped robot are used for subsequent natural resource survey and monitoring. Each of the three systems has its own specific division of labor for different tasks. First, the master vehicle and drone collaborate to enable automated, multi-dimensional on-site survey and verification of surface change patterns for natural resource surveys and monitoring. For the master vehicle, a multimodal RGB image enhancement method based on multimodal data was designed to address the issue of surround-view cameras being easily affected by lighting conditions, resulting in blurry and unusable images. This method leverages the geometric characteristics of LiDAR point cloud data and the spectral detail of hyperspectral image data to output RGB images with enhanced visual quality. Secondly, building on this image enhancement, a deep learning-based multimodal data fusion method for surface change pattern object detection was designed, leveraging the multimodal data collected by the vehicle's roof-mounted sensors. This method enhances the system's fine-grained detection capabilities on the ground and outputs object classification and location. The master vehicle only captures detailed information surrounding the objects in the change pattern and is unable to capture the broader aerial perspective. Therefore, a drone-based edge intelligence-based UAV change pattern field survey subsystem was designed. Leveraging the drone's high viewing angle, this system enables aerial data collection and processing, supplementing on-site verification efforts with an aerial dimension. The input is image data captured by the drone, and the output is the category and location of ground objects. To address the needs of natural resource surveys and monitoring for both on-site sampling (point-based) and large-area remote sensing monitoring (area-based), such as water quality sampling and analysis, and regional-level monitoring based on the large-scale data captured by the drone, a drone-quadruped robot coupled sampling and monitoring solution was designed. First, the quadruped robot uses its back-mounted robotic arm to collect water quality samples at designated locations. It then uses a water quality monitoring probe to measure the water quality at each sampling point and simultaneously record the latitude and longitude coordinates of the sampling points. For areas beyond the center of the water body, the drone can be equipped with a water quality monitoring probe for measurement. The drone then uses a lightweight hyperspectral camera to capture large-scale images of the designated water body. Next, the water quality data is spatially aligned with the drone-collected image data based on their latitude and longitude locations. A cubic polynomial is used to fit the pixel values ​​and water quality values ​​at the sampling points to create a water quality inversion model. Finally, this water quality inversion model is applied to the entire image to generate a large-scale, regional water quality monitoring image. Thus, the input is the water quality monitoring values ​​and image data acquired by the quadruped robot and drone, and the output is a large-scale, regional water quality monitoring image.

[0254] To address the challenges of insufficient storage space and efficient transmission caused by the excessive data volume collected by the master vehicle's rooftop sensor module, a timestamp-spacestamp-driven multimodal data collaborative compression method was designed. This method first generates a compressed bitstream, then transmits it to the cloud via a multimodal communication module. Decompression is then performed to recover the data, which is then backed up and stored in the cloud. For natural resource survey and monitoring work, which often involves harsh environments and lacks terrestrial network coverage, a multimodal communication module was designed that supports both local-area communications (LoRa, WiFi) and wide-area communications (4G, Tiantong high-orbit narrowband satellites, and Xingyun low-orbit narrowband satellites). Where terrestrial networks are available, 4G networks are used for data transmission; in areas without terrestrial network coverage, Tiantong high-orbit narrowband satellites or Xingyun low-orbit narrowband satellites are used for data transmission. Furthermore, a satellite communication quality prediction method based on multimodal data was designed to predict the effective window of satellite communication in advance, guiding the execution of the multimodal communication module. The inputs are sky images acquired by the master vehicle, meteorological data collected by its built-in meteorological monitoring equipment, and historical remote sensing cloud image data. The output is the effective communication window size. The satellite communication effective window predicted by the prediction method guides the aforementioned data compression model to dynamically adjust the data compression ratio in advance to maximize data transmission within the communication window. If the predicted communication quality is poor and communication is unavailable, the data is temporarily stored in the NAS cache queue, awaiting subsequent transmission. This cache queue is dynamically updated, meaning that if storage space is insufficient, the oldest historical data can be replaced by the latest data.

[0255] In this optional embodiment, the master vehicle, drone, and quadruped robot each have unique characteristics in terms of mobility, environmental adaptability, and the type of data they acquire. These advantages complement and complement each other in natural resource survey and monitoring operations. The master vehicle has strong ground mobility and a large carrying capacity. Using the multi-source sensors mounted on its roof, it can operate stably in relatively flat, paved areas, thereby acquiring detailed ground data. However, the master vehicle cannot access areas with large terrain undulations or narrow passages. With its high aerial perspective and wide data coverage, the drone can fly over areas difficult for the master vehicle to access, acquiring ground data, effectively filling in the master vehicle's data collection blind spots. However, due to limited battery life and a relatively fixed aerial perspective, the drone's operation is limited in some obstructed areas, such as woodlands. The quadruped robot, on the other hand, has excellent terrain adaptability and the ability to explore unknown and dangerous environments, enabling operations in complex areas such as forests. Therefore, the master vehicle, drone, and quadruped robot work in unison and collaborate, enabling the entire system to achieve all-terrain survey and monitoring capabilities.

[0256] In some more specific embodiments, the master vehicle subsystem consists of an onboard computer, a positioning module, a multimode communication module, a roof-mounted sensor module, a NAS data logger, and solar cells. The positioning module, which includes an IMU and GNSS positioning module, provides high-precision positioning and timing for the master vehicle and is connected to the onboard computer via a serial port. The multimode communication module includes Tiantong satellite communication, 4G / 5G, and image transmission modules to meet the system's data transmission requirements in various operating environments.

[0257] The rooftop sensor module includes three sensors of different modalities: a lidar, four surround-view cameras, and a hyperspectral camera. The lidar connects to the onboard computer via the ETH 1 interface, transmitting large volumes of collected 3D point cloud data at high speed. The four surround-view cameras connect to the onboard computer via a serializer-deserializer (Serializer-Deserializer) using the GMSL protocol, enabling high-bandwidth, low-latency, and robust data transmission. The hyperspectral camera connects to the onboard computer via a USB interface, transmitting stable data and accepting remote sampling from the onboard computer. A NAS data logger, with up to 40TB of storage, is used to store and back up data generated during system operation. It connects to the onboard computer via the ETH 2 interface. The solar cell stores sufficient energy to provide a stable power supply to the main control vehicle and can be recharged autonomously in good lighting conditions. The drone subsystem and quadruped robot subsystem communicate with the onboard computer via local area image transmission signals and Wi-Fi protocols, respectively, enabling information exchange and simultaneous operation.

[0258] In some more specific embodiments, the onboard computer receives a patch file from the backend business system via a cloud platform. Based on the patch's location and area, it automatically generates an initial flight path covering the survey area, including trajectory, altitude, and image capture frequency. Field personnel adjust the initial flight path based on specific site conditions and select the target detection model embedded in the onboard computer as needed. The drone survey and monitoring mission is then dispatched. Upon receiving the mission execution information from the cloud platform, the drone airport activates the drone within the cabin and begins the mission, transmitting the captured video stream in real time. The drone airport incorporates an AI computing module and deploys a trained target detection model trained on a sample library from a cloud server. This model intelligently processes the live video stream, performing frame-by-frame target detection to extract valuable information, fully utilizing edge computing power to reduce data transmission. The sample library includes nine types of features closely related to farmland and forestland protection: existing buildings, buildings under construction, simple sheds, towers, quarries (mines), aquaculture cages, earth piles, greenhouses, and ponds. In addition, the on-board computer deploys a drone image orthophoto stitching module and a 3D reconstruction module to support quick preview for on-site investigation and verification.

[0259] In some more specific embodiments, from a business process perspective, the user first issues instructions, such as "Please help me develop a mission plan based on survey requirements, local topographic maps, and other information," or "Please further organize and analyze the collected data according to the 'Basic Monitoring Field Verification Work Guidelines' to produce a field survey report for the area." The large-scale model system then comprehensively analyzes these instructions, relevant natural resource survey and monitoring information (geological atlases, local chronicles, etc.), and processed data transmitted in real time by the survey and monitoring vehicle system, and outputs a corresponding mission plan and survey report. Mission plans primarily fall into two categories. The first, for automated, multi-dimensional on-site verification of surface change patterns in natural resource surveys and monitoring, primarily includes the survey sequence and route for the master vehicle, the location and angle of data collection for the master vehicle, the flight path and frequency of image acquisition for the drone, and the type of ground object detection model. The second, for the natural resource survey and monitoring requirements of both on-site sampling (point-based) and large-scale remote sensing monitoring (area-based), primarily includes the types of sensors to be equipped on the drone and quadruped robot, as well as the locations of sampling points. Accordingly, two types of survey reports are also produced. One includes the types, locations and areas of surface changes monitored by natural resource surveys; the other includes the results of large-area remote sensing monitoring, such as the distribution of water quality in water bodies.

[0260] like Figure 6 As shown, a natural resource survey and monitoring device is based on a natural resource survey and monitoring system, wherein the natural resource survey and monitoring system includes a master control vehicle and a cloud device; the natural resource survey and monitoring device includes:

[0261] A mission planning scheme acquisition module 10 is used to acquire a mission planning scheme through a decision support system using the cloud device;

[0262] The original multimodal data acquisition module 20 is used to control the master vehicle to collect original multimodal data through the mission planning scheme;

[0263] a spatiotemporal calibration processing module 30 for obtaining raw multimodal data from the master vehicle and performing spatiotemporal calibration processing on the raw multimodal data to obtain target multimodal data, wherein the target multimodal data is used to represent multimodal data with spatial position information and time information;

[0264] The target detection module 40 is configured to obtain target detection image data by performing target detection on the target multimodal data;

[0265] A data compression module 50 is configured to compress the target detection image data to obtain bit stream data, and send the bit stream data to the cloud device;

[0266] The natural resource survey and monitoring report acquisition module 60 is used to use the cloud device to decode the bit stream data to obtain a natural resource survey and monitoring report.

[0267] The natural resource survey and monitoring device of this embodiment is used to implement the natural resource survey and monitoring method described above. Its advantages over the existing technology are the same as the advantages of the above-mentioned natural resource survey and monitoring method over the existing technology, and will not be repeated here.

[0268] Optionally, the spatiotemporal calibration processing module 30 is specifically configured to: obtain spatially calibrated multimodal data by performing spatial calibration processing on the original multimodal data;

[0269] Based on a time synchronization protocol, time calibration processing is performed on the original multimodal data to obtain time-calibrated multimodal data, wherein the time synchronization protocol includes a PPS pulse per second signal, a GPRMC data transmission format, and a master-slave PTP protocol;

[0270] The target multimodal data is obtained according to the spatially calibrated multimodal data and the temporally calibrated multimodal data.

[0271] Optionally, the spatiotemporal calibration processing module 30 is specifically configured to: perform unified coordinate system processing on the original surround-view camera data, the original lidar data, and the original hyperspectral data to obtain surround-view camera data in a unified coordinate system, lidar data in a unified coordinate system, and hyperspectral data in a unified coordinate system, wherein the original surround-view camera data includes first surround-view camera data, second surround-view camera data, third surround-view camera data, and fourth surround-view camera data;

[0272] The process of processing the original laser radar data into a unified coordinate system is as follows:

[0273] ,

[0274] Among them, P camera#1 is the coordinate of point P in the coordinate system of the first surround-view camera data, P LiDAR is the coordinate of point P in the coordinate system of the original lidar data, R is the lidar data rotation matrix, and T is the lidar data translation vector;

[0275] The process of processing the original hyperspectral data into a unified coordinate system is as follows:

[0276] ,

[0277] Among them, P H is the coordinate of point P in the coordinate system of the original hyperspectral data, R H is the hyperspectral data rotation matrix, TH is the hyperspectral data translation vector;

[0278] The process of processing the original surround view camera data into a unified coordinate system is as follows:

[0279] ,

[0280] ,

[0281] ,

[0282] Among them, P camera#2 is the coordinate of point P in the coordinate system of the second surround-view camera data, P camera#3 is the coordinate of point P in the coordinate system of the third surround-view camera data, P camera#4 is the coordinate of point P in the coordinate system of the fourth surround-view camera data, R camera#2 is the second surround camera rotation matrix, R camera#3 is the third surround camera rotation matrix, R camera#4 is the fourth surround camera rotation matrix, T camera#2 is the translation vector of the second surround camera, T camera#3 is the translation vector of the third surround camera, T camera#4 is the translation vector of the fourth surround camera;

[0283] Inputting GNSS positioning data, surround view camera data in the unified coordinate system, and lidar data in the unified coordinate system into a spatial precision positioning model for feature extraction and fusion to obtain predicted values ​​of the position, speed, and attitude of the master vehicle;

[0284] Wherein, the spatial precise positioning model is:

[0285] ,

[0286] ,

[0287] ,

[0288] ,

[0289] Among them, B 时空感知 B is the spatiotemporal perception feature processing branch of the spatial precise positioning model, 视觉感知 B is the visual perception feature processing branch of the spatial precise positioning model, 几何感知 is the geometric perception feature processing branch of the spatial precise positioning model, F 时空感知 is the feature vector obtained by the spatiotemporal perception feature processing branch, F 视觉感知is the feature vector obtained by the visual perception feature processing branch, F 几何感知 is the feature vector obtained by the geometric perception feature processing branch, GNSSRTK+IMU time series data is the GNSS positioning data, F 联合 is to fuse feature vectors, and Concat is the feature fusion module;

[0290] The spatial calibration multimodal data is obtained according to the surround view camera data in the unified coordinate system, the lidar data in the unified coordinate system, the hyperspectral data in the unified coordinate system, and the predicted values ​​of the position, speed, and posture of the master vehicle.

[0291] Optionally, the target detection module 40 is specifically configured to: perform RGB image enhancement on the target surround view camera data according to the target surround view camera data, the target lidar data, and the target hyperspectral data to obtain enhanced RGB image data;

[0292] The enhanced RGB image data, the target lidar data and the target hyperspectral data are respectively subjected to feature extraction by a feature extraction backbone network to obtain low-level features, mid-level features and high-level features;

[0293] Inputting the low-level features, the mid-level features and the high-level features into a low-level feature fusion module, a mid-level feature fusion module and a high-level feature fusion module respectively to obtain low-level feature fusion data, mid-level feature fusion data and high-level feature fusion data;

[0294] The low-level feature fusion data, the mid-level feature fusion data and the high-level feature fusion data are input into a PAN feature pyramid module to obtain the target detection image data.

[0295] Optionally, the target detection module 40 is specifically configured to: perform feature extraction on the target surround camera data through a convolution layer to obtain RGB image features;

[0296] Decomposing the RGB image features using a discrete wavelet transform module to obtain target high-frequency features and target low-frequency features;

[0297] Among them, the target high-frequency features are:

[0298] ,

[0299] Among them, F RGB is the RGB image feature, F High ' is the target high frequency feature, F Highis the high-frequency feature, DWT()_High is the high-frequency feature obtained by decomposing the discrete wavelet transform module, and Conv() is the convolution operation;

[0300] Wherein, the target low-frequency feature is:

[0301] ,

[0302] Among them, F Low ' is the target low-frequency feature, F Low It is a low-frequency feature, and DWT()_Low is a low-frequency feature obtained by decomposing it using the discrete wavelet transform module;

[0303] Extracting features of the target lidar data through a convolutional layer to obtain lidar image features;

[0304] Inputting the target high-frequency feature and the lidar image feature into a first adaptive attention mechanism fusion module to obtain a first fusion feature;

[0305] The first fusion feature is:

[0306] ,

[0307] Among them, F Fusion1 is the first fusion feature, fusion module #1 is the first adaptive attention mechanism fusion module, F LiDAR is the laser radar image feature;

[0308] Extracting features of the target hyperspectral data through a convolutional layer to obtain hyperspectral image features;

[0309] Inputting the target low-frequency feature and the hyperspectral image feature into a second adaptive attention mechanism fusion module to obtain a second fused feature;

[0310] The second fusion feature is:

[0311] ,

[0312] Among them, F Fusion2 is the second fusion feature, fusion module #2 is the second adaptive attention mechanism fusion module, F Hyperspectral is the hyperspectral image feature;

[0313] The enhanced RGB image data is obtained by fusing the first fusion feature and the second fusion feature and processing them through a convolution layer.

[0314] Optionally, the data compression module 50 is specifically configured to: perform mask processing on the target detection image data to obtain filtered target detection image data;

[0315] The target detection image data after screening is compressed by an entropy coding module to obtain the bit stream data.

[0316] Optionally, the natural resource investigation and monitoring device further comprises a satellite communication quality prediction module, which is configured to: acquire meteorological data, a cloud image and a sky image;

[0317] The meteorological data, the cloud image and the sky image are input into a satellite communication quality prediction model to obtain an effective communication window.

[0318] The data compression rate of the entropy coding module is adjusted according to the effective communication window.

[0319] Optionally, the natural resource investigation and monitoring device further comprises a UAV and quadruped robot control module, which is configured to: acquire UAV multi-modal data by the RGB camera, the hyperspectral camera and the laser radar of the UAV;

[0320] Acquire quadruped robot multi-modal data by the sampling mechanical arm, the RGB camera, the hyperspectral camera and the laser radar of the quadruped robot;

[0321] Acquire the natural resource investigation and monitoring report according to the original multi-modal data, the UAV multi-modal data and the quadruped robot multi-modal data.

[0322] As shown in Figure 7 The embodiment of the present application provides an electronic device 700, which comprises a memory 710 and a processor 720; the memory 710 is used for storing a computer program; the processor 720 is used for realizing the natural resource investigation and monitoring method as described above when the computer program is executed.

[0323] Alternatively, an electronic device 700 comprises a memory 710 and a processor 720 coupled to the memory 710; the memory 710 is configured to store a computer program; the processor 720 is configured to execute the following operations when the computer program is executed:

[0324] Acquire original multi-modal data by the master vehicle, and perform space-time calibration processing on the original multi-modal data to obtain target multi-modal data, wherein the target multi-modal data is used for representing multi-modal data with spatial position information and time information;

[0325] Acquire target detection image data by performing target detection on the target multi-modal data;

[0326] Compressing the target detection image data to obtain bit stream data, and sending the bit stream data to the cloud device;

[0327] The cloud device is used to decode the bit stream data to obtain a natural resources survey and monitoring report.

[0328] An electronic device 700 that can serve as a server or client of the present invention will now be described, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device 700 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 700 can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0329] Electronic device 700 includes a computing unit that can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) or loaded from a storage unit into a random access memory (RAM). The RAM can also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. An input / output (I / O) interface is also connected to the bus.

[0330] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM). In this application, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network elements. Some or all of these units can be selected based on actual needs to achieve the objectives of the embodiments of the present invention. Furthermore, the functional units in the various embodiments of the present invention can be integrated into a single processing unit, each unit can exist physically separately, or two or more units can be integrated into a single unit. These integrated units can be implemented in either hardware or software functional units.

[0331] Although the present application has been disclosed with reference to the above embodiments, the scope of the present application is not limited to the above. Various changes and modifications can be made thereto without departing from the spirit and scope of the present application, and such changes and modifications are intended to fall within the scope of the present application.

Claims

1. A natural resource survey and monitoring method, characterized in that: Based on the natural resources survey and monitoring system, the natural resources survey and monitoring system includes a master control vehicle and cloud equipment; The natural resources survey and monitoring method includes: Obtaining a mission planning solution through a decision support system using the cloud device; Controlling the master vehicle to collect original multimodal data through the mission planning scheme; Performing spatiotemporal calibration processing on the original multimodal data to obtain target multimodal data, wherein the target multimodal data is used to represent multimodal data with spatial position information and time information; Target detection image data is obtained by performing target detection on the target multimodal data, wherein the target multimodal data includes target surround camera data, target lidar data, and target hyperspectral data, including: Performing RGB image enhancement on the target surround view camera data according to the target surround view camera data, the target lidar data, and the target hyperspectral data to obtain enhanced RGB image data; The enhanced RGB image data, the target lidar data and the target hyperspectral data are respectively subjected to feature extraction by a feature extraction backbone network to obtain low-level features, mid-level features and high-level features; Inputting the low-level features, the mid-level features and the high-level features into a low-level feature fusion module, a mid-level feature fusion module and a high-level feature fusion module respectively to obtain low-level feature fusion data, mid-level feature fusion data and high-level feature fusion data; Inputting the low-level feature fusion data, the mid-level feature fusion data and the high-level feature fusion data into a PAN feature pyramid module to obtain the target detection image data; Compressing the target detection image data to obtain bit stream data, and sending the bit stream data to the cloud device; The cloud device is used to decode the bit stream data to obtain a natural resources survey and monitoring report.

2. The natural resource survey and monitoring method according to claim 1, characterized in that: The performing spatiotemporal calibration processing on the original multimodal data to obtain target multimodal data includes: Obtaining spatially calibrated multimodal data by performing spatial calibration processing on the original multimodal data; Based on a time synchronization protocol, time calibration processing is performed on the original multimodal data to obtain time-calibrated multimodal data, wherein the time synchronization protocol includes a PPS pulse per second signal, a GPRMC data transmission format, and a master-slave PTP protocol; The target multimodal data is obtained according to the spatially calibrated multimodal data and the temporally calibrated multimodal data.

3. The natural resource survey and monitoring method according to claim 2, characterized in that: The original multimodal data includes original surround camera data, original lidar data, and original hyperspectral data. The spatially calibrated multimodal data is obtained by performing spatial calibration on the original multimodal data, including: Performing unified coordinate system processing on the original surround-view camera data, the original lidar data, and the original hyperspectral data to obtain surround-view camera data in a unified coordinate system, lidar data in a unified coordinate system, and hyperspectral data in a unified coordinate system, wherein the original surround-view camera data includes first surround-view camera data, second surround-view camera data, third surround-view camera data, and fourth surround-view camera data; The process of processing the original laser radar data into a unified coordinate system is as follows: , Among them, P camera#1 is the coordinate of point P in the coordinate system of the first surround-view camera data, P LiDAR is the coordinate of point P in the coordinate system of the original lidar data, R is the lidar data rotation matrix, and T is the lidar data translation vector; The process of processing the original hyperspectral data into a unified coordinate system is as follows: , Among them, P H is the coordinate of point P in the coordinate system of the original hyperspectral data, R H is the hyperspectral data rotation matrix, T H is the hyperspectral data translation vector; The process of processing the original surround view camera data into a unified coordinate system is as follows: , , , Among them, P camera#2 is the coordinate of point P in the coordinate system of the second surround-view camera data, P camera#3 is the coordinate of point P in the coordinate system of the third surround-view camera data, P camera#4 is the coordinate of point P in the coordinate system of the fourth surround-view camera data, R camera#2 is the second surround camera rotation matrix, R camera#3 is the third surround camera rotation matrix, R camera#4 is the fourth surround camera rotation matrix, T camera#2 is the translation vector of the second surround camera, T camera#3 is the translation vector of the third surround camera, T camera#4 is the translation vector of the fourth surround camera; Inputting GNSS positioning data, surround view camera data in the unified coordinate system, and lidar data in the unified coordinate system into a spatial precision positioning model for feature extraction and fusion to obtain predicted values ​​of the position, speed, and attitude of the master vehicle; Wherein, the spatial precise positioning model is: , , , , Among them, B 时空感知 B is the spatiotemporal perception feature processing branch of the spatial precise positioning model, 视觉感知 B is the visual perception feature processing branch of the spatial precise positioning model, 几何感知 is the geometric perception feature processing branch of the spatial precise positioning model, F 时空感知 is the feature vector obtained by the spatiotemporal perception feature processing branch, F 视觉感知 is the feature vector obtained by the visual perception feature processing branch, F 几何感知 is the feature vector obtained by the geometric perception feature processing branch, the GNSS RTK+IMU time series data is the GNSS positioning data, and F 联合 is to fuse feature vectors, and Concat is the feature fusion module; The spatial calibration multimodal data is obtained according to the surround view camera data in the unified coordinate system, the lidar data in the unified coordinate system, the hyperspectral data in the unified coordinate system, and the predicted values ​​of the position, speed, and posture of the master vehicle.

4. The natural resource survey and monitoring method according to claim 1, characterized in that: The performing RGB image enhancement on the target surround view camera data according to the target surround view camera data, the target lidar data, and the target hyperspectral data to obtain enhanced RGB image data includes: Extracting features from the target surround camera data through a convolutional layer to obtain RGB image features; Decomposing the RGB image features using a discrete wavelet transform module to obtain target high-frequency features and target low-frequency features; Among them, the target high-frequency features are: , Among them, F RGB is the RGB image feature, F High ' is the target high frequency feature, F High is the high-frequency feature, DWT()_High is the high-frequency feature obtained by decomposing the discrete wavelet transform module, and Conv() is the convolution operation; Wherein, the target low-frequency feature is: , Among them, F Low ' is the target low-frequency feature, F Low It is a low-frequency feature, and DWT()_Low is a low-frequency feature obtained by decomposing it using the discrete wavelet transform module; Extracting features of the target lidar data through a convolutional layer to obtain lidar image features; Inputting the target high-frequency feature and the lidar image feature into a first adaptive attention mechanism fusion module to obtain a first fusion feature; The first fusion feature is: , Among them, F Fusion1 is the first fusion feature, fusion module #1 is the first adaptive attention mechanism fusion module, F LiDAR is the laser radar image feature; Extracting features of the target hyperspectral data through a convolutional layer to obtain hyperspectral image features; Inputting the target low-frequency feature and the hyperspectral image feature into a second adaptive attention mechanism fusion module to obtain a second fused feature; The second fusion feature is: , Among them, F Fusion2 is the second fusion feature, fusion module #2 is the second adaptive attention mechanism fusion module, F Hyperspectral is the hyperspectral image feature; The enhanced RGB image data is obtained by fusing the first fusion feature and the second fusion feature and processing them through a convolution layer.

5. The natural resource survey and monitoring method according to claim 1, characterized in that: The step of compressing the target detection image data to obtain bit stream data includes: Performing mask processing on the target detection image data to obtain filtered target detection image data; The bit stream data is obtained by compressing the filtered target detection image data through an entropy coding module.

6. The natural resource survey and monitoring method according to claim 5, characterized in that: Before compressing the filtered target detection image data by the entropy coding module to obtain the bit stream data, the method further includes: Obtain meteorological data, cloud images, and sky images; Inputting the meteorological data, the cloud image and the sky image into a satellite communication quality prediction model to obtain an effective communication window; The data compression rate of the entropy coding module is adjusted according to the effective communication window.

7. The natural resource survey and monitoring method according to any one of claims 1 to 6, characterized in that: The natural resource survey and monitoring system further includes a drone and a quadruped robot, wherein the drone and the quadruped robot are provided with a sensor module, wherein the sensor module includes an RGB camera, a hyperspectral camera and a laser radar, and the quadruped robot is further provided with a sampling mechanical arm; The natural resources survey and monitoring method further includes: Obtaining multimodal data of the drone through the RGB camera, the hyperspectral camera, and the lidar of the drone; Obtaining multimodal data of the quadruped robot through the sampling manipulator arm of the quadruped robot, the RGB camera, the hyperspectral camera and the laser radar; The natural resources survey and monitoring report is obtained based on the original multimodal data, the drone multimodal data and the quadruped robot multimodal data.

8. A natural resource survey and monitoring device, characterized in that: Based on the natural resources survey and monitoring system, the natural resources survey and monitoring system includes a master control vehicle and cloud equipment; The natural resources survey and monitoring device comprises: A mission planning scheme acquisition module is used to obtain a mission planning scheme through a decision support system using the cloud device; An original multimodal data acquisition module, configured to control the master vehicle to collect original multimodal data through the mission planning scheme; a spatiotemporal calibration processing module, configured to perform spatiotemporal calibration processing on the original multimodal data to obtain target multimodal data, wherein the target multimodal data is used to represent multimodal data with spatial position information and time information; A target detection module is configured to obtain target detection image data by performing target detection on the target multimodal data, wherein the target multimodal data includes target surround camera data, target lidar data, and target hyperspectral data, including: Performing RGB image enhancement on the target surround view camera data according to the target surround view camera data, the target lidar data, and the target hyperspectral data to obtain enhanced RGB image data; The enhanced RGB image data, the target lidar data and the target hyperspectral data are respectively subjected to feature extraction by a feature extraction backbone network to obtain low-level features, mid-level features and high-level features; Inputting the low-level features, the mid-level features and the high-level features into a low-level feature fusion module, a mid-level feature fusion module and a high-level feature fusion module respectively to obtain low-level feature fusion data, mid-level feature fusion data and high-level feature fusion data; Inputting the low-level feature fusion data, the mid-level feature fusion data and the high-level feature fusion data into a PAN feature pyramid module to obtain the target detection image data; a data compression module, configured to compress the target detection image data to obtain bit stream data, and send the bit stream data to the cloud device; The natural resource survey and monitoring report acquisition module is used to use the cloud device to decode the bit stream data to obtain a natural resource survey and monitoring report.

9. An electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the natural resource survey and monitoring method according to any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • Environmental protection monitoring equipment

    CN105300457A

  • GIS (Geographic Information System) monitoring device for investigating and monitoring natural resources

    CN118410287A