A method for verifying the accuracy of simulated point clouds based on offline querying of ICESat-2 ATL data

By establishing an offline database and using key-value pairs, red-black trees, CSV format, and the DBSCAN clustering algorithm, the network and hard disk management problems in ICESat-2 ATL data processing were solved, achieving efficient and secure data processing and accurate simulation results.

CN119988354BActive Publication Date: 2025-11-18XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510204996.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-11-18
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

Existing technologies suffer from slow network speeds, unstable connections, high risk of data loss, complex hard drive management, and low storage efficiency when downloading and processing ICESat-2 ATL data, making it difficult to achieve efficient and secure data processing and analysis.

Method used

By pre-downloading ICESat-2 ATL data and establishing an offline database, time and latitude/longitude information are stored using key-value pairs. Combined with red-black trees and CSV format, the data is quickly searched and parsed. The DBSCAN clustering algorithm is used to evaluate the simulation results.

Benefits of technology

It ensures data security and integrity, improves data processing efficiency and reliability, features fast searching, low memory usage and high portability, and ensures accurate data parsing and precision of simulation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988354B_ABST
    Figure CN119988354B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on offline query ICESat-2ATL data simulation point cloud precision checking method, belong to laser simulation technical field, its method includes: first, ICESat-2ATL is constructed database, and data name is stored to txt text, then data name is placed in red-black tree and constructs key-value pair;Again, standard ICESat-2 orbit data is obtained, and is stored as csv format with orbit start time naming, then orbit time and latitude and longitude are combined, and are driven in orbit time, search in the key-value pair constructed, then compile HDF5 library to parse the photon event in search result, and compile visual interface, add orbit in visual interface, query display function, obtain point cloud image;Finally, add laser radar point cloud map display interface, put the parsed point cloud image into chart, and use confidence algorithm based on DBSCAN clustering to evaluate and analyze simulation result;With data security, simple, search speed is fast, portability is strong, expandability is strong, and the effect of small space occupation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of laser simulation technology, specifically to a method for verifying the accuracy of simulated point clouds based on offline querying of ICESat-2 ATL data. Background Technology

[0002] ATL-level data plays a crucial role in lidar simulations, providing precise photon-level surface measurement information, supporting the verification and optimization of simulation models, and applicable to various data analyses such as ice thickness monitoring and forest height analysis. However, downloading massive amounts of ATL-level data in real-time from NASA's website suffers from slow network speeds and unstable connections, impacting data acquisition efficiency and increasing the risk of data loss. Furthermore, laboratories and similar facilities have very high requirements for network configurations; to ensure data security, experimental environment stability, and prevent external interference, physical network isolation is typically necessary. In contrast, building an offline database effectively avoids these problems, ensuring data integrity and security while improving data processing efficiency and reliability, and enabling accurate analysis of data from typical terrain areas.

[0003] Ji Jun et al. published the establishment of an alertness database based on HDF5 format (Ji Jun, Wang Jinghua, Zeng Yidong, et al. Establishment of an alertness database based on HDF5 format [J]. China Medical Equipment, 2021, 18(09):6-9.). This study uses HDF5 format to store alertness data in a hierarchical structure, combining header files, data files, and annotation files to facilitate efficient storage, management, and analysis of multi-source physiological data. However, due to the large amount of data to be stored, a large amount of hard disk storage is required. The root directory will change when reading data from different hard disks, and it will consume a lot of storage space when storing a large amount of hierarchical data. It has the disadvantages of data redundancy, reduced reading efficiency, complex hard disk management, and large memory consumption. Summary of the Invention

[0004] To overcome the shortcomings of the prior art, the present invention aims to provide a method for verifying the accuracy of simulated point clouds based on offline querying of ICESat-2 ATL data. This method involves pre-downloading and establishing an ATL03-level offline database, using key-value pairs to store time and latitude / longitude information, and then searching for time and latitude / longitude to quickly find the corresponding ATL03 data. This method features that the data does not occupy additional memory, is reliable and secure, has fast search speed, and is highly portable.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A method for verifying the accuracy of simulated point clouds based on offline querying of ICESat-2 ATL data includes the following steps:

[0007] Step 1: Acquire and store ICESat-2 ATL data to build the base database;

[0008] Step 2: Use a batch script to store the ICESat-2 ATL data name obtained in Step 1 into a text file;

[0009] Step 3: Put the ICESat-2 ATL data names stored in the txt text file in Step 2 into a red-black tree to construct key-value pairs;

[0010] Step 4: Obtain standard ICESat-2 orbital data and store it in a CSV file named after the orbital start time;

[0011] Step 5: Combine the orbit time and latitude / longitude of the ICESat-2 orbit data obtained in Step 4, and search the key-value pairs constructed in Step 3, using orbit time as the driving force.

[0012] Step 6: Compile the HDF5 library to parse the photon events in the search results of Step 5, compile the visualization interface, and add track and query display functions to the visualization interface to obtain the point cloud image;

[0013] Step 7: Add a lidar point cloud map display interface, display the point cloud image parsed in Step 6 in the chart, and use the DBSCAN clustering-based confidence algorithm to evaluate and analyze the simulation results.

[0014] Step 2 specifically includes:

[0015] Name the HDF5 format ICESat-2 ATL data file in the format of "Data Product Type_Acquisition Time_Track Number_Data Version Number_Number of Reprocessing", then store the corresponding ATL data file name in txt format, and use the acquisition time as the primary key to query the specific file path where the ICESat-2 ATL data is located, and parse the ICESat-2 ATL data.

[0016] The data product types include ATL01, ATL02, ATL03, ATL04, ATL05, ATL06, ATL07, ATL08, ATL09, ATL10, ATL11, ATL12, and ATL13 formats, depending on the content and processing level of the data.

[0017] The data collection time includes year, month, day, hour, minute, and second;

[0018] The track number serves as the identification label for the track;

[0019] Data version numbers 001, 002, 003, ..., 00n indicate the revision number;

[0020] Reprocessing is indicated by 00, 01, 02, ..., 0n, signifying whether the orbital data has been processed.

[0021] Where n is an integer.

[0022] Step 3 specifically includes:

[0023] Read the absolute paths of all ICESat-2 ATL data stored in the plaintext txt file generated in step 2, and store them in a red-black tree for fast access.

[0024] Step 4 specifically includes:

[0025] Download all ICESat-2 track files, which are in kml format;

[0026] The xml library in Python code is used to parse the ICESat-2 orbital files into CSV format. The CSV files store the data for each segment of the orbit, named in the format "day-month-year_hour-minute-second.csv" for easy retrieval later. The orbital data is stored in a folder according to the ICESat-2 Earth scanning cycle, and will be used as the orbital time parsing part later.

[0027] Step 5 specifically includes:

[0028] The folder for storing track data is determined by year and month. Then, the track data in the folder is sorted by time. The start time t1 of the track data is located by searching by date, hour, minute and second. The end time t2 of the track data is located by searching in the same way. After finding the track positions of t1 and t2, the track data in the intermediate time period between t1 and t2 is merged to form the track search result, which is used for subsequent track queries.

[0029] The search results for the orbit can be filtered by setting upper and lower limits for latitude and longitude:

[0030] First, set the latitude and longitude range, then iterate through each orbital point and check whether its latitude and longitude are within the given upper and lower limits. At the same time, add a function to search for nearby data for points on the edge of the latitude and longitude area. If the nearest data is found and is within the set range, the orbital point will be included in the final query results. The final generated orbit is also within the latitude and longitude range and is stored.

[0031] Step 6 specifically includes:

[0032] By using C++ and the HDF5 library, the photon position, timestamp, and distance along the track information in ATL03 level data are directly read and parsed. Point cloud images are obtained by parsing the distance along the track and elevation corresponding to photon events in ATL03 level data.

[0033] Step 7 specifically includes:

[0034] Step 7.1: Construct a kd-tree based on the DBSCAN clustering algorithm to accelerate neighborhood query and improve efficiency. The algorithm first initializes the core parameters required for density clustering through input parameters, and converts the original data into a data format that adapts to the algorithm to generate a point set. The core parameters include: neighborhood radius and core point threshold.

[0035] Step 7.2: Construct a kd-tree to find the neighbors of each point within the neighborhood radius, and record the neighbor information of the point set in Step 7.1. In the main loop, initialize the cluster label, noise label and visit label, and check whether each point has been visited.

[0036] Step 7.3: For the points not visited in Step 7.2, obtain their neighbor set. If the number of neighbors is less than the core point threshold, mark the point as noise.

[0037] Step 7.4: For the points visited in Step 7.2, take them as core points and start a new cluster expansion. The expanded cluster recursively processes the neighbors of the core points, assigns unclassified points to the current cluster, and merges neighbors that meet the conditions until all points are classified or are marked as noise.

[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0039] 1. This invention processes typical key information in the database, such as using red-black trees to construct key-value pair searches in steps 3 to 5. It does not process the database itself and does not occupy database storage space, thus achieving the effects of data simplicity and fast search.

[0040] 2. This invention uses absolute paths to store ICESat-2 ATL data. The absolute path serves as an index, enabling the processing of data in different directories and providing high portability.

[0041] 3. This invention uses CSV format to store photon events and ICESat-2 orbital files. Since CSV format files use commas as delimiters, they are easy to read, highly scalable, and occupy little space.

[0042] In summary, this invention utilizes an offline database to ensure a stable scientific research data processing environment. It can complete and better handle some functions of the offline database, and features data security, simplicity, fast search speed, strong portability, strong scalability, and small space footprint. Attached Figure Description

[0043] Figure 1 This is a flowchart of the simulation point cloud accuracy verification method according to an embodiment of the present invention.

[0044] Figure 2 The DBSCAN clustering algorithm of this invention.

[0045] Figure 3 This refers to some information contained in the H5 format file of this invention embodiment.

[0046] Figure 4 This is a ground scanning foot point trajectory diagram according to an embodiment of the present invention.

[0047] Figure 5 This is a standard photon point cloud image from an embodiment of the present invention. Detailed Implementation

[0048] The present invention will now be described in detail with reference to the accompanying drawings.

[0049] This embodiment uses ICESat-2 ATL level 03 data as an example to further illustrate the present invention.

[0050] like Figure 1 As shown, a method for verifying the accuracy of simulated point clouds based on offline querying of ICESat-2 ATL data includes the following steps:

[0051] Step 1: Acquire and store ICESat-2 ATL data to build the base database;

[0052] Step 2: Use a batch script to store the ICESat-2 ATL data name obtained in Step 1 into a text file;

[0053] Step 3: Put the ICESat-2 ATL data names stored in the txt text file in Step 2 into a red-black tree to construct key-value pairs;

[0054] Step 4: Obtain standard ICESat-2 orbital data and store it in a CSV file named after the orbital start time;

[0055] Step 5: Combine the orbit time and latitude / longitude of the ICESat-2 orbit data obtained in Step 4, and search the key-value pairs constructed in Step 3, using orbit time as the driving force.

[0056] Step 6: Compile the HDF5 library to parse the photon events in the search results of Step 5, compile the visualization interface, and add track and query display functions to the visualization interface to obtain the point cloud image;

[0057] Step 7: Add a lidar point cloud map display interface, display the point cloud image parsed in Step 6 in the chart, and use the DBSCAN clustering-based confidence algorithm to evaluate and analyze the simulation results.

[0058] The method for obtaining ATL03 level data in step 1 is as follows:

[0059] The ATL03 level data from 2019 to 2021 is of high quality and has a wide coverage, reflecting the dynamic changes in the Earth's surface environment in recent years. It provides reliable benchmark data for analyzing ice thickness, forest cover, and sea level height, helping researchers to conduct more accurate research and trend prediction on environmental changes. The open-source data information for ICESat-2 was found on website A, and the ATL03 level data was downloaded.

[0060] Step 2 involves processing the data stored on the hard disk array using a batch script, as follows:

[0061] Because the ATL03 class data volume is very large, it requires dozens of 60TB high-capacity hard drives for storage, and the corresponding hard drive enclosures need to be connected to the computer when reading data. Different connection orders and drive letters mean that the corresponding hard drive enclosures need to be re-stored after each connection. Therefore, databases such as MySQL will lose their data storage locations, and the hard drives must be rescanned and their text information re-stored every time a hard drive is inserted.

[0062] ATL03 level data files are named in the format "Data Product Type_Acquisition Time_Track Number_Data Version Number_Reprocessing Count" and stored in HDF5 data format. For example, consider an H5 file named ATL03_20191130141011_09880502_006_01.h5. Here, ATL03 represents the data product type, 20191130141011 represents the acquisition time (November 30, 2019, 14:10:11), 09880502 represents the track number, 006 represents the data version number, 01 represents the reprocessing count, and .h5 indicates the data storage format (HDF5). Since the H5 file name contains time information, the acquisition time is used as the primary key to query the specific file path of the ATL03 level data for parsing. Some information contained in the H5 format file includes... Figure 3 As shown.

[0063] The process of constructing key-value pairs in step 3 is as follows:

[0064] A red-black tree is a self-balancing binary search tree that maintains its balance through color and rotation rules, ensuring that the time complexity of search, insertion, and deletion operations is O(log n). In C++, red-black trees use std::map and std::set to provide fast ordered data storage and access.

[0065] The txt text generated in step 2 is in plaintext and stores the absolute paths of all data in the ATL03 level database. The ATL03 level data paths are read in and stored in a red-black tree for fast access.

[0066] Step 4 involves obtaining data from the standard ICESat-2 orbit as follows:

[0067] You can download all the ICESat-2 orbital files from website B. The files are in kml format, which is the standard Google Earth format and can be imported into Google Earth for viewing.

[0068] A KML file describes the geographic data of a track at a specific moment. The "LineString" element represents the geographic location of the path, defined in coordinates of longitude, latitude, and altitude within the "coordinates" tag, where each set of numbers represents longitude, latitude, and altitude in that order. Setting "altitudeMode" to "clampToGround" indicates that the path's altitude is fixed relative to the ground. This file is used to display the precise location of the geographic path on a map.

[0069] CSV files use commas as delimiters, can be quickly read using the iostream library in C++, and have the advantages of easy parsing and wide compatibility. CSV files have a simple structure, can efficiently store and transmit tabular data, and occupy less space, making them easier to manage and debug compared to complex database formats.

[0070] The KML database was parsed into CSV format using the xml library in Python code. The CSV file stores the data for each orbital segment and is named "01-Feb-2019_001131.csv" based on the time. To facilitate subsequent retrieval, the orbital data was stored in a folder as ICESat-2 Earth scan cycle data (approximately 91 days) and used as the subsequent orbital time parsing part.

[0071] Step 5 specifically involves:

[0072] The user wants to find a track with a start time of t1 and an end time of t2. Since tracks are stored by name based on time, the search can be done by year-month-day-hour-minute-second. The year and month determine which folder the track data is stored in. Then, the track data in the folder is sorted by time. The start time t1 is located by searching by date and hour-minute-second. The end time t2 is located using the same search method. After finding the track locations for t1 and t2, the track data in the intermediate time period between t1 and t2 is merged to form the track search result, which is used for subsequent track queries.

[0073] Furthermore, to further limit the range of track query results, upper and lower limits of latitude and longitude can be set to filter the search results. In this embodiment, the longitude range is set to 119.45995°E to 119.46145°E, and the latitude range is set to 34.69238°N to 34.70000°N. Based on these ranges, all track points are filtered, and only track data within these ranges are included in the query results. Each track point is iterated over, and its latitude and longitude are checked against the given upper and lower limits. If the conditions are met, the track point is included in the final query results, and the final generated track also falls within the given latitude and longitude range.

[0074] The data spans a large latitudinal range, making it difficult to accurately locate specific time periods using red-black trees. Finding points at the edges of latitude and longitude regions requires searching for the nearest neighbor data. Real-time searching of ATL data results and storage are crucial. Figure 4 As shown.

[0075] The method for parsing H5 format standard point clouds using the HDF5 library in step 6 is as follows:

[0076] HDF5 is a file format and library for efficiently managing large-scale data. It supports multiple data types and parallel reading and writing, and is widely used in scientific computing and other fields. By combining C++ with the HDF5 library, it can directly read and parse information such as photon position, timestamp, and distance along the orbit in ATL03 level data. The structured access of HDF5 improves reading efficiency and provides reliable support for the accurate analysis and simulation of Earth observation data. It can also parse the distance along the orbit and elevation corresponding to photon events in ATL03 level data.

[0077] like Figure 5 As shown, step 7 saves the along-track distance and elevation of all photon events from step 6 and plots them in a two-dimensional chart. Specifically, it includes the following steps:

[0078] Step 7.1: Construct a kd-tree based on the DBSCAN clustering algorithm to accelerate neighborhood queries and improve efficiency. The core parameters of the DBSCAN clustering algorithm are as follows: Figure 2 As shown, the algorithm first initializes the core parameters (neighborhood radius and core point threshold) required for density clustering through input parameters, and converts the original data into a data format that adapts to the algorithm to generate a point set;

[0079] Step 7.2: Construct a kd-tree to efficiently find the neighbors of each point within the neighborhood radius, and record the neighbor information of the point set in Step 7.1. In the main loop, initialize the clustering label, noise label and visit label, and check whether each point has been visited.

[0080] Step 7.3: For the points not visited in Step 7.2, obtain their neighbor set. If the number of neighbors is less than the core point threshold, mark the point as noise.

[0081] 7.4 For the points visited in step 7.2, treat them as core points and initiate a new cluster expansion. The expanded cluster recursively processes the neighbors of the core points, assigning unclassified points to the current cluster and merging neighbors that meet the conditions, until all points are classified or labeled as noise.

[0082] This invention enables rapid lookup of all ATL03 data from 2019 to 2021, totaling 245.4TB of data. It generates all ground trajectory tracks from 12:00:51 on October 28, 2021 to 7:40:28 on October 29, 2021, taking 5.9 seconds to generate the tracks. A total of 70,776 points were retrieved and plotted on the interface. The database query took 15 seconds, yielding 92 h5 results. Parsing one selected ATL03 data point to generate a point cloud and track took 5.9 seconds, resulting in 79,428 photon signals and 202,500 track points.

[0083] The standard point cloud result was obtained by analyzing the ATL03 level data at 3:13:20 on September 22, 2019. The result is as follows. Figure 5 As shown, relying on a certain laser end-to-end simulation software, the point cloud results of the ATL03 level data analysis trajectory simulation can be obtained. Comparing the point cloud confidence of the two, the error of the average signal photon number is 6.22%, and the noise rate error is 8.70%. The time spent searching the database is less than that of the online database, which shows that the present invention has real-time and efficient search efficiency. When verifying the point cloud accuracy, the point cloud extracted by the denoising algorithm has a high accuracy.

Claims

1. A method for verifying the accuracy of simulated point clouds based on offline querying of ICESat-2 ATL data, characterized in that, Includes the following steps: Step 1: Acquire and store ICESat-2 ATL data to build the base database; Step 2: Use a batch script to store the ICESat-2 ATL data name obtained in Step 1 into a text file; Step 3: Put the ICESat-2 ATL data names stored in the txt text file in Step 2 into a red-black tree to construct key-value pairs; Step 4: Obtain standard ICESat-2 orbital data and store it in a CSV file named after the orbital start time; Step 5: Combine the orbit time and latitude / longitude of the ICESat-2 orbit data obtained in Step 4, and search within the key-value pairs constructed in Step 3, using orbit time as the driving force. The specific steps are as follows: The folder for storing track data is determined by year and month. Then, the track data in the folder is sorted by time. The start time t1 of the track data is located by searching by date, hour, minute and second. The end time t2 of the track data is located by searching in the same way. After finding the track positions of t1 and t2, the track data in the intermediate time period between t1 and t2 is merged to form the track search result, which is used for subsequent track queries. The search results for the orbit can be filtered by setting upper and lower limits for latitude and longitude: First, set the latitude and longitude range, then iterate through each orbit point and check whether its latitude and longitude are within the given upper and lower limits. At the same time, add a function to search for nearby data for points on the edge of the latitude and longitude area. If the nearest data is found and is within the set range, the orbit point will be included in the final query results. The final generated orbit is also within the latitude and longitude range and is stored. Step 6: Compile the HDF5 library to parse the photon events in the search results of Step 5, compile the visualization interface, and add track and query display functions to the visualization interface to obtain the point cloud image; Step 7: Add a lidar point cloud map display interface, display the point cloud image parsed in Step 6 in the chart, and use the DBSCAN clustering-based confidence algorithm to evaluate and analyze the simulation results.

2. The method for verifying the accuracy of simulated point clouds based on offline query of ICESat-2 ATL data according to claim 1, characterized in that, Step 2 specifically includes: Name the HDF5 format ICESat-2 ATL data file in the format "Data Product Type_Acquisition Time_Track Number_Data Version Number_Reprocessing Times", then store the corresponding ATL data file name in txt format, and use the acquisition time as the primary key to query the specific file path where the ICESat-2 ATL data is located, and parse the ICESat-2 ATL data. The data product types include ATL01, ATL02, ATL03, ATL04, ATL05, ATL06, ATL07, ATL08, ATL09, ATL10, ATL11, ATL12, and ATL13 formats, depending on the content and processing level of the data. The data collection time includes year, month, day, hour, minute, and second; The track number serves as the identification label for the track; Data version numbers 001, 002, 003, ..., 00n indicate the revision number; Reprocessing is indicated by 00, 01, 02, ..., 0n, signifying whether the orbital data has been processed. Where n is an integer.

3. The method for verifying the accuracy of simulated point clouds based on offline query of ICESat-2 ATL data according to claim 1, characterized in that, Step 3 specifically includes: Read the absolute paths of all ICESat-2 ATL data stored in the plaintext txt file generated in step 2, and store them in a red-black tree for fast access.

4. The method for verifying the accuracy of simulated point clouds based on offline query of ICESat-2 ATL data according to claim 1, characterized in that, Step 4 specifically includes: Download all ICESat-2 track files, which are in kml format; The Python code uses the xml library to parse the ICESat-2 orbital files into CSV format. The CSV files store the data for each orbital segment and are named in the format "day-month-year_hour-minute-second.csv" for easy retrieval later. The orbital data is stored in a folder according to the ICESat-2 Earth scanning cycle and will be used as the orbital time parsing part later.

5. The method for verifying the accuracy of simulated point clouds based on offline query of ICESat-2 ATL data according to claim 1, characterized in that, Step 6 specifically includes: By using C++ and the HDF5 library, the photon position, timestamp, and distance along the track information in ATL03 level data are directly read and parsed. Point cloud images are obtained by parsing the distance along the track and elevation corresponding to photon events in ATL03 level data.

6. The method for verifying the accuracy of simulated point clouds based on offline query of ICESat-2 ATL data according to claim 1, characterized in that, Step 7 specifically includes: Step 7.1: Construct a kd-tree based on the DBSCAN clustering algorithm to accelerate neighborhood query and improve efficiency. The algorithm first initializes the core parameters required for density clustering through input parameters, and converts the original data into a data format that adapts to the algorithm to generate a point set. The core parameters include: neighborhood radius and core point threshold. Step 7.2: Construct a kd-tree to find the neighbors of each point within the neighborhood radius, and record the neighbor information of the point set in Step 7.

1. In the main loop, initialize the cluster label, noise label and visit label, and check whether each point has been visited. Step 7.3: For the points not visited in Step 7.2, obtain their neighbor set. If the number of neighbors is less than the core point threshold, mark the point as noise. Step 7.4: For the points visited in Step 7.2, take them as core points and start a new cluster expansion. The expanded cluster recursively processes the neighbors of the core points, assigns unclassified points to the current cluster, and merges neighbors that meet the conditions until all points are classified or are marked as noise.

Citation Information

Patent Citations

  • An efficient organization and query method for global ICESat / GLAS point cloud

    CN109947884A

  • Satellite-borne photon counting laser radar ground elevation extraction method for complex terrain

    CN117876888A