Geological survey data modeling system based on data lake and storage method, device and equipment
By using a data lake-based geological survey data modeling system, a file hierarchy structure is constructed and stored in MinIO and PostgreSQL, solving the problem of a single standard for geological survey data submission and achieving differentiated support for different geological work methods and efficient data management.
Patent Information
- Application Number
- CN202510801760.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-10-28
AI Technical Summary
The existing geological survey data submission standards are uniform and lack differentiated support for different geological work methods, resulting in information loss, incomplete expression, or difficulty in reuse.
The geological survey data modeling system based on a data lake acquires the geological survey work methods corresponding to the file sets, constructs a file hierarchy structure, manages the file sets hierarchically, stores the standardized file sets in MinIO and PostgreSQL, extracts application data and stores it in a thematic application database, and supports multi-source heterogeneous data management.
It enables differentiated support for different geological survey methods, improves the integrity and readability of data submission, and supports unified management and efficient application of multi-source heterogeneous data.
Smart Images

Figure CN120849356A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of data management technology, and in particular to a geological survey data modeling system and storage method, apparatus and equipment based on a data lake. Background Art
[0002] Currently, the mechanisms for the submission and management of various geological survey data generally suffer from a high degree of standardization but insufficient adaptability. Existing systems typically employ uniform data formats and submission requirements, failing to adequately consider the differences in data types, acquisition methods, and accuracy requirements among various geological work methods (such as geophysical exploration, geochemical sampling, and drilling sampling). This leads to situations where some specialized data suffers from information loss, incomplete representation, or difficulty in reuse during the submission process. Therefore, it is evident that existing geological survey data submission standards are singular and lack differentiated support for different geological work methods. Summary of the Invention
[0003] The purpose of this invention is to provide at least one geological survey data modeling system and storage method, device and equipment based on data lake, which can at least solve the technical problem that the existing geological survey data submission standards are single and lack differentiated support for different geological work methods, and can at least achieve the technical effect of differentiated support for different geological work methods when submitting geological data.
[0004] To address the aforementioned technical problems, at least one embodiment of this application provides a geological survey data modeling system and storage method based on a data lake, comprising: acquiring a file set of geological survey projects to be submitted; determining the geological survey work method corresponding to the file set, wherein the file set corresponds one-to-one with the geological survey projects to be submitted; acquiring a file hierarchy structure corresponding to the geological survey work method, wherein the file hierarchy structure is used to construct hierarchical relationships between files in the file set in a hierarchical manner, so as to construct a modeling system for the file set based on the hierarchical relationships; constructing hierarchical relationships between files in the file set in a hierarchical manner based on the file hierarchy structure, thereby obtaining a standardized file set corresponding to the geological survey work method; storing the standardized file set in MinIO; and storing at least one subset of files of a preset file type in the standardized file set in PostgreSQL.
[0005] This solution standardizes the file set corresponding to each project by using the file hierarchy structure corresponding to the geological survey methods of each project. This ensures that the file organization structure of the file set is uniform for the same geological survey method, which facilitates unified management. Furthermore, the file organization structure of the file set is different for different geological survey methods, thus providing differentiated support for different geological survey methods.
[0006] In some examples, the document hierarchy includes a primary document hierarchy and a secondary document hierarchy. The secondary document hierarchy is determined based on the characteristics of geological survey methods, building upon the primary document hierarchy. This hierarchical structure constructs the hierarchical relationships between documents in the document set, resulting in a standardized document set corresponding to the geological survey methods. This includes: dividing the documents in the document set according to the primary document hierarchy to obtain a primary raw document set and a primary result document set; dividing the documents in the primary raw document set according to the secondary document structure corresponding to the primary raw document set to obtain multiple secondary raw document sets; dividing the documents in the primary result document set according to the secondary document structure corresponding to the primary result document set to obtain multiple secondary result document sets; and finally, the document set consisting of the primary raw document set, the primary result document set, multiple secondary raw document sets, and multiple secondary result document sets, conforming to the document hierarchy structure, is considered the standardized document set corresponding to the geological survey methods.
[0007] In some examples, storing at least one subset of files of a preset file type from a normalized file set into PostgreSQL includes: selecting at least one subset of files of a preset file type from the normalized file set; selecting spatial data files and attribute data files from the at least one subset of files; storing the spatial data files and attribute data files into a geological database of PostgreSQL; extracting application data from the at least one subset of files based on preset application requirements to form application data files; and storing the application data files into a thematic application database of PostgreSQL, wherein the thematic application database stores data files for different application requirements.
[0008] In some examples, the method also includes: extracting tile data from a spatial data file to form a tile data file; and storing the tile data file in MongoDB.
[0009] In some examples, the method also includes: extracting vector data and attribute table data from spatial data files and attribute data files; and storing the vector data and attribute table data in ElasticSearch.
[0010] In some examples, the method also includes: acquiring the geological survey data files to be submitted, and determining at least one set of files for a geological survey project based on the geological survey data files; dividing the set of files for each geological survey project based on multiple target file types to obtain a file subset that corresponds one-to-one with the multiple target file types; and using the file subset that corresponds one-to-one with the multiple target file types as the set of files for the geological survey project to be submitted.
[0011] In some examples, the method further includes: determining the set of file set element values corresponding to the file set based on the file set metadata; determining the set of file element values corresponding to each file in the file set based on the file set metadata, wherein the file metadata inherits from the file set metadata; for each file, adding the file's unique identifier to the file element value set corresponding to the file to obtain a new set of file element values; and storing the new set of file element values and the set of file set element values in PostgreSQL, where the file's unique identifier is used to identify the file corresponding to the set of file element values.
[0012] At least one embodiment of this application also provides a geological survey data modeling system and storage device based on a data lake, comprising: an acquisition unit, configured to acquire a file set of geological survey projects to be submitted, and determine the geological survey work method corresponding to the file set, wherein the file set corresponds one-to-one with the geological survey projects to be submitted; the acquisition unit is further configured to acquire a file hierarchy structure corresponding to the geological survey work method, wherein the file hierarchy structure is used to construct hierarchical relationships between files in the file set in a hierarchical manner, so as to construct a modeling system for the file set based on the hierarchical relationships; a construction unit, configured to construct hierarchical relationships between files in the file set in a hierarchical manner based on the file hierarchy structure, thereby obtaining a standardized file set corresponding to the geological survey work method; a storage unit, configured to store the standardized file set in MinIO; and the storage unit is further configured to store at least one subset of files of a preset file type in the standardized file set in PostgreSQL.
[0013] At least one embodiment of this application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the above-described data lake-based geological survey data modeling system and storage method.
[0014] At least one embodiment of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described data lake-based geological survey data modeling system and storage method. Attached Figure Description
[0015] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.
[0016] Figure 1 This is a flowchart illustrating a geological survey data modeling system and storage method based on a data lake, as provided in one embodiment of this application.
[0017] Figure 2This is a schematic diagram of the file organization structure of a file set provided in one embodiment of this application;
[0018] Figure 3 This is a schematic diagram of a geological data file organization and storage architecture based on a data lake, provided in one embodiment of this application;
[0019] Figure 4 This is a schematic diagram of the metadata organization structure provided in one embodiment of this application;
[0020] Figure 5 This is a schematic diagram of a standardized file set's file organization structure provided in one embodiment of this application;
[0021] Figure 6 This is a schematic diagram of a geological survey data modeling system and storage device based on a data lake, provided in another embodiment of this application;
[0022] Figure 7 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the various embodiments of this application to help readers better understand this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for the convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.
[0024] It should be noted that the acquisition or use of data in the embodiments of this application requires the user's consent. The relevant data can only be obtained after the user's authorization, and the acquisition or use of the data complies with the provisions of relevant laws and regulations.
[0025] To facilitate understanding of the embodiments of this application, the relevant content regarding the geological survey data modeling system and storage method based on data lake will be introduced first.
[0026] Currently, there are already some standardized and policy requirements for the management of geological survey data submission, which have played an important role in practical management. The General Office of the Ministry of Natural Resources has made standardized requirements for the management of geological data submission. When submitting geological data, the types of results and original data must be uniformly registered. Results include text, approvals, maps, tables, attachments, databases, software, multimedia, and others. Original data includes base data, survey data, observation data, exploration data, sample data, test data, recording data, images, comprehensive data, and text data. A list of required documents has also been specified. The Geological Cloud Platform has formulated technical requirements for geological survey data aggregation. Data aggregation is organized according to a three-level directory. The first-level directory is divided into nine major categories based on the types of data and products of "Geological Cloud": geological spatial data, geological maps, geological science popularization, geological databases, publications, technical methods and standards, software, instruments and equipment, and databases. The aggregated data file types are divided into four categories: "ordinary files, GIS engineering files, MapGIS engineering files, digital mapping data packages, and data products," including various data files, data tables, GIS layer files, and metadata. In addition, the geological cloud platform integrates data services from multiple geological professional fields in the management and application of core geological data resources (including structured data such as spatial data and attribute data), providing basic browsing and query capabilities.
[0027] However, current mechanisms for the collection and management of various geological survey data generally suffer from high standardization but insufficient adaptability. Existing systems typically employ uniform data formats and submission requirements, failing to adequately consider the differences in data types, acquisition methods, and accuracy requirements among various geological work methods. This leads to situations where some specialized data suffers from information loss, incomplete representation, or difficulty in reuse during the collection process. Furthermore, current metadata systems for geological survey data mostly only describe data up to the project level, lacking fine-grained descriptions of specific datasets, data records, and even fields, making it difficult to support refined management and efficient application of geological data. Regarding data aggregation, structured spatial and attribute data are mostly presented in the form of service interfaces, lacking centralized storage and integration at the unified entity database level. While some systems have achieved thematic data aggregation by professional category, they have still failed to form a cross-professional overall database system, resulting in severe data silos and hindering the improvement of data sharing and comprehensive analysis capabilities.
[0028] In summary, the existing mechanisms for the collection and management of geological survey data suffer from the following main problems: data collection standards are uniform, lacking differentiated support for different geological work methods; metadata granularity is too coarse, failing to achieve precise description and retrieval of data content in various files; structured data mostly exists in the form of service interfaces, lacking unified and scalable entity database support; and data aggregation is limited to professional categories, failing to form an overall database architecture covering multi-source heterogeneous data. Therefore, a new geological survey data collection and management scheme is urgently needed to address these problems.
[0029] To address the technical problems of existing geological survey data submission standards being singular and lacking differentiated support for different geological work methods, this invention proposes a geological survey data modeling system and storage method based on a data lake. The implementation details of the geological survey data modeling system and storage method based on a data lake in this embodiment are described below. The following content is only for ease of understanding and is not necessary for implementing this solution.
[0030] Example 1:
[0031] The geological survey data modeling system and storage method based on a data lake described in this embodiment can be applied to electronic devices with communication, computing, and data storage capabilities. The specific process can be as follows: Figure 1 As shown, it includes:
[0032] Step 110: Obtain the document set of the geological survey projects to be submitted, determine the geological survey working method corresponding to the document set, and ensure that the document set corresponds one-to-one with the geological survey projects to be submitted.
[0033] Specifically, "geological survey projects to be submitted" refers to geological survey projects that need to be submitted. A geological survey project is a task undertaken with a clear objective. Examples include mineral exploration projects aimed at acquiring mineral resources, and geological hazard assessment projects aimed at solving specific geological problems.
[0034] Specifically, the document set contains multiple documents, each of which refers to documents obtained when conducting the same geological survey project to be submitted in the same work area.
[0035] Specifically, geological survey methodology refers to the work methods used to execute a geological survey project to be submitted within a specific area and obtain the corresponding document set for that project. For example, Figure 2 As shown, geological survey methods are categorized into over a hundred methods, including route geological mapping, geological profile measurement, remote sensing interpretation, surveying (topographic mapping), drilling engineering, mountain engineering, monitoring, geophysics, geochemistry, experimental testing, and databases. When executing corresponding geological survey projects based on various geological survey methods, the acquired files belong to the file sets corresponding to the geological survey projects under each method, such as... Figure 2 The actual material maps, geological profile summaries, remote sensing interpretation maps, control network distribution maps, borehole columnar sections, engineering sketches, observation point distribution maps, gravity and gas measurements, sample analyses, etc. shown are all file data obtained from the file set when carrying out geological survey projects to be submitted, based on the geological survey work methods of various regions.
[0036] In some examples, the process involves obtaining the document set of a geological survey project to be submitted, determining the geological survey methodology corresponding to the document set, including: obtaining the document set of a geological survey project to be submitted, obtaining the geological survey methodology for the geological survey project to be submitted, determining the geological survey methodology corresponding to the document set, and establishing the correspondence between the geological survey methodology and the document set.
[0037] In some cases, the geological survey project to be submitted is often one of multiple projects within a single geological survey, typically contained within a geological survey data file comprised of multiple projects. To facilitate the management of these data files and improve their readability, the method further includes: acquiring the geological survey data files to be submitted and identifying at least one set of files for each geological survey project based on these files; dividing the file sets for each geological survey project based on multiple target file types to obtain file subsets that correspond one-to-one with each of the target file types; and using these file subsets as the file set for the geological survey project to be submitted. Therefore, dividing the data in the geological survey data files into multiple file sets according to different projects, resulting in file sets that correspond one-to-one with each project, facilitates project-based management of geological survey data and improves its readability. Furthermore, further fine-grained division of file sets under different projects using target file types yields a more granular file organization structure, which is beneficial for managing the file sets.
[0038] Specifically, a geological survey data file is a collection of data files from geological survey projects conducted for one or more purposes. If a geological survey data file is a collection of data files from a geological survey project conducted for one purpose, then the geological survey data file is a collection of files from a geological survey project to be submitted. If a geological survey data file is a collection of data files from geological survey projects conducted for multiple purposes, then the geological survey data file contains multiple collections of files from geological survey projects to be submitted, i.e., there are multiple collections of files.
[0039] In some examples, multiple target file types include document types, multimedia file types, image file types, vector file types, professional file types, and table file types. The corresponding file subsets for each of these target file types are document file sets, multimedia file sets, image file sets, vector file sets, professional file sets, and table file sets. Specifically, document type files refer to text-based files, and files of this type in the file set constitute a document file set. Multimedia type files refer to files that store audio, video, or mixed media content, and files of this type in the file set constitute a multimedia file set. Image type files refer to files that store still images or raw sensor data, and files of this type in the file set constitute an image file set. Vector type files refer to files that describe graphics using mathematical formulas, and files of this type in the file set constitute a vector file set. Professional file type files refer to files that store data in a specific domain using a data storage format specific to that domain, and files of this type in the file set constitute a professional file set. Table type files refer to files that primarily use structured data, storing information organized by rows and columns, and files of this type in the file set constitute a table file set.
[0040] In some examples, the file formats in the document collection include, but are not limited to, .docx, .doc, .pdf, and .txt. The file formats in the multimedia collection include, but are not limited to, .jpg, .png, and .tif. The file formats in the vector collection include, but are not limited to, .shp and .gbd. The file formats in the image collection include, but are not limited to, .img and .tiff. The file formats in the table collection include, but are not limited to, .xls, .xlsx, and .db. The file formats in the professional collection include, but are not limited to, .grd and .srf.
[0041] For example, such as Figure 3 As shown, taking geological survey data files comprising n projects as an example, the file sets corresponding to each project are divided based on multiple target file types, resulting in file sets for each project that include document sets, multimedia file sets, vector file sets, image file sets, table file sets, and professional file sets. Each file subset may contain files of different file formats.
[0042] Step 120: Obtain the file hierarchy structure corresponding to the geological survey working method. The file hierarchy structure is used to construct the hierarchical relationship between the files in the file set in a hierarchical manner, so as to construct the modeling system of the file set based on the hierarchical relationship.
[0043] Specifically, a modeling system refers to an abstract modeling of the hierarchical relationships between files in a file set, which is used to construct a data model to support the storage of the file set, as well as its systematic organization and structured management.
[0044] Specifically, the file hierarchy is used to manage files in a file set hierarchically. By organizing and managing the file set in a hierarchical manner, the file set has a unified organizational structure, which facilitates centralized management.
[0045] Specifically, different geological survey methods have different document hierarchical structures, while the same geological survey method has the same document hierarchical structure.
[0046] Step 130: Based on the file hierarchy structure, construct the hierarchical relationship between the files in the file set to obtain a standardized file set corresponding to the geological survey work method.
[0047] In some examples, in step 130 above, the document hierarchy includes a first-level document hierarchy and a second-level document hierarchy. The second-level document hierarchy is determined based on the characteristics of the geological survey methodology, building upon the first-level document hierarchy. This involves constructing hierarchical relationships between documents in the document set according to the hierarchical structure, resulting in a standardized document set corresponding to the geological survey methodology. This includes: dividing the documents in the document set according to the first-level document hierarchy to obtain a first-level original document set and a first-level result document set; dividing the documents in the first-level original document set according to the second-level document structure corresponding to the first-level original document set to obtain multiple second-level original document sets; dividing the documents in the first-level result document set according to the second-level document structure corresponding to the first-level result document set to obtain multiple second-level result document sets; and using the document set composed of the first-level original document set, the first-level result document set, multiple second-level original document sets, and multiple second-level result document sets, conforming to the document hierarchy structure, as the standardized document set corresponding to the geological survey methodology.
[0048] Specifically, the first-level document hierarchy includes two main categories: original documents and output documents. The second-level document hierarchy refers to a more granular document set division structure, based on the characteristics of geological survey methodologies, and further refined from the original and output document sets. Specifically, it is based on... Figure 2 For example, the working method of the file set corresponds to the file hierarchy structure A. The file hierarchy structure A includes two levels. The first level (i.e., the first-level file level) is divided into the original file set and the result file set. The second level (i.e., the second-level file level) groups the original file set and the result file set respectively, resulting in n second-level original file sets for the original file set and n second-level result file sets for the result file set.
[0049] Specifically, if Figure 2 As shown, for each file set at the second-level file hierarchy, specifications can be defined according to the characteristics of geological survey methods. These specifications can include whether each file set can be empty and the file formats it can contain. Each file can be a single file or a composite file. A composite file mainly refers to a complete professional data file composed of multiple file formats. For example, a complete vector file in spatial data is composed of multiple file entities such as *.shp, *.prj, *.shx, *.pbf, *.cpg, *.sbn, and *.sbx.
[0050] Specifically, the original file set contains at least one original file. An original file refers to first-hand data obtained directly from the data source, without any processing, transformation, or modification. Specifically, original files may include several categories such as baseline, survey, observation, exploration, sample, test, recording, image, summary, and document. For relevant content regarding baseline, survey, observation, exploration, sample, test, recording, image, summary, and document, please refer to existing technologies, which will not be elaborated further here.
[0051] Specifically, the deliverables collection contains at least one deliverable, which refers to a data product generated after a series of processing steps. These processing steps may include, but are not limited to, data cleaning, format conversion, analysis and calculation, and visualization. Specifically, deliverables include several categories of documents such as main text, approvals, figures, tables, attachments, databases, software, multimedia, and others. For details regarding main text, approvals, figures, tables, attachments, databases, software, and multimedia, please refer to existing technologies; further details will not be elaborated here.
[0052] For example, taking the geological survey methodology corresponding to File Set 1 as route geological mapping, the first-level file hierarchy of this file is the first-level original file set and the first-level result file set. Therefore, File Set 1 is hierarchically divided based on the first-level file hierarchy. The original files in File Set 1 are taken as files in the first-level original file set, resulting in the first-level original file set. The result files in File Set 1 are taken as files in the first-level result file set, resulting in the first-level result file set. Based on the characteristics of the geological survey methodology, the first-level original file set is divided into a second-level file structure corresponding to the first-level original file set, consisting of an observation point record file set, a GPS trajectory data file set, a field photo file set, and a sample number list file set. The observation point record file set contains at least one observation point file, which is a file that records the location coordinates, lithological description, structural characteristics, and other information of the observation point. The GPS trajectory data file set contains at least one GPS trajectory data file, which is used to record the coordinates of the route walked by the surveyors to generate path lines or verify location accuracy. The field photo file set contains at least one field photo, which is a photo taken on-site showing rocks, structural features, and stratigraphic contact relationships. The sample number list file set contains at least one sample number list file, which is a file that records the collected rock or mineral samples and their location information. Based on the secondary file structure corresponding to the primary raw file set, the primary raw file set of File Set 1 is divided into the observation point record file set, the GPS trajectory data file set, the field photograph file set, and the sample number list file set. Specifically, the files belonging to the observation point record file set constitute the observation point record file set. The files belonging to the GPS trajectory data file set constitute the GPS trajectory data file set. The files belonging to the field photograph file set constitute the field photograph file set. The files belonging to the sample number list file set constitute the sample number list file set. The process of dividing the primary result file set of File Set 1 into multiple secondary result file sets based on the secondary file structure of the primary result file set can be found in the process of dividing the primary raw file set, and will not be elaborated further here.
[0053] Step 140: Store the normalized file set in MinIO.
[0054] Specifically, MinIO is a high-performance, distributed object storage system.
[0055] Based on the file fragment upload method, standardized geological survey data files are uploaded to the MinIO project library. Within the project library, the file sets corresponding to each project within the geological survey data files are stored, ensuring a one-to-one correspondence between each file set and the project. Specifically, storing the file sets corresponding to each project in MinIO is based on the modeling system corresponding to each file set.
[0056] like Figure 3 As shown, in the MinIO project library, each file set corresponds to a MinIO storage space. A storage space stores a file set corresponding to one project, including document sets, multimedia file sets, vector file sets, image file sets, spreadsheet file sets, and professional file sets.
[0057] For example, such as Figure 3 As shown, the file sets corresponding to each project in the geological survey data file are processed hierarchically based on the working methods (geological survey working methods) of each project to obtain standardized file sets. The standardized file sets corresponding to each project in the geological survey data file are stored in MinIO using a file fragment upload method. Specifically, they are stored in the project library within MinIO. Specifically, the standardized file sets are organized by project, and the hierarchical organization of the standardized file sets is determined based on the working methods corresponding to each project. Therefore, the file sets corresponding to each project are stored in the project library by project, and the storage of file sets for each project needs to be determined based on the working methods corresponding to different file sets. In other words, the MinIO project library needs to implement the storage of file sets by project and by working method.
[0058] Step 150: Store at least one subset of files of the preset file types in the normalized file set to PostgreSQL.
[0059] Specifically, the file subset contains at least one file, and the file format of the file subset is a preset file type.
[0060] Specifically, PostgreSQL is a powerful, open-source object-relational database management system that emphasizes scalability and standards compliance.
[0061] It should be understood that the data collected in the file sets differ depending on the geological survey methodology and the geological survey project. This difference is particularly prominent in terms of data structure and expression. Therefore, in order to improve the flexibility of data organization structure and achieve unified management of the differences in file sets of different geological survey projects, this solution proposes a data lake architecture that meets the requirements of multi-source heterogeneous data management. This data lake architecture stores the file sets in MinIO and stores the structured core data of the file sets in PostgreSQL.
[0062] In some examples, the preset file types include vector file types, professional file types, and table file types, and at least one subset of files includes a vector file set corresponding to the vector file type, a professional file set corresponding to the professional file type, and a table file set corresponding to the table file type.
[0063] For example, such as Figure 3 As shown, the vector file set, specialized file set, and tabular file set stored in the normalized file set in MinIO are stored in PostgreSQL. Specifically, a corresponding geological database exists for each different project, and at least one subset of files is stored in the geological database corresponding to that project. The geological database includes multiple geological databases for multiple projects.
[0064] In some cases, different application scenarios have their own specific business needs, thus exhibiting differentiated characteristics in data usage. File subsets often contain a large amount of multi-source heterogeneous data, much of which is irrelevant to the current specific application requirements. During analysis and processing, this redundant information not only increases the computational burden but may also interfere with the analysis process and reduce efficiency. Therefore, to better address specific application needs and improve the accuracy and efficiency of data usage, this solution proposes an application data extraction method based on the aforementioned step 150. Specifically, in step 150, at least one subset of files of a preset file type from the normalized file set is stored in PostgreSQL, including: selecting at least one subset of files of a preset file type from the normalized file set; selecting spatial data files and attribute data files from the at least one subset of files; storing the spatial data files and attribute data files in the PostgreSQL geological database; extracting application data from the at least one subset of files based on preset application requirements to form application data files; and storing the application data files in a PostgreSQL thematic application database, where the thematic application database stores data files for different application requirements. Therefore, by extracting application data relevant to preset application requirements from a subset of files to form an application data file and eliminating irrelevant data, a precise response to preset application requirements can be achieved. Simultaneously, this application data file is stored in a thematic application database, supplementing the database with additional data. The thematic application database contains various application data files, enabling precise response and efficient support for diverse application requirements. It also provides a clearly structured and content-focused data foundation for subsequent spatial modeling, visualization, and intelligent decision-making.
[0065] The spatial data file and attribute data file contain structured spatial data and attribute data.
[0066] Specifically, preset application requirements are used to address the needs of solving a certain type of problem, achieving a certain purpose, or completing a certain task. Preset application requirements include, but are not limited to, layer requirements, attribute table requirements, and data range requirements. For example, if preset application requirements include layer requirements, attribute table requirements, and data range requirements, when extracting application data from at least one subset of files to form an application data file based on the preset application requirements, the specific steps are as follows: based on the preset application requirements, application data is extracted from at least one subset of files according to the layer requirements, attribute table requirements, and data range requirements of the preset application requirements, respectively, to form the application data file.
[0067] There can be one or more preset application requirements, with each preset application requirement corresponding to a specific application database. If there are multiple preset application requirements, for each requirement, the following steps are performed: extracting application data from at least one subset of files based on the preset application requirement to form an application data file; and storing the application data file in the PostgreSQL specific application database. Specifically, storing the application data file in the PostgreSQL specific application database means storing the application data file in the specific application database corresponding to its preset application requirement.
[0068] In some cases, if there are multiple new pre-defined application requirements, the specific implementation process for constructing a new thematic application database corresponding to each new pre-defined application requirement is as follows: for each new pre-defined application requirement, multiple data corresponding to the application requirement are obtained from each geological database, and multiple data from the same geological database are combined to form a thematic application file, thus obtaining the thematic application file corresponding to the geological database; a new thematic application database corresponding to the new pre-defined application requirement is formed based on the thematic application files corresponding to each geological database.
[0069] For example, such as Figure 3 As shown, multiple thematic application databases are composed of data extracted from various geological databases.
[0070] In some examples, the method also includes: extracting tile data from a spatial data file to form a tile data file; and storing the tile data file in MongoDB.
[0071] Specifically, storing tile data files in MongoDB includes building indexes on the tile data files and storing indexed tile data files in MongoDB.
[0072] Specifically, the tile data file can be in .pbf or .png format.
[0073] Specifically, tile data includes vector tile data and raster tile data. Vector tile data refers to data in which geographic information is divided into vector formats and organized according to a certain hierarchical structure; vector tile data refers to data of vector tile type. Raster tile data refers to a series of small images of fixed size.
[0074] Specifically, MongoDB is an open-source, document-oriented database management system (DBMS) designed to provide flexible and scalable data storage solutions for modern applications.
[0075] In some cases, to support efficient data browsing, the aforementioned data lake architecture also includes MongoDB. See also... Figure 3 It extracts tile data from a PostgreSQL geological database, converts the tile data into tile data files (e.g., map tiles), and stores the tile data files in MongoDB.
[0076] In some examples, the method also includes: extracting vector data and attribute table data from spatial data files and attribute data files; and storing the vector data and attribute table data in ElasticSearch.
[0077] Specifically, storing attribute table data in Elasticsearch includes building an index on the attribute table data and storing the indexed attribute table data in Elasticsearch.
[0078] In some cases, to support efficient data querying and statistics, the aforementioned data lake architecture also includes ElasticSearch. See also... Figure 3 It extracts vector data and attribute table data from a PostgreSQL geological database and stores the vector data and attribute table data in ElasticSearch.
[0079] Specifically, ElasticSearch is a distributed, open-source search and analytics engine suitable for all types of data, including text, numbers, geospatial, structured, and unstructured data.
[0080] In some examples, to achieve a more granular metadata management approach, the following steps are also taken: determining the set of file set element values corresponding to the file set based on the file set metadata dataset; determining the set of file element values corresponding to each file in the file set based on the file set metadata dataset, wherein the file metadata dataset is inherited from the file set metadata dataset; for each file, adding the file's unique identifier to the file element value set corresponding to the file, resulting in a new set of file element values; and storing the new set of file element values and the set of file set element values in PostgreSQL, where the file's unique identifier is used to identify the file corresponding to the set of file element values.
[0081] Specifically, the metadata center stores a large amount of metadata, which refers to data that describes data. More specifically, the file set metadata center stores a large amount of file set metadata, which refers to data that describes the file set; the file metadata center stores a large amount of file metadata, which refers to data that describes the file.
[0082] Specifically, a file unique identifier refers to the element value that identifies the file to which a set of file element value values belongs; that is, it is used to determine the file described by the set of file element value values. Specifically, a file unique identifier can be a filename, a file extension, or other identifying symbols; no specific limitations are made here.
[0083] Specifically, the file metadata dataset inherits from the file set metadata dataset. The initial element values of each file's file metadata dataset are inherited from the file set element values of its respective file set, and then the inherited file set element values are modified as needed based on the actual situation of the file. In addition, the file metadata dataset adds a file unique identifier metadata on top of the file set metadata dataset. The element values of the file unique identifier metadata may contain CGSDOI.
[0084] Specifically, the metadata set is based on the "Geological Information Metadata Standard (DD 2006-05)" and can be appropriately tailored and expanded according to actual management needs. Figure 4 For example, based on the "Geological Information Metadata Standard (DD2006-05)," multiple file set metadata were identified. These metadata sets were then divided into eight subsets: metadata information subset, project information subset, identification information subset, dataset quality information subset, spatial reference system information subset, content information subset, service information subset, and reference and responsible unit information subset. The file metadata dataset inherits from this file set metadata dataset and includes these eight file metadata subsets: metadata information subset, project information subset, identification information subset, dataset quality information subset, spatial reference system information subset, content information subset, service information subset, and reference and responsible unit information subset. Additionally, the file metadata dataset includes an extra subset of file unique identifier metadata. For example... Figure 4 As shown, each subset of file set metadata can contain one or more file set metadata. Each subset of file metadata can contain one or more file metadata. Figure 4 Taking the metadata information subset as an example, the metadata information subset includes at least two file set metadata: language metadata and metadata creation date metadata. For information on the relationship between other file set metadata subsets and file set metadata, as well as the relationship between file metadata subsets and file metadata, please refer to [link to relevant documentation]. Figure 4 This will not be elaborated further here.
[0085] For example, to better understand the data organization structure of the file set in this scheme, please refer to... Figure 5A file set is a dataset of files acquired for a specific project within a specific work area from a geological survey data file. Different projects employ different working methods. The file set corresponding to a project in that work area is grouped according to its working method, resulting in multiple file groups (file groups are either secondary raw file sets or secondary result file sets). Each file group may contain zero, one, or more files, and each file corresponds to a set of file element values. Simultaneously, each file set corresponds to a set of file set element values. Specifically, the working method includes a primary working method and at least one secondary working method corresponding to the primary working method. The primary working method is used for the first grouping of the file set, resulting in multiple primary groups (primary groupings are either primary raw file sets or primary result file sets). Based on each primary group, for each primary working method corresponding to at least one secondary working method, the primary group is further grouped, resulting in multiple secondary groups (secondary groupings are either secondary raw file sets or secondary result file sets).
[0086] In summary, this solution obtains the file sets of geological survey projects to be submitted, determines the geological survey methodology corresponding to each file set, and establishes a one-to-one correspondence between the file sets and the geological survey projects to be submitted. It obtains the file hierarchy structure corresponding to the geological survey methodology, which is used to construct the hierarchical relationships between files in the file set, thereby building a modeling system for the file set based on these hierarchical relationships. Based on the file hierarchy structure, it constructs the hierarchical relationships between files in the file set, resulting in a standardized file set corresponding to the geological survey methodology. The standardized file set is stored in MinIO. At least one subset of files of preset file types in the standardized file set is stored in PostgreSQL. By standardizing the file sets corresponding to the geological survey methodology for each project through the file hierarchy structure, the file organization structure of the file set is unified for the same geological survey methodology, facilitating unified management. Furthermore, the file organization structure of the file set differs for different geological survey methodology, achieving differentiated support for different geological survey methodology, improving the flexibility of the file organization structure of different file sets, and enhancing the adaptability of the file organization structure of each file set to the business of different projects.
[0087] Furthermore, the use of diverse and heterogeneous data management ensures the integrity of geological survey data after submission. Combined with the multi-source heterogeneous data lake architecture set up in this solution, differentiated support for different geological survey methodologies is achieved. For geological survey projects of different professional types to be submitted, the file organization structure of the project's file set is standardized according to the corresponding geological survey methodology. This ensures that when managing geological data submission, dimensions such as the professional category of each geological data point, geological survey project, and geological survey methodology for each project are taken into account. By constructing file set metadata datasets and file metadata datasets, more granular metadata management of the file set is achieved, which is conducive to refined management of geological data. Precise description of the data content of the files facilitates the retrieval of specific files. Moreover, the structured core data of the file set is stored in PostgreSQL, providing entity database support for structured data.
[0088] Example 2:
[0089] Another embodiment of this application relates to a geological survey data modeling system and storage device based on a data lake. The implementation details of this embodiment's geological survey data modeling system and storage device are described below. The following details are provided for ease of understanding and are not essential for implementing this solution. A schematic diagram of the geological survey data modeling system and storage device 60 based on a data lake in this embodiment can be seen as follows: Figure 6 As shown, it includes an acquisition unit 601, a construction unit 602, and a storage unit 603.
[0090] The acquisition unit 601 is used to acquire the document set of the geological survey project to be submitted, determine the geological survey work method corresponding to the document set, and the document set corresponds one-to-one with the geological survey project to be submitted.
[0091] The acquisition unit 601 is also used to acquire the file hierarchy structure corresponding to the geological survey work method. The file hierarchy structure is used to construct the hierarchical relationship between each file in the file set in a hierarchical manner, so as to construct a modeling system for the file set based on the hierarchical relationship.
[0092] Construction unit 602 is used to construct the hierarchical relationship between files in the file set based on the file hierarchy structure, so as to obtain a standardized file set corresponding to the geological survey work method.
[0093] Storage unit 603 is used to store the normalized file set into MinIO.
[0094] Storage unit 603 is also used to store at least one subset of files of a preset file type in a normalized file set to PostgreSQL.
[0095] In some examples, the document hierarchy includes a first-level document hierarchy and a second-level document hierarchy. The second-level document hierarchy is determined based on the characteristics of the geological survey methodology, building upon the first-level document hierarchy. Specifically, when the aforementioned construction unit 602 is used to construct the hierarchical relationships between documents in the document set based on the document hierarchy to obtain a standardized document set corresponding to the geological survey methodology, it is used to: divide the documents in the document set according to the first-level document hierarchy to obtain a first-level original document set and a first-level result document set; divide the documents in the first-level original document set according to the second-level document structure corresponding to the first-level original document set to obtain multiple second-level original document sets; divide the documents in the first-level result document set according to the second-level document structure corresponding to the first-level result document set to obtain multiple second-level result document sets; and use the document set that conforms to the document hierarchy structure, composed of the first-level original document set, the first-level result document set, multiple second-level original document sets, and multiple second-level result document sets, as the standardized document set corresponding to the geological survey methodology.
[0096] In some examples, when the aforementioned storage unit 603 is used to store at least one subset of files of a preset file type from the normalized file set to PostgreSQL, it is specifically used to: select at least one subset of files of a preset file type from the normalized file set; select spatial data files and attribute data files from the at least one subset of files; store the spatial data files and attribute data files to the geological database of PostgreSQL; extract application data from the at least one subset of files based on preset application requirements to form an application data file; and store the application data file to a thematic application database of PostgreSQL, wherein the thematic application database stores data files for different application requirements.
[0097] In some examples, the aforementioned storage unit 603 is also used to: extract tile data from a spatial data file to form a tile data file; and store the tile data file in MongoDB.
[0098] In some examples, the aforementioned storage unit 603 is also used to: extract vector data and attribute table data from spatial data files and attribute data files; and store the vector data and attribute table data in ElasticSearch.
[0099] In some examples, the aforementioned acquisition unit 601 is also used to: acquire geological survey data files to be submitted, and determine at least one set of files for a geological survey project based on the geological survey data files; divide the set of files for each geological survey project based on multiple target file types to obtain a file subset that corresponds one-to-one with the multiple target file types; and use the file subset that corresponds one-to-one with the multiple target file types as the set of files for the geological survey project to be submitted.
[0100] In some examples, the aforementioned storage unit 603 is also used to: determine the file set element value set corresponding to the file set based on the file set metadata dataset; determine the file element value set corresponding to each file in the file set based on the file set metadata dataset, wherein the file metadata dataset is inherited from the file set metadata dataset; for each file, add the file unique identifier to the file element value set corresponding to the file to obtain a new file element value set; and store the new file element value set and the file set element value set in PostgreSQL, wherein the file unique identifier is used to determine the file corresponding to the file element value set.
[0101] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units are absent in this embodiment.
[0102] Example 3:
[0103] Another embodiment of this application relates to an electronic device, such as... Figure 7 As shown, it includes: at least one processor 901; and a memory 902 communicatively connected to at least one processor 901; wherein the memory 902 stores instructions executable by at least one processor 901, the instructions being executed by at least one processor 901 to enable at least one processor 901 to execute the data lake-based geological survey data modeling system and storage method in the above embodiments.
[0104] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0105] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0106] Example 4:
[0107] Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.
[0108] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0109] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of this application.
Claims
1. A geological survey data modeling system and storage method based on a data lake, characterized in that, include: Obtain the document set of the geological survey projects to be submitted, determine the geological survey working method corresponding to the document set, and the document set corresponds one-to-one with the geological survey projects to be submitted; Obtain the file hierarchy structure corresponding to the geological survey method. The file hierarchy structure is used to construct the hierarchical relationship between the files in the file set in a hierarchical manner, so as to construct the modeling system of the file set based on the hierarchical relationship. Based on the file hierarchy structure, the hierarchical relationship between the files in the file set is constructed in a hierarchical manner to obtain a standardized file set corresponding to the geological survey method. Store the standardized file set in MinIO; Store at least one subset of files of a preset file type in the normalized file set to PostgreSQL.
2. The geological survey data modeling system and storage method based on data lake according to claim 1, characterized in that, The document hierarchy includes a primary document hierarchy and a secondary document hierarchy. The secondary document hierarchy is determined based on the primary document hierarchy and the characteristics of the geological survey methodology. The hierarchical relationship between documents in the document set is constructed hierarchically based on this document hierarchy to obtain a standardized document set corresponding to the geological survey methodology, including: The files in the file set are divided according to the first-level file hierarchy to obtain the first-level original file set and the first-level result file set; Each file in the first-level original file set is divided according to the second-level file structure corresponding to the first-level original file set to obtain multiple second-level original file sets; Each file in the primary deliverable file set is divided according to the secondary file structure corresponding to the primary deliverable file set to obtain multiple secondary deliverable file sets; A file set that conforms to the file hierarchy structure, consisting of the first-level original file set, the first-level result file set, multiple second-level original file sets, and multiple second-level result file sets, is considered a standardized file set corresponding to the geological survey work method.
3. The geological survey data modeling system and storage method based on a data lake according to claim 1, characterized in that, The step of storing at least one subset of files of a preset file type in the normalized file set to PostgreSQL includes: Select at least one subset of files of a preset file type from the normalized file set; Select spatial data files and attribute data files from the at least one subset of files; The spatial data file and the attribute data file are stored in a PostgreSQL geological database; Based on preset application requirements, application data is extracted from at least one subset of files to form an application data file; The application data files are stored in a PostgreSQL topic application database, which contains data files for different application needs.
4. The geological survey data modeling system and storage method based on data lake according to claim 3, characterized in that, The method further includes: Tile data is extracted from the spatial data file to form a tile data file; The tile data file is stored in MongoDB.
5. The geological survey data modeling system and storage method based on data lake according to claim 3, characterized in that, The method further includes: Extract vector data and attribute table data from the spatial data file and the attribute data file; The vector data and the attribute table data are stored in ElasticSearch.
6. The geological survey data modeling system and storage method based on data lake according to claim 1, characterized in that, The method further includes: Obtain the geological survey data files to be submitted, and determine the file set of at least one geological survey project based on the geological survey data files; For the document sets of various geological survey projects, the document sets are divided based on multiple target document types to obtain document subsets that correspond one-to-one with multiple target document types; The subset of files that correspond one-to-one with the multiple target file types will be used as the file set for the geological survey project to be submitted.
7. The geological survey data modeling system and storage method based on data lake according to claim 6, characterized in that, The method further includes: The set of file set element values corresponding to the file set is determined based on the file set metadata dataset; The file element value set corresponding to each file in the file set is determined based on the file meta dataset, wherein the file meta dataset is inherited from the file set meta dataset; For each of the aforementioned files, a unique file identifier is added to the file element value set corresponding to that file, resulting in a new file element value set; The new set of file element values and the set of file element values are stored in the PostgreSQL database, and the file unique identifier is used to identify the file corresponding to the set of file element values.
8. A geological survey data modeling system and storage device based on a data lake, characterized in that, include: The acquisition unit is used to acquire the file set of geological survey projects to be submitted, determine the geological survey working method corresponding to the file set, and the file set corresponds one-to-one with the geological survey projects to be submitted. The acquisition unit is also used to acquire the file hierarchy structure corresponding to the geological survey method. The file hierarchy structure is used to construct the hierarchical relationship between each file in the file set in a hierarchical manner, so as to construct the modeling system of the file set based on the hierarchical relationship. The construction unit is used to construct the hierarchical relationship between the files in the file set based on the file hierarchy structure, so as to obtain a standardized file set corresponding to the geological survey method. A storage unit is used to store the normalized file set into MinIO; The storage unit is also used to store at least one subset of files of a preset file type in the normalized file set to PostgreSQL.
9. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to execute the geological survey data modeling system and storage method based on data lake as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the geological survey data modeling system and storage method based on data lake as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Field geological survey data real-time converging method and system
CN110716898A
Remote sensing image data management service system and method
CN118113894A