Geological data management method and system based on project path and semantic association

CN122817344APending Publication Date: 2026-09-25CHINA GEOLOGICAL SURVEY NATURAL RESOURCES COMPREHENSIVE SURVEY COMMAND CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611013519.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

但是,虽然相关部门对汇交文件要求统一,却未考虑不同地质工作方法的差异,且元数据多针对整个项目,粒度较粗,对于地质数据的精细化管理和应用支撑不足

Benefits of technology

[0039]1.通过在项目路径解析过程中提取层级特征向量并生成全局唯一的数据集编码,将原本用于人工浏览的目录树转化为可计算的编码空间,使得任意地质数据集均可通过编码前缀快速定位其所属项目层级、工作方法和空间位置;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817344A_ABST
    Figure CN122817344A_ABST
Patent Text Reader

Abstract

The application relates to the field of data processing of geological data, and specifically discloses a geological data management method and device based on project path and semantic association, which comprises the following steps: acquiring a plurality of geological data sets to be submitted; performing project path analysis on each geological data set, generating a data set code corresponding to the geological data set, and constructing metadata corresponding to the geological data set; determining a format parser corresponding to the geological data set according to the metadata, performing element-by-element analysis on the geological data set by using the format parser, and obtaining standardized element features corresponding to each element; taking the standardized element features corresponding to each element as nodes, constructing a semantic association graph based on a work area and a work method in a hierarchical feature vector; and constructing a three-layer structure geological data lake according to the plurality of geological data sets, the standardized element features and the semantic association graph, and establishing a multi-scale significance index for each element, so that fine management of the geological data is effectively realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of geological data processing, and specifically to a geological data management method and system based on project path and semantic association. Background Technology

[0002] In geological work such as regional geological surveys, mineral resource exploration, hydrogeological evaluation, engineering geological investigation, and geological hazard investigation, a large amount of geological data from multiple sources, methods, and scales is generated daily. This includes borehole logging data, geophysical exploration data (such as seismic, electromagnetic, gravity, and magnetic methods), geochemical sampling data, remote sensing interpretation data, analytical testing data, geological maps, and three-dimensional geological models. As geological survey work develops towards greater refinement, integration, and digitization, the types, formats, and quantities of this data are growing exponentially. How to efficiently organize, store, retrieve, and reuse this geological data has become a critical technical problem that urgently needs to be solved in the field of geological data management.

[0003] Currently, relevant departments have established standardized requirements for the management of geological data submission. When submitting geological data, the types of output data and original data must be uniformly registered, and a list of required documents has been specified. However, while these requirements for submission documents are standardized, they do not consider the differences in various geological work methods. Furthermore, the metadata is mostly project-wide, with a coarse granularity, which is insufficient for the refined management and application of geological data. Summary of the Invention

[0004] In view of the above-mentioned defects or deficiencies in the existing technology, it is desirable to provide a geological data management method and system based on project path and semantic association, so as to effectively realize the refined management of geological data.

[0005] In a first aspect, embodiments of this application provide a geological data management method based on project path and semantic association, including:

[0006] Obtain multiple geological datasets to be submitted;

[0007] For each geological dataset, project path parsing is performed, and the hierarchical feature vector corresponding to the geological dataset is obtained based on the geological organization tree. The dataset code corresponding to the geological dataset is generated based on the hierarchical feature vector. The geological organization tree is a directory tree constructed based on "project-method-dataset".

[0008] Based on the geological dataset, the hierarchical feature vector, and the dataset encoding, construct the metadata corresponding to the geological dataset;

[0009] Based on the metadata, determine the format parser corresponding to the geological dataset, and use the format parser to perform element-by-element parsing on the geological dataset to obtain the standardized element features corresponding to each element;

[0010] Using the standardized feature characteristics corresponding to each element as nodes, a semantic association graph is constructed based on the working area and working method in the hierarchical feature vector;

[0011] Based on multiple geological datasets, standardized feature characteristics, and semantic association graphs, a three-layer geological data lake is constructed, and a multi-scale saliency index is established for each feature; the three-layer geological data lake includes at least an original lake layer and a standard lake layer.

[0012] In some embodiments, constructing the metadata corresponding to the geological dataset based on the geological dataset, the hierarchical feature vector, and the dataset encoding includes:

[0013] Based on the geological dataset, the hierarchical feature vector, and the dataset encoding, determine the physical attribute information, overall feature description information, and element description information corresponding to the geological dataset;

[0014] Based on the physical attribute information, the overall feature description information of the dataset, and the feature description information, file mirror-level metadata, dataset-level metadata, and feature-level metadata are constructed respectively in the metadata.

[0015] In some embodiments, determining the format parser corresponding to the geological dataset based on the metadata includes:

[0016] Based on the file format type in the file mirror-level metadata of the aforementioned metadata, determine the format parser corresponding to the geological dataset.

[0017] In some embodiments, the step of using the format parser to perform element-by-element parsing of the geological dataset to obtain the standardized feature characteristics corresponding to each element includes:

[0018] The geological dataset is parsed element by element using the format parser. For each element, spatial geometric representation standardization and attribute information standardization are performed to obtain the standardized spatial geometric representation and standardized attribute set corresponding to the element.

[0019] After the element is standardized, a source tuple corresponding to the element is generated based on the unique identifier of the geological dataset, the location information of the element in the geological dataset, and the identifier of the standardization steps performed on the element.

[0020] Based on the standardized spatial geometric representation of the element, the standardized attribute set, and the source tuple, the standardized element feature corresponding to the element is generated.

[0021] In some embodiments, the step of constructing a semantic association graph based on the working area and working method in the hierarchical feature vector, using the standardized feature features corresponding to each element as nodes, includes:

[0022] For any two elements, based on the standardized element features corresponding to the elements, determine the spatial overlap, semantic similarity of working methods, and consistency of geological age between the two elements respectively;

[0023] The edge weights between the two elements are determined based on the spatial overlap, the semantic similarity of the working methods, and the consistency of the geological ages.

[0024] The semantic association graph is constructed based on the standardized element features and weights corresponding to each element.

[0025] In some embodiments, establishing a multi-scale saliency index for each of the features includes:

[0026] For each of the aforementioned elements, obtain the spatial geometric quantity, effective scale, and attribute information gain of the element;

[0027] Based on the spatial geometric quantity, the effective scale, and the attribute information gain, a multi-scale saliency index is established for the element.

[0028] Secondly, embodiments of this application provide a geological data management system based on project path and semantic association, including:

[0029] The acquisition module is used to acquire multiple geological datasets to be submitted.

[0030] The first parsing module is used to parse the project path for each of the geological datasets, obtain the hierarchical feature vectors corresponding to the geological datasets based on the geological organization tree, and generate the dataset code corresponding to the geological datasets based on the hierarchical feature vectors; the geological organization tree is a directory tree constructed based on "project-method-dataset";

[0031] The first construction module is used to construct the metadata corresponding to the geological dataset based on the geological dataset, the hierarchical feature vector, and the dataset encoding;

[0032] The second parsing module is used to determine the format parser corresponding to the geological dataset based on the metadata, and to use the format parser to perform element-by-element parsing on the geological dataset to obtain the standardized element features corresponding to each element.

[0033] The second construction module is used to construct a semantic association graph based on the working area and working method in the hierarchical feature vector, using the standardized feature features corresponding to each element as nodes.

[0034] The third construction module is used to construct a three-layer geological data lake based on multiple geological datasets, standardized feature characteristics, and semantic association graphs, and to establish a multi-scale saliency index for each feature; the three-layer geological data lake includes at least an original lake layer and a standard lake layer.

[0035] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in embodiments of this application.

[0036] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in embodiments of this application.

[0037] Fifthly, embodiments of this application provide a computer program product, including a computer program, characterized in that, when the computer program is executed by a processor, it implements the method described in embodiments of this application.

[0038] The above-described technical solutions of the embodiments of the present invention have the following beneficial technical effects:

[0039] 1. By extracting hierarchical feature vectors and generating globally unique dataset codes during the project path parsing process, the directory tree originally used for manual browsing is transformed into a computable coding space, enabling any geological dataset to quickly locate its project level, working method, and spatial location through the coding prefix;

[0040] 2. By constructing a three-tiered metadata system at the file mirror level, dataset level, and feature level, the granularity of metadata description is refined from the traditional project level to the spatial feature level that can be independently referenced, supporting multi-level data description and retrieval from macro-level project-level queries to micro-level feature-level queries;

[0041] 3. By constructing a weighted semantic association graph with standardized elements as nodes and work area spatial overlap and work method semantic similarity as core dimensions, the originally isolated project data is woven into a logically integrated knowledge network, realizing proactive data association and recommendation across projects and work methods;

[0042] 4. By introducing a multi-scale significance index, the data organization is naturally adapted to the multi-scale application scenarios of geological work from macro to micro, realizing the hierarchical loading and on-demand presentation of large-scale overview data and small-scale fine data;

[0043] 5. By constructing a three-layer physical architecture of "original lake - standard lake - application lake", complete data lifecycle management is achieved, which enables traceability of original data, computation of standard data, and reusability of application data.

[0044] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0045] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0046] Figure 1 A flowchart illustrating a geological data management method based on project path and semantic association provided in an embodiment of this application is shown.

[0047] Figure 2 This illustration shows a schematic diagram of the structure of a geological data management system based on project path and semantic association provided in an embodiment of this application;

[0048] Figure 3 A schematic diagram of the structure of a computer system suitable for implementing an electronic device or server according to embodiments of this application is shown. Detailed Implementation

[0049] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0050] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0051] To further illustrate the technical solutions provided in the embodiments of this application, a detailed description is provided below in conjunction with the accompanying drawings and specific implementation methods. Although the embodiments of this application provide method operation instruction steps as shown in the following embodiments or drawings, the method may include more or fewer operation instruction steps based on conventional or non-creative effort. In steps where there is no logically necessary causal relationship, the execution order of these steps is not limited to the execution order provided in the embodiments of this application. In actual processing or system execution, the method may be executed sequentially or in parallel according to the method order shown in the embodiments or drawings.

[0052] It should be noted that the acquisition or use of data in the embodiments of this application requires the user's consent. The relevant data can only be obtained after the user's authorization, and the acquisition or use of the data complies with the laws and regulations of the relevant regions.

[0053] Please refer to Figure 1 , Figure 1 This illustration shows a flowchart of a geological data management method based on project path and semantic association, provided in an embodiment of this application. Figure 1 As shown, the method includes:

[0054] Step 101: Obtain multiple geological datasets to be submitted.

[0055] It should be noted that in the geological data submission and management scenario, the geological datasets to be submitted can come from different geological survey projects, different working methods, and different work areas. Specific types of geological datasets include, but are not limited to: borehole logging data (stored in tabular or database format), geophysical exploration data (such as seismic data, electromagnetic data, gravity data, and magnetic data, usually existing in gridded raster or measurement point sequence format), geochemical sampling data (existing in sampling point-element content table format), remote sensing interpretation data (existing in raster image or vector interpretation layer format), analytical testing data (existing in laboratory report or data table format), geological maps (existing in vector graphics file format), and 3D geological model data, etc. These data file formats include, but are not limited to, .shp, .gdb, .img, .db, .grd, .xlsx, and .docx. It should be understood that when the submitting unit submits a data file package, the system obtains the geological datasets to be submitted.

[0056] Step 102: Perform project path parsing for each geological dataset, obtain the hierarchical feature vector corresponding to the geological dataset based on the geological organization tree, and generate the dataset code corresponding to the geological dataset based on the hierarchical feature vector; the geological organization tree is a directory tree built based on "project-method-dataset".

[0057] It should be noted that the geological organization tree is a directory tree built based on the "project-method-dataset" structure. From top to bottom, the geological organization tree includes first-level project nodes, second-level project nodes, third-level project nodes, first-level working method nodes, second-level working method nodes, work area nodes, file grouping nodes, file nodes, and dataset nodes. It should be understood that the geological organization tree effectively reflects the actual management structure of geological survey work: a large-scale geological survey project (first-level project) encompasses several regional survey projects (second-level projects), each region and survey project contains several specific sub-projects (third-level projects), each sub-project employs one or more working methods, and each working method conducts data collection in a specific work area, forming several datasets.

[0058] It should be understood that the geological datasets to be submitted will contain corresponding project path identification information, such as the storage path of the data files corresponding to the geological datasets, accompanying metadata description files, or submission lists. After obtaining multiple geological datasets to be submitted, project path parsing can be performed on the geological datasets based on the hierarchical structure of the geological organization tree to obtain the hierarchical information corresponding to each geological dataset, including first-level projects, second-level projects, third-level projects, first-level work methods, second-level work methods, work areas, file groups, and dataset names. Then, hierarchical feature vectors are generated based on the hierarchical feature vectors, and dataset codes corresponding to the geological datasets are generated based on the hierarchical feature vectors. The hierarchical feature vectors are multi-dimensional feature vectors obtained by combining the hierarchical information corresponding to each geological dataset, and the dataset codes are encoding information obtained by encoding calculation based on the hierarchical feature vectors using a preset algorithm.

[0059] In one feasible embodiment, hierarchical information (e.g., name information) of each level node is extracted from the storage path of the geological dataset level by level using regular expression matching and path parsing algorithms. Then, the extracted hierarchical information is combined according to the structural order of the geological tissue tree to generate the hierarchical feature vector corresponding to the geological dataset. The hierarchical feature vectors are then concatenated and encoded to obtain the dataset code corresponding to the geological dataset.

[0060] For example, the hierarchical feature vector corresponding to geological dataset E can be represented as:

[0061]

[0062] in, This represents the hierarchical feature vector corresponding to the geological dataset E. For the primary project information corresponding to geological dataset E, For the secondary project information corresponding to geological dataset E, For the third-level project information corresponding to geological dataset E, The primary working method corresponding to geological dataset E, The secondary working method corresponding to geological dataset E, The work area number corresponding to geological dataset E, For the file grouping type corresponding to geological dataset E, This is the name of the dataset corresponding to geological dataset E.

[0063] The dataset encoding corresponding to geological dataset E can be represented as:

[0064]

[0065] in, is the dataset code corresponding to geological dataset E, is the hash function algorithm name, and is the sequential numbering function.

[0066] It should be understood that this application does not simply store file paths as strings as in existing technologies, but rather explicitly extracts the project hierarchy relationships contained in the file paths into computable feature vectors. Each component of this feature vector corresponds to a specific level node in the geological organization tree, possessing a clear geological management meaning. Based on this, the system further... The component representing the "project-method-workspace" combination (i.e.) to A SHA-256 hash operation is performed to generate a fixed-length encoding prefix. This prefix is ​​then concatenated with the dataset name and the sequential number corresponding to the file grouping type to generate the final dataset encoding. Through this hierarchical feature vector extraction and hash encoding mechanism, the directory tree, originally intended only for manual, step-by-step browsing, is transformed into an encoding space that can be efficiently indexed and retrieved by computer programs. Any geological dataset can be accessed through this... The prefix part can locate its project level, working method and spatial location in constant time without traversing the entire directory tree, thus providing a coding-level index foundation for subsequent multi-dimensional cross-project retrieval.

[0067] Step 103: Construct the metadata corresponding to the geological dataset based on the geological dataset, hierarchical feature vectors, and dataset encoding.

[0068] It's important to note that metadata is "data about data." The metadata for a geological dataset is descriptive data for that dataset, providing descriptive information about its content without requiring the dataset file to be opened. In geological data management scenarios, raw data files (such as .gdb, .shp, .img, etc.) are typically in binary or proprietary formats, with complex internal structures and large data volumes. If the raw file had to be opened and parsed every time an overview of a dataset was needed, it would result in unacceptable I / O overhead in large-scale data aggregation scenarios. Therefore, it is necessary to pre-build metadata as the index foundation for rapid retrieval and filtering.

[0069] In this embodiment, the metadata corresponding to the geological dataset is a three-level structure, including file-mirror level metadata, dataset-level metadata, and feature-level metadata. File-mirror level metadata records the physical attributes of the original dataset file, enabling rapid location and verification of its integrity without accessing the file content. Dataset-level metadata records the overall characteristics of the dataset, i.e., the overall feature description information, allowing for rapid filtering at the dataset level based on preset descriptive features without opening the dataset file. Feature description information records the descriptive information of at least some features in the dataset, further penetrating the metadata management granularity from the "dataset level" to the "feature level," supporting precise feature-level retrieval and tracing. A feature refers to the smallest independently identifiable and referenceable spatial data unit within the dataset, such as a single pixel or identifiable raster region in mountain song data. This can be set according to the data lake storage requirements, and this application does not impose specific limitations on it.

[0070] In a feasible embodiment, based on the geological dataset, hierarchical feature vectors, and dataset encoding, the physical attribute information, overall dataset feature description information, and feature description information corresponding to the geological dataset are determined; based on the physical attribute information, overall dataset description information, and feature description information, file mirror-level metadata, dataset-level metadata, and feature-level metadata in the metadata are constructed respectively.

[0071] It should be understood that the information types used to construct file-mirror-level metadata, dataset-level metadata, and feature-level metadata can be set according to the data lake's data computing capabilities or application requirements, and this application does not impose specific limitations on this.

[0072] For example, the physical attribute information of a geological dataset may include, but is not limited to, the unique identifier of the original geological dataset file, file format type, file size, digital fingerprint of the file, and file location pointer in the storage system. The overall characteristic description information of the dataset may include, but is not limited to, the spatial extent, scale, acquisition time, and method parameters corresponding to the geological dataset. Feature description information may include, but is not limited to, the geometric type of the feature, attribute pattern, and attribute value range.

[0073] Among them, in the dataset-level metadata, the spatial extent corresponding to the geological dataset can be the spatial envelope rectangle of the dataset, used to determine the corresponding geological dataset when searching the spatial extent; the scale can be the effective scale of the dataset, used to index the geological dataset when indexing at multiple scales and loading in a hierarchical manner; the acquisition time can be the data acquisition time window, used for time-series retrieval and geological age consistency judgment; the method parameters can be the technical parameters of the working method, used for method similarity calculation and data applicability evaluation.

[0074] Step 104: Based on the metadata, determine the format parser corresponding to the geological dataset, and use the format parser to perform element-by-element parsing on the geological dataset to obtain the standardized feature characteristics corresponding to each element.

[0075] It should be noted that, due to the diverse original file formats of geological datasets, in order to effectively parse geological datasets, it is necessary to first parse the geological datasets using a matching format parser.

[0076] Specifically, the format parser corresponding to the geological dataset is determined based on the file format type in the file mirror-level metadata.

[0077] Furthermore, the geological dataset is parsed element by element using a format parser to obtain the standardized feature characteristics corresponding to each element. This includes: parsing the geological dataset element by element using the format parser; performing spatial geometric representation standardization and attribute information standardization for each element to obtain the standardized spatial geometric representation and standardized attribute set corresponding to the element; after the element standardization is completed, a source tuple corresponding to the element is generated based on the unique identifier of the geological dataset, the location information of the element in the geological dataset, and the identifier of the standardization steps performed by the element; and the standardized feature characteristics corresponding to the element are generated based on the standardized spatial geometric representation, standardized attribute set, and source tuple corresponding to the element.

[0078] The standardization of spatial geometric representation includes, but is not limited to, converting vector data into geometric objects in the CGCS2000 coordinate system and raster data into a regular grid matrix with geocoding, to standardize the spatial geometric information of each feature in the geological dataset. Attribute information standardization involves mapping attribute field names from different sources to a unified field dictionary and standardizing their dimensions. Source tuples serve as a precise mapping between standardized features and geological dataset files. When anomalies occur in the data of the standard lacustrine layer, the specific record location in the geological dataset containing the feature can be directly located using source tuples, achieving accurate data correction instead of having to reprocess the entire file.

[0079] It should be understood that the standardized feature characteristics corresponding to multiple elements constitute the standardized feature characteristics corresponding to the geological dataset.

[0080] Step 105: Using the standardized feature characteristics corresponding to each element as nodes, construct a semantic association graph based on the working area and working method in the hierarchical feature vector.

[0081] In other words, a semantic association graph is constructed by combining the standardized feature nodes corresponding to each element with the degree of association between the work areas and work methods of each element.

[0082] In a feasible embodiment, for any two elements, based on the standardized element features corresponding to the elements, the spatial overlap, semantic similarity of working methods, and consistency of geological age between the two elements are determined respectively; the edge weights between the two elements are determined according to the spatial overlap, semantic similarity of working methods, and consistency of geological age; and a semantic association graph is constructed according to the standardized element features and edge weights corresponding to each element.

[0083] Specifically, the spatial overlap between elements can be determined by extracting the standardized spatial geometric expressions from the standardized feature characteristics of each element, and then determining the overlap based on the standardized spatial geometric expressions of the two elements.

[0084] Semantic similarity of working methods can be based on standardized feature characteristics of elements, and pass-through can be used to obtain the first-level working method corresponding to the geological dataset to which the element belongs. and secondary working methods Then, the shortest path distance is calculated on the geological method ontology tree:

[0085]

[0086] in, The semantic similarity of the working methods between element i and element j. The attenuation coefficient is... Working method for element i Working method with element j The shortest path distance on the ontology tree of geological methods.

[0087] Among them, the geological method ontology tree can be an ontology tree structure constructed based on the working methods used in multiple geological datasets.

[0088] It should also be noted that, in the embodiments of this application, and Level 1 working method and secondary working methods The composite result. The composite method can be splicing, i.e. = When calculating the shortest path distance, one can... Starting from the node, traverse upwards to the root node, recording the sequence of nodes along the path. Starting from a node, traverse upwards to the root node, record the sequence of path nodes, find the lowest common ancestor of two nodes in the geological method ontology tree, and obtain the shortest path distance. .

[0089] For example, the following expression can be used:

[0090]

[0091] in, Working method for element i Working method with element j The shortest path distance on the ontology tree of geological methods For nodes Depth to the root node, For nodes Depth to the root node, Lowest common node Depth to the root node.

[0092] Geological age consistency can be determined by extracting the collection time window from the standardized attribute set corresponding to the standardized feature of each element, then determining the geological age of each element based on the collection time window of each element, and finally determining the consistency of the geological age of the two elements based on the difference between the two geological ages.

[0093] Furthermore, the edge weight between two elements can be expressed as:

[0094]

[0095] in, Let i be the edge weight between element i and element j. The spatial overlap between element i and element j. The semantic similarity of the working methods between element i and element j. For the consistency of geological age between element i and element j, , and These are the weighting coefficients.

[0096] Step 106: Based on multiple geological datasets, standardized feature characteristics, and semantic association graphs, construct a three-layer geological data lake and establish a multi-scale saliency index for each feature; the three-layer geological data lake includes at least the original lake layer and the standard lake layer.

[0097] In other words, after obtaining the relevant materials for constructing a data lake from the geological dataset, a three-layer geological data lake is built. The three-layer geological data lake includes at least an original lake layer and a standard lake layer. The original lake layer stores the original files of the geological dataset, and the standard lake layer stores standardized feature data. The standard lake layer also stores a semantic association graph. For example, the semantic association graph can be stored in the standard lake layer using an association table, or the standardized feature data can be directly stored using the relationships within the semantic association graph; this application does not specifically limit this approach.

[0098] It should be understood that the original lake layer uses file paths to express "where the data is stored," thus solving the problem of data assets being "storable and retrievalable." The standard lake layer uses feature tables, edge tables, and index lists to express "what each feature is, what the relationships are between features, and which features are loaded first," thus solving the problem of data being "understood, correlated, and used efficiently."

[0099] The third-layer geological data lake can be constructed according to needs. In a preferred embodiment, the third-layer data lake can be an application lake layer, which records the source traceability information of each application data. Specifically, the application lake layer uses data lineage to express "which elements these application results originate from," solving the problems of data "reuse and traceability."

[0100] In one feasible embodiment, geological data inherently possesses multi-scale characteristics, ranging from regional 1:250,000 scale geological maps to local 1:10,000 scale detailed survey data, exhibiting a vast scale span. Furthermore, geological datasets and their constituent elements also correspond to different scales. Based on this, this application constructs a multi-scale saliency index for each element, based on the idea that the "degree of need" for each standardized element differs at different spatial scales, to measure the position of each standardized element package within a continuous spectrum from "macro" to "micro" levels. High-value data belongs to "fine-grained data" (which should be loaded at a micro scale). Data with low values ​​belongs to "overview data" (which should be loaded at a macro scale).

[0101] Specifically, a multi-scale saliency index is established for each element, including: for each element, obtaining the element's spatial geometric quantity, effective scale, and attribute information gain; and establishing a multi-scale saliency index corresponding to the element based on the spatial geometric quantity, effective scale, and attribute information gain.

[0102] For example, the following expression can be used:

[0103]

[0104] in, as elements The corresponding multi-scale feature index, The spatial geometric quantity corresponding to element i. as elements The corresponding effective scale, as elements The corresponding attribute gain information, This is the adjustment coefficient.

[0105] It should be noted that spatial geometric quantities Used to reflect the richness of detail in an element. Attribute information gain is a measure of the richness of information contained in the set of attributes of an element, reflecting the amount of information in the attribute dimension. For an element whose set of attributes contains multiple fields (such as lithology code, resistivity value, depth, sampling date, etc.), InfoGain measures how much "additional information" the combination of these attribute fields can provide.

[0106] It should be noted that although the operation of the method of the present invention is described in a specific order in the accompanying drawings, this does not require or imply that the operations must be performed in that specific order, or that all the operations shown must be performed in order to achieve the desired result.

[0107] Figure 2 A schematic diagram of the structure of a geological data management system based on project path and semantic association provided in an embodiment of this application is shown.

[0108] like Figure 2 As shown, the geological data management system 10 based on project path and semantic association includes:

[0109] Module 11 is used to acquire multiple geological datasets to be submitted.

[0110] The first parsing module 12 is used to perform project path parsing for each of the geological datasets, obtain the hierarchical feature vectors corresponding to the geological datasets based on the geological organization tree, and generate the dataset code corresponding to the geological datasets based on the hierarchical feature vectors; the geological organization tree is a directory tree constructed based on "project-method-dataset";

[0111] The first construction module 13 is used to construct the metadata corresponding to the geological dataset based on the geological dataset, the hierarchical feature vector, and the dataset encoding;

[0112] The second parsing module 14 is used to determine the format parser corresponding to the geological dataset based on the metadata, and use the format parser to perform element-by-element parsing on the geological dataset to obtain the standardized element features corresponding to each element.

[0113] The second construction module 15 is used to construct a semantic association graph based on the working area and working method in the hierarchical feature vector, using the standardized feature features corresponding to each element as nodes.

[0114] The third construction module 16 is used to construct a three-layer geological data lake based on multiple geological datasets, standardized feature characteristics, and semantic association graphs, and to establish a multi-scale saliency index for each feature; the three-layer geological data lake includes at least an original lake layer and a standard lake layer.

[0115] In some embodiments, the first construction module 13 is specifically used for:

[0116] Based on the geological dataset, the hierarchical feature vector, and the dataset encoding, determine the physical attribute information, overall feature description information, and element description information corresponding to the geological dataset;

[0117] Based on the physical attribute information, the overall feature description information of the dataset, and the feature description information, file mirror-level metadata, dataset-level metadata, and feature-level metadata are constructed respectively in the metadata.

[0118] In some embodiments, the second parsing module 14 is specifically used for:

[0119] Based on the file format type in the file mirror-level metadata of the aforementioned metadata, determine the format parser corresponding to the geological dataset.

[0120] In some embodiments, the second parsing module 14 is specifically used for:

[0121] The geological dataset is parsed element by element using the format parser. For each element, spatial geometric representation standardization and attribute information standardization are performed to obtain the standardized spatial geometric representation and standardized attribute set corresponding to the element.

[0122] After the element is standardized, a source tuple corresponding to the element is generated based on the unique identifier of the geological dataset, the location information of the element in the geological dataset, and the identifier of the standardization steps performed on the element.

[0123] Based on the standardized spatial geometric representation of the element, the standardized attribute set, and the source tuple, the standardized element feature corresponding to the element is generated.

[0124] In some embodiments, the second building module 15 is specifically used for:

[0125] For any two elements, based on the standardized element features corresponding to the elements, determine the spatial overlap, semantic similarity of working methods, and consistency of geological age between the two elements respectively;

[0126] The edge weights between the two elements are determined based on the spatial overlap, the semantic similarity of the working methods, and the consistency of the geological ages.

[0127] The semantic association graph is constructed based on the standardized element features and weights corresponding to each element.

[0128] In some embodiments, the third construction module 16 is specifically used for:

[0129] For each of the aforementioned elements, obtain the spatial geometric quantity, effective scale, and attribute information gain of the element;

[0130] Based on the spatial geometric quantity, the effective scale, and the attribute information gain, a multi-scale saliency index is established for the element.

[0131] It should be understood that the modules or modules recorded in the geological data management system 10 based on project paths and semantic associations are related to the reference modules. Figure 1 The steps in the described method correspond accordingly. Therefore, the operations and features described above for the method are also applicable to the geological data management system 10 based on project path and semantic association, and its included modules, and will not be repeated here. The geological data management system 10 based on project path and semantic association can be pre-implemented in the browser or other secure applications of an electronic device, or can be loaded into the browser or its secure applications of an electronic device through download or other means. The corresponding modules in the geological data management system 10 based on project path and semantic association can cooperate with the modules in the electronic device to implement the solution of the embodiments of this application.

[0132] The division of modules or units mentioned in the detailed description above is not mandatory. In fact, according to the embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0133] The following is for reference. Figure 3 , Figure 3 A schematic diagram of the structure of a computer system suitable for implementing the embodiments of this application is shown.

[0134] like Figure 3 As shown, the computer system 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 302 or programs loaded from storage section 308 into random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the system's operating instructions. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0135] The following components are connected to I / O interface 305: an input section 306 including a keyboard, mouse, etc.; an output section 307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, modem, etc. The communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0136] Specifically, according to embodiments of this application, the flowchart above refers to... Figure 2 The described process can be implemented as a computer software program. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program contains program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the functions defined in the system of this application.

[0137] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operational instructions of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two connected blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified functions or operational instructions, or using a combination of dedicated hardware and computer instructions.

[0139] The units or modules described in the embodiments of this application can be implemented in software or hardware. The described units or modules can also be housed in a processor; for example, a processor can be described as including an acquisition module, a first parsing module, a first construction module, a second parsing module, a second construction module, and a third construction module. The names of these units or modules do not necessarily limit the specific unit or module itself; for example, the acquisition module can also be described as "acquiring multiple geological datasets to be submitted".

[0140] In another aspect, this application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments, or may exist independently and not assembled into the electronic device. The aforementioned computer-readable storage medium stores one or more programs that, when used by one or more processors, execute the geological data management method based on project path and semantic association described in this application.

[0141] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A geological data management method based on project path and semantic association, characterized in that, include: Obtain multiple geological datasets to be submitted; For each geological dataset, project path parsing is performed, and the hierarchical feature vector corresponding to the geological dataset is obtained based on the geological organization tree. The dataset code corresponding to the geological dataset is generated based on the hierarchical feature vector. The geological organization tree is a directory tree constructed based on "project-method-dataset". Based on the geological dataset, the hierarchical feature vector, and the dataset encoding, construct the metadata corresponding to the geological dataset; Based on the metadata, determine the format parser corresponding to the geological dataset, and use the format parser to perform element-by-element parsing on the geological dataset to obtain the standardized element features corresponding to each element; Using the standardized feature characteristics corresponding to each element as nodes, a semantic association graph is constructed based on the working area and working method in the hierarchical feature vector; Based on multiple geological datasets, standardized feature characteristics, and semantic association graphs, a three-layer geological data lake is constructed, and a multi-scale saliency index is established for each feature; the three-layer geological data lake includes at least an original lake layer and a standard lake layer.

2. The geological data management method based on project path and semantic association according to claim 1, characterized in that, The construction of metadata corresponding to the geological dataset based on the geological dataset, the hierarchical feature vector, and the dataset encoding includes: Based on the geological dataset, the hierarchical feature vector, and the dataset encoding, determine the physical attribute information, overall feature description information, and element description information corresponding to the geological dataset; Based on the physical attribute information, the overall feature description information of the dataset, and the feature description information, file mirror-level metadata, dataset-level metadata, and feature-level metadata are constructed respectively in the metadata.

3. The geological data management method based on project path and semantic association according to claim 1, characterized in that, The step of determining the format parser corresponding to the geological dataset based on the metadata includes: Based on the file format type in the file mirror-level metadata of the aforementioned metadata, determine the format parser corresponding to the geological dataset.

4. The geological data management method based on project path and semantic association according to claim 1, characterized in that, The process of using the format parser to perform element-by-element parsing of the geological dataset to obtain the standardized feature characteristics corresponding to each element includes: The geological dataset is parsed element by element using the format parser. For each element, spatial geometric representation standardization and attribute information standardization are performed to obtain the standardized spatial geometric representation and standardized attribute set corresponding to the element. After the element is standardized, a source tuple corresponding to the element is generated based on the unique identifier of the geological dataset, the location information of the element in the geological dataset, and the identifier of the standardization steps performed on the element. Based on the standardized spatial geometric representation of the element, the standardized attribute set, and the source tuple, the standardized element feature corresponding to the element is generated.

5. The geological data management method based on project path and semantic association according to claim 1, characterized in that, The step of constructing a semantic association graph by using the standardized feature features corresponding to each element as nodes and based on the working area and working method in the hierarchical feature vector includes: For any two elements, based on the standardized element features corresponding to the elements, determine the spatial overlap, semantic similarity of working methods, and consistency of geological age between the two elements respectively; The edge weights between the two elements are determined based on the spatial overlap, the semantic similarity of the working methods, and the consistency of the geological ages. The semantic association graph is constructed based on the standardized element features and weights corresponding to each element.

6. The geological data management method based on project path and semantic association according to claim 1, characterized in that, The step of establishing a multi-scale saliency index for each of the aforementioned elements includes: For each of the aforementioned elements, obtain the spatial geometric quantity, effective scale, and attribute information gain of the element; Based on the spatial geometric quantity, the effective scale, and the attribute information gain, a multi-scale saliency index is established for the element.

7. A geological data management system based on project path and semantic association, characterized in that, include: The acquisition module is used to acquire multiple geological datasets to be submitted. The first parsing module is used to parse the project path for each of the geological datasets, obtain the hierarchical feature vectors corresponding to the geological datasets based on the geological organization tree, and generate the dataset code corresponding to the geological datasets based on the hierarchical feature vectors; the geological organization tree is a directory tree constructed based on "project-method-dataset"; The first construction module is used to construct the metadata corresponding to the geological dataset based on the geological dataset, the hierarchical feature vector, and the dataset encoding; The second parsing module is used to determine the format parser corresponding to the geological dataset based on the metadata, and to use the format parser to perform element-by-element parsing on the geological dataset to obtain the standardized element features corresponding to each element. The second construction module is used to construct a semantic association graph based on the working area and working method in the hierarchical feature vector, using the standardized feature features corresponding to each element as nodes. The third construction module is used to construct a three-layer geological data lake based on multiple geological datasets, standardized feature characteristics, and semantic association graphs, and to establish a multi-scale saliency index for each feature; the three-layer geological data lake includes at least an original lake layer and a standard lake layer.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the geological data management method based on project path and semantic association as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the geological data management method based on project path and semantic association as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the geological data management method based on project path and semantic association as described in any one of claims 1-6.