Natural resource multi-source heterogeneous data analysis method and system

By constructing multi-dimensional metadata mapping relationship and semantic alignment, it converts it into spatiotemporal grid data, the inefficiency and error problems in the analysis of multi-source heterogeneous data of natural resources are solved, and efficient and accurate data analysis is achieved.

CN120372249AActive Publication Date: 2025-07-25GUANGDONG JINGDI PLANNING TECH CO LTD

Patent Information

Application Number
CN202510485726.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-25
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Existing methods for multi-source heterogeneous data analysis of natural resources are inefficient when processing massive data, and due to semantic and structural differences, the analysis results are lagging and errors, making it difficult to meet the real-time and accuracy requirements of decision support.

Method used

By constructing a multi-dimensional metadata mapping relationship, a metadata knowledge graph is generated for semantic alignment, and the data is converted into spatiotemporal grid data, and analyzed in combination with the optimal processing mode.

Benefits of technology

It improves data processing speed, reduces preprocessing time, eliminates errors caused by semantic deviations and structural differences, and improves the accuracy and calculation efficiency of analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372249A_ABST
    Figure CN120372249A_ABST
Patent Text Reader

Abstract

The invention provides a natural resource multi-source heterogeneous data analysis method and system, and the method comprises the steps: carrying out the extraction and recognition of obtained natural resource multi-source heterogeneous data, and obtaining structural features and semantic features; constructing a multi-dimensional metadata mapping relationship based on the structural features and the semantic features, and performing association mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph; performing semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantic unified data; performing semantic information and structure information analysis on the semantic unified data, determining a data structure conversion strategy, and converting the multi-source heterogeneous data into target space-time grid data based on the data structure conversion strategy; and analyzing the type characteristics of the target space-time grid data to obtain an analysis result, and determining an optimal processing mode based on the analysis result. According to the invention, the calculation speed and the accuracy of an analysis result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular, to a method and system for analyzing multi-source heterogeneous data of natural resources. Background Art

[0002] At present, with the rapid development of big data and information technology, the field of natural resource management and research has been paying increasing attention to the analysis of multi-source heterogeneous data. The analysis of multi-source heterogeneous data can not only provide rich and comprehensive basis for natural resource management decisions. In addition, in-depth analysis of such data helps to explore the characteristics and laws of the spatio-temporal distribution of natural resources, clarify the internal relationships between different resources, and thus improve the resource utilization efficiency.

[0003] However, when dealing with actual situations, the current multi-source heterogeneous data analysis methods need to invest a large amount of time and resources in both the data preprocessing stage and the subsequent calculation process due to the extremely complex natural resource data sources, which are not only diverse in types, but also have huge differences in data structures and formats. As a result, the analysis results are lagged, making it difficult to provide strong support for decision-making in a timely manner. In addition, the data generated by different data sources have differences in semantics and structures, which makes it extremely easy to have deviations when integrating data for unified analysis, thus affecting the accuracy of the analysis results. Summary of the Invention

[0004] The present invention provides a method and system for analyzing multi-source heterogeneous data of natural resources, which can improve the calculation speed, effectively solve the problem of lagged analysis results of the existing methods, and avoid the analysis error problems caused by semantic deviations and data structure differences.

[0005] In a first aspect, the present invention provides a method for analyzing multi-source heterogeneous data of natural resources, including:

[0006] Extracting and identifying the obtained multi-source heterogeneous data of natural resources to obtain structural features and semantic features;

[0007] Based on the structural features and the semantic features, constructing a multi-dimensional metadata mapping relationship, and performing associative mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph;

[0008] Performing semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantically unified data;

[0009] Analyzing semantic information and structural information of the semantically unified data respectively to determine a data structure conversion strategy, and converting the multi-source heterogeneous data into spatio-temporal grid data based on the data structure conversion strategy;

[0010] Analyze the type characteristics of the target spatio-temporal grid data to obtain an analysis result, and determine an optimal processing mode based on the analysis result.

[0011] In a second aspect, the present invention further provides a natural resource multi-source heterogeneous data analysis system, which is applied to the natural resource multi-source heterogeneous data analysis method as described in the first aspect; the natural resource multi-source heterogeneous data analysis system includes:

[0012] An extraction and recognition module, configured to extract and recognize the obtained natural resource multi-source heterogeneous data to obtain structural features and semantic features;

[0013] A knowledge graph construction module, configured to construct a multi-dimensional metadata mapping relationship based on the structural features and the semantic features, and perform an association mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph;

[0014] A semantic alignment module, configured to perform semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantically unified data;

[0015] A data conversion module, configured to analyze semantic information and structural information of the semantically unified data respectively, determine a data structure conversion strategy, and convert the multi-source heterogeneous data into spatio-temporal grid data based on the data structure conversion strategy;

[0016] An analysis module, configured to analyze the type characteristics of the target spatio-temporal grid data to obtain an analysis result, and determine an optimal processing mode based on the analysis result.

[0017] In a third aspect, the present invention further provides an electronic device, including: a memory for storing a computer software program; a processor for reading and executing the computer software program, thereby implementing the natural resource multi-source heterogeneous data analysis method as described in any one of the above.

[0018] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, in which a computer software program is stored, and when the computer software program is executed by a processor, the natural resource multi-source heterogeneous data analysis method as described in any one of the above is implemented.

[0019] In a fifth aspect, the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the natural resource multi-source heterogeneous data analysis method as described in any one of the above is implemented.

[0020] The natural resource multi-source heterogeneous data analysis method provided by the embodiments of the present invention constructs a multi-dimensional data mapping relationship through the structural features and semantic features in the multi-source heterogeneous data, integrates the scattered data source information together, and associates different data in the multi-source heterogeneous data according to the multi-dimensional data mapping relationship to form a unified and dynamic metadata knowledge graph, ensuring that the data from different data sources can accurately correspond in time and space, reducing the time and workload of data preprocessing, improving the calculation speed, and effectively solving the problem of lagging analysis results in the existing methods; further, semantic alignment is performed on the multi-source heterogeneous data according to the metadata knowledge graph to determine the semantically unified data with a unified semantic level, making the data form an organic whole at the semantic level, and cooperating with the dynamically generated data structure conversion strategy to convert data with different formats and structures into standardized spatio-temporal grid data, so that the data also forms a unity in structure, thereby avoiding the analysis error problem caused by semantic deviation and data structure differences and improving the accuracy of the analysis results; in addition, type feature analysis is performed on the spatio-temporal grid data with unified structure standardization, and the optimal data processing mode is automatically selected based on the analysis results, giving full play to the advantages of different processing modes, improving the calculation efficiency and further improving the accuracy at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic flowchart of the natural resource multi-source heterogeneous data analysis method provided by the embodiments of the present invention;

[0022] Figure 2 is a schematic structural diagram of the natural resource multi-source heterogeneous data analysis system provided by the embodiments of the present invention;

[0023] Figure 3 is an embodiment diagram of the electronic device provided by the embodiments of the present invention;

[0024] Figure 4 is an embodiment diagram of the computer-readable storage medium provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0026] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.

[0027] In the description of the present invention, the term "such as" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "such as" in the present invention is not necessarily to be construed as more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the present invention. In the following description, details are set forth for purposes of explanation. It should be understood that those skilled in the art can recognize that the present invention can be implemented without the use of these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0028] Refer to Figure 1 , Figure 1 which is a schematic flow chart of the method for analyzing multi-source heterogeneous data of natural resources provided by the present invention. In the embodiments of the present invention, the execution subject of the method for analyzing multi-source heterogeneous data of natural resources is a data analysis system. Therefore, the method for analyzing multi-source heterogeneous data of natural resources includes:

[0029] Step 10: Extract and identify the obtained multi-source heterogeneous data of natural resources to obtain structural features and semantic features.

[0030] Optionally, the data analysis system extracts and identifies data from natural resource data of different sources and formats, such as satellite remote sensing image data (e.g., in.tif format), vector data of geographic information systems (GIS) (e.g., in.shp format), and text data of meteorological monitoring stations, etc. For image data, the OpenCV library in Python can be used to read image files, and its structural features are extracted by analyzing the pixel arrangement rules, band information, etc. of the image, such as the resolution of the image. For semantic feature extraction, a convolutional neural network (CNN) model in deep learning can be used to train the image to identify natural resource types in the image, such as forests, water areas, etc., to obtain semantic features. For text data, natural language processing (NLP) techniques are used, such as based on word vector models (e.g., Word2Vec), to map the words in the text to a low-dimensional vector space, and semantic features are extracted by analyzing the vector relationships between words, that is, to determine the natural resource-related concepts represented by the data, such as land use types, mineral resource types, meteorological elements, etc., and the relationships between these concepts. For structured data such as database tables, structural features such as field names, types, and constraints are obtained by querying database metadata, and semantic features are determined based on data dictionaries and business rules; for semi-structured data (e.g., XML, JSON), tags and key-value pairs are parsed to obtain structural features, and semantic features are determined according to the content and context. By comprehensively and accurately extracting the structural and semantic features of the data, a solid foundation is provided for subsequent metadata mapping and semantic alignment, which can effectively reduce information loss and misunderstanding in the data processing process and improve the reliability of the analysis results. In addition, presenting different types of data in a unified feature form also facilitates the processing and analysis of the data in subsequent steps, improving the pertinence and accuracy of data processing.

[0031] Step 20: Based on the structural features and semantic features, construct a multi-dimensional metadata mapping relationship, and perform an association mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph.

[0032] Optionally, the data analysis system constructs a multi-dimensional metadata mapping relationship according to the association relationship between the structural features and semantic features extracted in step 10, and in different dimensions such as space and semantics of the natural resource data. For data items representing the same or related concepts in different data sources, they are associated through the mapping relationship, and finally these mapping relationships are integrated into a knowledge graph to form a unified and dynamic metadata knowledge graph, providing comprehensive semantic and structural information for subsequent data processing, as specifically described in steps 201 - 205.

[0033] Furthermore, through the construction of multi-dimensional metadata mapping relationships, the data analysis system can comprehensively integrate multi-source heterogeneous data, establishing an organic connection between the originally isolated data. Moreover, the complex relationships between the data are intuitively displayed in the form of a knowledge graph, facilitating subsequent data query and analysis.

[0034] Step 30: Semantically align the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantically unified data.

[0035] Optionally, based on the obtained metadata knowledge graph, the data analysis system can use alignment techniques (such as ontology alignment techniques, etc.) to semantically align the same concepts and the relationships between concepts in different data sources to obtain semantically unified data, as specifically described in Steps 301 - 305. Through semantic alignment, the multi-source data is made consistent at the semantic level, eliminating the understanding and processing obstacles caused by semantic expression differences and improving the usability of the data.

[0036] Furthermore, in one embodiment, the "cultivated land area" in the land resource data source is conceptually matched with the "scale of land used for farming" in the agricultural statistics data source.

[0037] Step 40: Analyze the semantic information and structural information of the semantically unified data respectively, determine the data structure conversion strategy, and convert the multi-source heterogeneous data into spatio-temporal grid data based on the data structure conversion strategy.

[0038] Optionally, the data analysis system analyzes the semantic information and structural information of the received semantically unified data again, combines the information in the spatio-temporal dimension, and determines the data structure conversion strategy for different data sources, as specifically described in Steps 4011 - 4014. Then, the data analysis system converts the multi-source heterogeneous data (i.e., structured data, unstructured data, and semi-structured data) into standardized spatio-temporal grid data according to the determined data structure conversion strategy, as specifically described in Steps 4021 - 4024.

[0039] Furthermore, the standardized spatio-temporal grid is a unified data representation method that divides time and space into regular grid cells, and each grid cell contains specific attribute values. For example, for satellite remote sensing image data, it is gridded according to a certain spatio-temporal resolution, and the information of each pixel is mapped to the corresponding grid cell; for meteorological monitoring data, the monitoring values at different locations and times are interpolated into the grid cells. Through this mapping, data in different formats and structures is unified into standardized spatio-temporal grid data for subsequent calculation and analysis.

[0040] Step 50: Analyze the type characteristics of the target spatio-temporal grid data to obtain an analysis result, and determine the optimal processing mode based on the analysis result.

[0041] Optionally, the data analysis system performs a comprehensive analysis of data type characteristics based on spatio-temporal grid data. On the one hand, it analyzes the overall type to clarify the characteristics of the data in the spatio-temporal dimension. On the other hand, it deeply analyzes the internal data types and automatically selects the optimal processing mode according to the analysis results, as specifically described in steps 501 - 504.

[0042] Furthermore, in one embodiment, for data with complex network relationships, such as the mutual relationship data between species in an ecosystem, the graph calculation mode is selected to analyze the connection and propagation rules between nodes through graph algorithms; for high-dimensional data, such as multi-band satellite remote sensing image data, the tensor decomposition mode is selected to decompose the high-dimensional data into low-dimensional tensors to reduce the data complexity; for the case where there are missing values in spatio-temporal data, the spatio-temporal interpolation mode is selected to predict the values of missing points based on the values of known data points. By automatically selecting the optimal calculation mode, the advantages of different calculation modes are fully utilized, the calculation efficiency and accuracy are improved, and at the same time, the flexibility and adaptability of the system are enhanced, enabling it to better handle complex and diverse multi-source heterogeneous data of natural resources.

[0043] In the embodiment of the present invention, a multi-dimensional data mapping relationship is constructed through the structural features and semantic features in multi-source heterogeneous data, integrating the scattered data source information together, and associating different data in the multi-source heterogeneous data according to the multi-dimensional data mapping relationship to form a unified and dynamic metadata knowledge graph, ensuring that the data from different data sources can accurately correspond in space and time, reducing the time and workload of data preprocessing, improving the calculation speed, and effectively solving the problem of lag in the analysis results of existing methods; further, semantic alignment of the multi-source heterogeneous data is performed according to the metadata knowledge graph to determine the semantically unified data with a unified semantic level, making the data form an organic whole at the semantic level, and cooperating with the dynamically generated data structure conversion strategy to convert data with different formats and structures into standardized spatio-temporal grid data, so that the data also forms a unity in structure, thereby avoiding the analysis error problem caused by semantic deviation and data structure differences and improving the accuracy of the analysis results; in addition, type feature analysis is performed on the spatio-temporal grid data with unified structure standardization, and the optimal data processing mode is automatically selected based on the analysis results, fully utilizing the advantages of different processing modes, improving the calculation efficiency and further improving the accuracy.

[0044] In one embodiment, the descriptions of steps 201 - 205 are as follows:

[0045] In step 201, a structural feature topology graph is constructed with structural features as nodes and the dependency relationships and hierarchical relationships between structural features as edges.

[0046] Optionally, after the data analysis system obtains the structural features of multi-source heterogeneous data, it sorts out the structural features and analyzes the dependency relationships and hierarchical relationships between the structural features. Among them, the dependency relationship can be a causal relationship or an association relationship. For example, in the association relationship, the structural features include fields such as land type, area, and ownership. Among them, the calculation of the area may depend on the boundary information of the land type, which constitutes a dependency relationship. For example, in the causal relationship, the change in air pressure in meteorological data may cause a change in wind direction. The hierarchical relationship includes the inclusion relationship. For example, a large geographical area contains multiple small sub-areas; the general land use table contains multiple sub-area land use tables. According to the dependency relationships and hierarchical relationships between the structural features, a structural feature topology graph is constructed using the method of graph theory. Each structural feature serves as a node, and the relationship serves as an edge. The structural feature topology graph intuitively shows the organization and association method of the data structure.

[0047] Further, in one embodiment, taking a forest resource monitoring system as an example, the data sources include a tree species database, a forest area statistics document, and a pest record form. There is a dependency relationship between the tree species field in the tree species database and the covered area of different tree species in the forest area statistics document, that is, the area of the corresponding tree species can be counted through the tree species information. Taking structural features such as tree species and the covered area of different tree species as nodes, and this dependency relationship as an edge, a structural feature topology graph is constructed.

[0048] Step 202, classify the semantic features and hierarchically arrange the classification results to obtain a semantic feature hierarchy tree.

[0049] Optionally, the data analysis system uses natural language processing (NLP) technology to classify the extracted semantic features. A pre-trained word vector model (such as BERT) can be used to represent the semantic features as vectors, and then similar semantic features are clustered into one category through a clustering algorithm (such as K-means clustering). For example, forest, grassland, wetland, etc. can be classified into the ecological environment category; coal mine, iron mine, petroleum, etc. can be classified into the mineral resource category. After classification, hierarchical arrangement is carried out according to the degree of abstraction of the semantics. The category with a higher degree of abstraction is located at the top layer of the hierarchy tree, and the specific semantic features are located at the bottom layer. The hierarchical tree can be represented using the data structure of a tree, where each node contains a category or semantic feature, and there is a hierarchical relationship between the child nodes and the parent node.

[0050] Step 203, starting from the root node of the structural feature topology graph, associate the root node with the top layer node of the semantic feature hierarchy tree to obtain an initial mapping relationship framework.

[0051] Optionally, after determining the structural feature topology map and the semantic feature hierarchical tree, the data analysis system first determines the root node of the structural feature topology map, which is usually the top-level, most comprehensive structural feature, and then determines the top-level node of the semantic feature hierarchical tree, and associates the root node with the top-level node to obtain an initial mapping relationship framework. For example, in a comprehensive natural resource management data system, the root node of the structural feature topology map may be a natural resource basic information table, and the top-level node of the semantic feature hierarchical tree may be a resource type, resource status, etc. The natural resource basic information table is associated with the top-level nodes such as resource type and resource status based on the corresponding relationship of business logic or domain knowledge, thereby forming an initial mapping relationship framework.

[0052] Step 204 , based on the initial mapping relationship framework, association expansion is performed along the edges of the structural feature topology graph and the branches of the semantic feature hierarchy tree to determine a multi-dimensional metadata mapping relationship.

[0053] Optionally, the data analysis system is based on the initial mapping relationship framework, starting from the root node of the structural feature topology map, traversing the topology map along the direction of the edge, and starting from the top node of the semantic feature hierarchy tree, expanding downward along the branch. When traversing to a certain node of the structural feature topology map, according to its relationship with the root node and the established mapping, the corresponding semantic node is found in the semantic feature hierarchy tree for association, and a new mapping relationship is established. For example, the land use sub-table node in the structural feature topology map is connected to the root node of the land use general table through an edge. Based on the association between the land use general table and the land use semantic category in the initial mapping, the specific semantic branch corresponding to the content of the land use sub-table (such as cultivated land use details, etc.) is found under the land use semantic category for association, and this process is repeated until all nodes of the structural feature topology map and all branches of the semantic feature hierarchy tree are traversed to obtain a multi-dimensional metadata mapping relationship.

[0054] Step 205, perform association mapping on multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph.

[0055] Optionally, the data analysis system performs association mapping on the multi-source heterogeneous data according to the obtained multi-dimensional metadata mapping relationship, and more accurately determines the association relationship between the data based on the context information of the multi-source heterogeneous data, thereby constructing a metadata knowledge graph with better context-awareness capabilities, as described in steps 2051 to 2054.

[0056] The embodiments of the present invention can organically combine the structural features and semantic features in multi-source heterogeneous data, and mine the deep-level associations between data, thereby improving the integration and availability of data and providing strong support for the analysis and application of natural resource data.

[0057] In one embodiment, the descriptions of steps 2051 - 2054 are as follows:

[0058] Step 2051, extract context information from multi-source heterogeneous data to obtain context data.

[0059] Optionally, the data analysis system uses data mining and natural language processing technologies to extract context information from multi-source heterogeneous data of different sources and formats. For structured data (such as tabular data in a database), analyze information such as the associations between fields, the timestamps and geographical locations of the data. For example, for meteorological data, in addition to basic data such as temperature and humidity, the collection time, collection location, etc. can also be extracted as context information. For semi-structured data (such as XML documents), the hierarchical structure of the tags, the overall theme of the document, etc. constitute context information. For unstructured data (such as geological exploration reports), through methods such as topic models and named entity recognition in natural language processing technologies, these context information are extracted from the text.

[0060] Step 2052, perform similarity processing on each associated mapping relationship in the metadata mapping relationship based on the context data and a preset similarity function to obtain context similarity.

[0061] Optionally, for each associated mapping relationship in the metadata mapping relationship, the data analysis system extracts the time information contained in the two associated data, such as the collection time and update time of the acquired data, and converts it into the form of a timestamp. Calculate the difference between the two timestamps, and convert the difference into years to obtain the time distance value S time , with a value in the range of [0 - 1]. The smaller the difference, the closer the two data are in the time dimension, and the stronger the correlation in terms of time may be. Then, extract the spatial information in the associated data, that is, geographical coordinates (such as longitude and latitude). And by means of the spherical distance formula (such as calculating the shortest distance between two points on the sphere using parameters such as the radius of the earth), calculate the interval between the two data in the actual space to obtain S space . With a value in the range of [0 - 1]. The smaller the spatial interval, the closer the connection between the two in the spatial dimension. Finally, evaluate the sources of the associated data, considering factors such as the nature of the data source (such as released by government departments, research by scientific research institutions, enterprise statistics, etc.), authority (whether peer-reviewed, whether widely cited, etc.). Set certain evaluation criteria. If the similarity degree of the source nature and authority is high, then S source has a high value, with a value in the range of [0 - 1]. Finally, obtain the preset similarity function where S xsdIt is expressed as a context similarity value, with a value range of [0-1]. α, β, and γ are all expressed as weight coefficients, and α + β + γ = 1, which are set according to specific business requirements and the degree of emphasis on different dimensions.

[0062] Step 2053: Based on the context similarity, screen and sort the association mapping relationships in the multi-dimensional data mapping relationships to obtain candidate association mapping relationships.

[0063] Optionally, the data analysis system presets a similarity threshold according to the obtained context similarity value, and screens the association mapping relationships with a context similarity higher than the similarity threshold, indicating that the screened association mapping relationships belong to relatively reliable and close associations in the current context environment. Then, sort the screened association mapping relationships from high to low according to the context similarity to obtain candidate association mapping relationships. By means of screening and sorting, low-quality and irrelevant association mapping relationships can be removed, reducing noise and redundancy in the process of constructing the knowledge graph. And the setting of the similarity threshold is set according to the actual situation, and manual judgment can also be made on the associations that are screened out to adjust the similarity threshold.

[0064] Step 2054: Based on the candidate association mapping relationships, perform association mapping on the multi-source heterogeneous data to obtain an initial knowledge graph, and based on the context data, annotate the attributes of the nodes and edges of the initial knowledge graph to obtain a metadata knowledge graph.

[0065] Optionally, the data analysis system associates each data element in the multi-source heterogeneous data according to the screened and sorted candidate association mapping relationships to construct an initial knowledge graph. In the initial knowledge graph, data elements (such as database fields, key concepts in text, etc.) serve as nodes, and the association mapping relationships serve as edges. Then, use the context data to annotate the attributes of the nodes and edges in the initial knowledge graph. For nodes, annotate attributes such as the data source, data type, and context theme to which they belong. For edges, annotate attributes such as the basis of the association, the context similarity value, and the description of the closeness of the association to obtain a metadata knowledge graph.

[0066] Furthermore, in one embodiment, taking the processing of multi-source heterogeneous data regarding natural resources in a certain region as an example, data elements such as forests, minerals, and rivers are associated according to the candidate association mapping relationships to construct an initial knowledge graph; then, according to the context data (such as the forest fire occurrence time, the mineral exploitation history, etc.), the attributes of the nodes and edges in the knowledge graph are annotated. For example, it is annotated on the XX forest node that a fire occurred on XX year XX month XX day, and it is annotated on the XX forest - ecological protection edge that it was affected to a certain extent due to the fire.

[0067] In an embodiment of the present invention, by introducing context information and similarity processing, meaningful association mapping relationships can be accurately screened out from complex multi-source heterogeneous data, and a more accurate, rich, and practical metadata knowledge graph can be constructed.

[0068] In one embodiment, the descriptions of steps 301 - 305 are as follows:

[0069] Step 301: For each data source in the multi-source heterogeneous data, node extraction is performed on the metadata knowledge graph based on the features involved in each data source to obtain feature nodes.

[0070] Optionally, after the data analysis system receives the multi-source heterogeneous data and determines the metadata knowledge graph, due to the diverse sources of the multi-source heterogeneous data, for each data source, its data is first analyzed to analyze the concept information of natural resource features it contains. For example, in the forestry resource data source, concept information reflecting forestry characteristics such as forest coverage rate and the number of tree species; in addition, the data analysis system pre-constructs a knowledge base of concepts related to natural resources. After extracting the concept information of natural resource features, its standard definition, attribute information, etc. are obtained from the pre-constructed knowledge base of concepts related to natural resources. For example, the definition of forest coverage rate is the percentage of forest area in the total land area, and its attributes include statistical time, statistical area, etc. Then, starting from each data source, based on the concept information of natural resource features it involves (i.e., the features involved in the data source), corresponding nodes are matched in the metadata knowledge graph as feature nodes. Ensure that the extracted feature nodes have accurate and complete semantic information. Combining the knowledge base information can avoid concept misunderstandings and ensure the accuracy and consistency of node semantics, laying a foundation for constructing reliable semantically unified data.

[0071] Step 302: For the relationships between the feature nodes in each data source, a feature topology structure is constructed.

[0072] Optionally, after the data analysis system obtains the feature nodes of each data source, it analyzes the relationships between the feature nodes, including causal relationships, hierarchical relationships, etc. For example, in the meteorological monitoring data source, there may be a certain correlation between temperature and humidity. An increase in temperature may lead to a decrease in humidity, and wind speed and wind direction are in a mutually related parallel relationship; in the geographic information database, there may be an inclusion relationship between land use types and landforms. For example, mountains may correspond to land use types such as forest land, and the forest ecosystem includes hierarchical relationships of sub-concepts such as tree species and soil quality. Then, with the feature nodes as vertices and their relationships as edges, a directed or undirected graph is constructed to represent the feature (concept) topology structure. It can visually display the relationships between the feature nodes within each data source, helping to deeply understand the internal logic of the data source.

[0073] Step 303: Based on a preset structural similarity function, match the feature nodes between the feature topological structures of two different data sources to obtain a matching similarity.

[0074] Optionally, the preset structural similarity function of the data analysis system is used to measure the similarity degree of the feature topological structures of different data sources. Considering the weight relationship between nodes and edges comprehensively, match the feature nodes between the feature topological structures of two different data sources.

[0075] Furthermore, for two feature topological structures G1 = (V1, E1, W1) and G2 = (V2, E2, W2) (where V represents the node set, E represents the edge set, and W represents the weight set), when calculating the structural similarity, first match the nodes. For node v i1 ∈ V1 and v i2 ∈ V2, if they represent similar concepts (judged by the previous semantic similarity), they are regarded as matching nodes, denoted as δ i1i2 = 1, otherwise δ i1i2 = 0. The matching of edges is similar. If the nodes connected by edge e j1 ∈ E1 and e j2 ∈ E2 are matched and the weight difference of the edges is within a certain range (set as Δw), then the edges are considered to be matched. Then the matching similarity formula is:

[0076] where n1 and n2 respectively represent the number of nodes of G1 and G2; ω ij is the weight of the relationship between nodes i and j (if the nodes are not connected, ω ij = 0); δ ij represents a binary variable, used to indicate whether nodes i and j are matched. If nodes i and j represent similar concepts, then δ ij = 1, otherwise δ ij = 0; ω ij1 is the weight of the edge connecting nodes i and j in G1 (if nodes i and j are not connected, the value has no practical meaning); ω ij2 is the weight of the edge connecting the corresponding nodes (the nodes matching nodes i and j in G1) in G2 (if the corresponding nodes are not connected, this value has no practical meaning), both of which are used to measure the influence of the weight difference of the edges on the structural similarity.

[0077] Furthermore, in an embodiment, for two data sources regarding natural resources, their feature topological structures are G1 and G2 respectively. In G1, there are forest area, tree species nodes and an edge with a weight of 0.6 between them. In G2, there are forest land scope (semantically similar to forest area), tree species type (semantically similar to tree species) nodes and an edge with a weight of 0.5 between them. Through calculation using the structural similarity formula, ω ij corresponding edge weights, δij = 1 (node matching), |ω ij1 - ω ij2 | = |0.6 - 0.5| = 0.1. Substitute it into the formula to calculate the similarity contribution value of these two feature topologies in this part. Then, considering the situations of other nodes and edges, the overall matching similarity is obtained.

[0078] Step 304: If the matching similarity is greater than or equal to the preset threshold, map the two feature nodes corresponding to the matching similarity to obtain the node mapping result.

[0079] Optionally, the data analysis system determines the preset threshold according to the criteria determined by the actual application requirements and data characteristics, which is used to judge whether the feature nodes can be mapped. When the matching similarity of the feature topologies of two data sources reaches or exceeds this threshold, it indicates that the corresponding feature nodes have a high consistency in semantics and structure, and a mapping relationship is established. For example, in the water resource monitoring data sources in different regions, the matching similarity of the water quality nodes of the drinking water sources is higher than the threshold, indicating that they are semantically similar and can be mapped to the same semantic concept.

[0080] Step 305: Integrate the multi-source heterogeneous data based on the node mapping result to obtain semantically unified data.

[0081] Optionally, the data analysis system integrates the multi-source heterogeneous data according to the obtained node mapping result. First, for the data of different data sources mapped to the same concept, the data formats are unified. For example, the unit of forest area in one data source is hectares, while in another data source it is square kilometers, and they are unified to hectares. Then, the data values are merged. If there are measurement values of the same concept from multiple data sources, operations such as weighted averaging can be performed according to factors such as data reliability and timeliness. Next, the data content is merged. For example, the precipitation monitoring data from different regions are merged according to the time series. At the same time, the semantic descriptions of the data are also unified and standardized. For example, different names for forests in different data sources are unified to forest. After these processes, the differences in multi-source heterogeneous data are eliminated to obtain semantically unified data.

[0082] Through the systematic processes of node extraction, topology construction, similarity calculation, node mapping, and data integration in the embodiments of the present invention, the metadata knowledge graph and the internal structure information of the data sources are fully utilized, effectively solving the problem of semantic inconsistency in multi-source heterogeneous data. And a unique matching similarity function is designed to enable accurate identification of feature nodes with similar semantics and structures in different data sources, providing a reliable basis for data integration.

[0083] In one embodiment, the descriptions of steps 4011 - 4014 are as follows:

[0084] Step 4011: Analyze the semantic elements of the semantically unified data to obtain the semantic element analysis result, and construct a semantic association network based on the semantic element analysis result.

[0085] Optionally, the data analysis system deeply analyzes the semantic elements of the semantically unified data, disassembling them into specific semantic elements. For example, in natural resource data, forests, tree species, and water resource reserves are all semantic elements. Mainly through natural language processing techniques such as part-of-speech tagging and named entity recognition, the type and attributes of each semantic element are determined. Then, based on the logical relationship between the semantic elements, the direct semantic association between the elements is judged. For example, between forests and tree species, since a forest is composed of multiple tree species, there is a direct semantic association between forests and tree species. Connect the elements with direct semantic associations by edges to construct a semantic association network. In the semantic association network, nodes represent semantic elements, and edges represent the semantic association relationships between the elements.

[0086] Step 4012: Analyze the structural feature information of the semantically unified data to obtain the structural feature categories, and analyze the semantically unified data in terms of time and space dimensions to determine the spatio-temporal dimension information.

[0087] Optionally, the data analysis system analyzes the result features of the semantically unified data and analyzes its structural features according to the data presentation form and storage method. During the analysis, if the data describes the natural resource situation in the form of a text paragraph, it belongs to the text structure; if the data is a series of numerical records, such as the numerical values of natural resource indicators monitored, it belongs to the numerical structure; if the data is recorded in chronological order, such as the monthly monitoring data of a river flow, it belongs to the time series structure. At the same time, the data analysis system also analyzes the semantically unified data in terms of time and space. In the time dimension, information such as the time point and time period corresponding to the data is determined, such as the monitoring time in the forest resource monitoring data; in the space dimension, the geographical location and regional scope involved in the data are determined through geographical coordinates, place names, etc., and finally the spatio-temporal dimension information is obtained.

[0088] Step 4013: Map the nodes in the semantic association network to the structural feature categories to obtain the association mapping result, and determine the initial structure conversion strategy based on the association mapping result.

[0089] Optionally, the data analysis system maps each node (i.e., semantic element) in the semantic association network to the structural feature categories. For example, for the semantic node of the forest area value, since it involves specific numerical values, it is mapped to the numerical structure category; through this mapping method, traverse the nodes in the semantic association network to establish the connection between semantics and structure, that is, obtain the association mapping result.

[0090] Further, after the data analysis system determines the associated mapping result, it formulates a corresponding initial structure conversion strategy. For example, if the semantic node belongs to the text structure and is related to the spatial position, the strategy is to extract the description of the spatial position in the text and convert it into grid coordinates. If it belongs to the numerical structure and is related to time, the data is sorted according to the time sequence. It should be noted that the initial structure conversion strategy provides preliminary guidance for subsequent data conversion, can specifically process data with different semantics and structures, and improves the efficiency and accuracy of data processing.

[0091] Step 4014, update the initial structure conversion strategy based on the spatio-temporal dimension information to obtain the data structure conversion strategy.

[0092] Optionally, after the data analysis system determines the initial structure conversion strategy, it combines the spatio-temporal dimension information determined in step 4012 to optimize and update the initial structure conversion strategy. On the one hand, considering the time factor, for spatial data that changes over time, such as urban land use change data at different time points, the conversion strategy can be refined to sequentially convert the spatial data at different times into grid data on the corresponding time slices according to the time sequence. On the other hand, considering the spatial factor, according to the geographical location and regional scope involved in the data, the conversion strategy is adjusted to make it more in line with the actual spatial distribution. Therefore, by adding spatio-temporal dimension information, the initial structure conversion strategy can be made more in line with the requirements of spatio-temporal grid data, and it is ensured that the converted data can accurately reflect the spatio-temporal characteristics of the data.

[0093] Through semantic element analysis, structural feature analysis, and spatio-temporal dimension analysis of the semantically unified data in the embodiments of the present invention, a semantic association network is gradually constructed, the structural feature category and spatio-temporal dimension information are determined, and the data structure conversion strategy is formulated and updated through semantic-structure association mapping. The semantic, structural, and spatio-temporal characteristics of the data are comprehensively considered, the data connotation can be deeply understood, the data structure conversion strategy is made more scientific and reasonable, and it lays a foundation for subsequent processing such as converting multi-source heterogeneous data into spatio-temporal grid data, which helps to improve the analysis and utilization efficiency of natural resource data.

[0094] In one embodiment, the multi-source heterogeneous data includes structured data, semi-structured data, and unstructured data. The descriptions of steps 4021 - 4025 are as follows:

[0095] Step 4021, preprocess the structural features and semantic features of the structured data, and perform data conversion on the preprocessed data based on the data structure conversion strategy to obtain the first spatio-temporal grid data.

[0096] Optionally, the data analysis system performs preprocessing based on the structural and semantic features of the structured data extracted in step 10. For example, for a database table recording land resource information, based on the semantic features of the fields, it determines which fields are related to spatial locations (such as plot coordinate fields), which are related to resource attributes (such as land type fields), and if there are data missing or errors, it fills or corrects them appropriately according to the semantic features and business rules of the data. After the preprocessing is completed, the data analysis system maps the spatio-temporal related information in the structured data to the spatio-temporal grid according to the determined data structure conversion strategy. For example, if there is a clear geographical coordinate field (such as longitude and latitude) in the table, according to the grid division rules, the coordinates are mapped to the corresponding spatial grid cells. For time information (such as data recording time), it is mapped to the grid in the time dimension according to the time granularity (such as divided by day or month). Other attributes (such as land type, area, etc.) between the spatio-temporal information in the structured data are filled into the corresponding spatio-temporal grid cells. Specifically, the storage method and format of the attributes can be determined according to the semantic features of the data to ensure that their meanings can be accurately expressed in the spatio-temporal grid data. For example, the land type is stored in text form and the area is stored in numerical form in the attribute fields of the corresponding grid cells, and finally the first spatio-temporal grid data of the structured data conversion is obtained.

[0097] Step 4022, analyze and extract the structural and semantic features of the semi-structured data, and perform data conversion on the parsed data based on the data structure conversion strategy to obtain the second spatio-temporal grid data.

[0098] Optionally, the data analysis system uses parsing tools to perform in-depth parsing and processing according to the structural features (such as the tag hierarchy of XML and the key-value pair structure of JSON) and semantic features (such as the meanings represented by the tags or key-value pairs) of the semi-structured data extracted in step 10. For example, for a JSON file describing forest resources, according to the structural features of the key-value pairs, it determines the organization form of the data, and based on the semantic features, it clarifies the forest resource information represented by each key-value pair (such as tree species, tree height, etc.). After that, it extracts the spatio-temporal related information from the parsed data, such as the geographical location of the forest (represented by longitude and latitude key-value pairs) and the monitoring time (such as the monitoring date key-value pair). After the analysis and extraction, the data analysis system maps the spatio-temporal information to the corresponding grid cells according to the grid division rules of the spatio-temporal grid according to the data structure conversion strategy. For example, it converts the geographical location coordinates of the forest into spatial grid coordinates and maps the monitoring time to the time grid. And according to the semantic features, it maps other attribute information (such as tree species name, tree height value, etc.) in the semi-structured data to the corresponding attribute fields of the spatio-temporal grid cells. Ensure that the mapping of the attributes accurately reflects the semantics of the data, such as mapping the tree species name to the vegetation type attribute field and the tree height value to the height attribute field, and obtain the second spatio-temporal grid data.

[0099] Step 4023: Perform an association process on the structural features and semantic features of the unstructured data, and perform data conversion on the processed data based on the data structure conversion strategy to obtain the third spatio-temporal grid data.

[0100] Optionally, the data analysis system performs an association process based on the structural features (such as the keyword positions in the text, the object distribution in the image, etc.) and semantic features (such as the text theme, image semantics) of the unstructured data extracted in step 10. For example, for text data, natural language processing techniques can be used to further extract key information, such as extracting key semantic information such as forest area changes and protection measures from a forest protection report text; for image data, image recognition algorithms are used to determine the categories and positions of natural resource objects in the image, etc. Then, the semantic features are associated with the structural features to complement each other in semantics and structure. For example, in the text, the keywords are associated with the theme; in the image, the feature points are associated with the detected target semantics. Finally, according to the data structure conversion strategy, the processed unstructured data is converted into the third spatio-temporal grid data. For example, for text data, if it contains time and space related descriptions, the spatio-temporal information is extracted and mapped to the spatio-temporal grid, and the semantic and structural features are used as the attributes of the grid cells. For image data, its spatial position is determined through geolocation technology, the time dimension is determined in combination with the shooting time, and the semantic and structural features of the image are used as the attribute values.

[0101] Step 4024: Perform data fusion on the first spatio-temporal grid data, the second spatio-temporal grid data, and the third spatio-temporal grid data to obtain the target spatio-temporal grid data.

[0102] Optionally, the data analysis system takes the spatio-temporal coordinates as the benchmark, matches the grid cells with the same spatio-temporal coordinates in the first spatio-temporal grid data, the second spatio-temporal grid data, and the third spatio-temporal grid data, and for the matched grid cells, fuses their attribute information. If the attribute is numerical, methods such as weighted average and median can be used for fusion. If the attribute is text or categorical, methods such as merging or voting can be used for fusion. For example, for the land use type attribute, if it is cultivated land, farmland, and planting land in the three datasets respectively, it can be merged into cultivated land (farmland, planting land) to form complete spatio-temporal grid data. And perform quality inspection on the fused data to ensure the accuracy, consistency, and integrity of the data, and correct or supplement the data with problems.

[0103] In an embodiment of the present invention, for multi-source heterogeneous structured, semi-structured, and unstructured data, preprocessing, parsing, or correlation processing is performed respectively starting from structural features and semantic features, and then it is converted into spatio-temporal grid data based on a data structure conversion strategy, and finally data fusion is carried out. The characteristics of different types of data are fully considered, the structural and semantic information of the data is comprehensively mined, the effective conversion and fusion of multi-source heterogeneous data into spatio-temporal grid data are realized, and the quality and integrity of the target spatio-temporal grid data are improved.

[0104] In one embodiment, the descriptions of steps 501 - 504 are as follows:

[0105] Step 501: Divide the target spatio-temporal grid data according to the time sequence to obtain multiple time slices.

[0106] Optionally, after the data analysis system determines the target spatio-temporal grid data, it divides the target spatio-temporal grid data according to the time sequence, and the time interval of the division can be determined according to the data characteristics and analysis requirements. For example, for meteorological data with high-frequency monitoring, it may be divided at hourly intervals; for natural resource data with annual statistics, it can be divided by year. The continuous time series is segmented into multiple discrete time periods, and each time period corresponds to a time slice. Each time slice contains the spatial data and other relevant data features within that time period. For example, when analyzing the spatio-temporal grid data of forest resources in a certain area, the time slices are divided quarterly, and each slice contains the spatial data such as the forest area and vegetation type in that area within that quarter, as well as the relevant monitoring index data. By dividing the time slices, the data can be decomposed according to the time granularity, and data features can be mined from a more microscopic time scale.

[0107] Step 502: Analyze the type characteristics of the data within each time slice to obtain a feature analysis result.

[0108] Optionally, the data analysis system analyzes the type characteristics of the data within each time slice from multiple dimensions. In terms of dimensions, it is determined whether the data is one-dimensional, two-dimensional, or multi-dimensional. For example, the temperature in meteorological data may be one-dimensional data, while the soil moisture data containing longitude and latitude information is two-dimensional data. In terms of spatial distribution, the uniformity index of the data in space is calculated, which can be measured by calculating the variance of the data in the spatial grid. In terms of data correlation, the correlation coefficient between different variables is calculated. For example, in the spatio-temporal grid data of the ecological environment, the Pearson correlation coefficient between the vegetation coverage rate and the soil moisture is calculated to judge the degree of linear correlation between the two. Through these analyses, the feature analysis result of the data in each time slice is obtained. Comprehensively analyzing the type characteristics of the data within the time slice can deeply understand the characteristics of the data in different time periods, and provide detailed and accurate information for constructing the feature evolution trajectory and determining the processing mode.

[0109] Step 503: Construct an evolution trajectory of data type features based on the feature analysis results, and perform identification based on the evolution trajectory of data type features to determine the feature evolution pattern.

[0110] Optionally, the data analysis system plots a curve of feature variation over time with time as the horizontal axis and the value or index of each feature as the vertical axis according to the feature analysis results of each time slice, so as to construct an evolution trajectory of data type features. For example, plot a curve of the change of the spatial distribution uniformity index over time, a curve of the correlation coefficient between different variables over time, etc. By observing the shape and trend of these curves, identify the evolution patterns therein. Evolution patterns include periodic changes (such as the output of certain natural resources showing periodic fluctuations in specific seasons of each year), trend changes (such as the forest coverage rate in a certain area showing a gradually increasing or decreasing trend over time), mutations (such as sudden natural disasters causing a sudden deterioration of ecological indicators in a certain area), etc. Constructing the feature evolution trajectory and identifying the evolution pattern can intuitively show the dynamic change trend of data type features over time and grasp the change law of data from a macroscopic perspective.

[0111] Step 504: Based on the feature evolution pattern, match it with the features of different processing patterns in the preset processing pattern library to determine the optimal processing pattern.

[0112] Optionally, the data analysis system is connected to the constructed preset processing pattern library. The preset processing pattern library contains various processing patterns, such as spatio-temporal interpolation mode, graph calculation mode, tensor decomposition mode, etc. Each mode has its applicable data features and application scenarios. Compare and match the identified feature evolution pattern with the features of each mode in the preset processing pattern library. If the data shows periodic changes and involves spatio-temporal relationships, the spatio-temporal interpolation mode is more appropriate, and the periodic characteristics can be used to interpolate and predict missing data; for data with complex network relationships, such as the relationship between pixel points in an image, the graph calculation mode can be used to analyze the connection and propagation rules between nodes through graph algorithms; for high-dimensional data, the tensor decomposition mode can be used to decompose it into low-dimensional tensors to reduce data complexity. It should be noted that during the matching process, the stability of the processing effect of the processing mode on different time slices also needs to be considered to ensure the optimal overall processing effect. By matching with the preset processing pattern library, the advantages of different processing patterns can be fully utilized, and the most suitable processing method can be selected according to the feature evolution pattern of the data, improving the calculation efficiency and accuracy.

[0113] In the embodiments of the present invention, by performing time slicing division, feature analysis, constructing an evolution trajectory, and pattern recognition on spatio-temporal grid data, and finally matching with a preset processing mode library to determine the optimal processing mode. The change characteristics of the data in the time dimension are analyzed, and the evolution law of the data can be accurately grasped, so as to select the processing mode most suitable for the data characteristics. The advantages of different processing modes are fully utilized, the efficiency and accuracy of data processing are improved, the ability of the system to handle multi-source heterogeneous data of complex and changeable natural resources is enhanced, and strong support is provided for subsequent data analysis and decision-making.

[0114] Further, the multi-source heterogeneous data analysis system for natural resources provided by the present invention will be described below. The multi-source heterogeneous data analysis system described below can be mutually corresponding and referred to the multi-source heterogeneous data analysis method described above.

[0115] Optionally, referring to Figure 2 , Figure 2 is a schematic structural diagram of the multi-source heterogeneous data analysis system for natural resources provided by the present invention. The multi-source heterogeneous data analysis system includes.

[0116] An extraction and recognition module, configured to extract and recognize the obtained multi-source heterogeneous data of natural resources to obtain structural features and semantic features;

[0117] A knowledge graph construction module, configured to construct a multi-dimensional metadata mapping relationship based on the structural features and semantic features, and perform associated mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph;

[0118] A semantic alignment module, configured to perform semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantically unified data;

[0119] A data conversion module, configured to analyze the semantic information and structural information of the semantically unified data respectively, determine a data structure conversion strategy, and convert the multi-source heterogeneous data into spatio-temporal grid data based on the data structure conversion strategy;

[0120] An analysis module, configured to analyze the type characteristics of the target spatio-temporal grid data to obtain an analysis result, and determine an optimal processing mode based on the analysis result.

[0121] In the embodiments of the present invention, a multi-dimensional data mapping relationship is constructed through the structural features and semantic features in multi-source heterogeneous data, integrating the information of scattered data sources, and correlating different data in the multi-source heterogeneous data according to the multi-dimensional data mapping relationship to form a unified and dynamic metadata knowledge graph, ensuring that the data from different data sources can accurately correspond in space and time, reducing the time and workload of data preprocessing, improving the calculation speed, and effectively solving the problem of lagging analysis results in existing methods; further, semantic alignment is performed on the multi-source heterogeneous data according to the metadata knowledge graph to determine the semantically unified data with a unified semantic level, making the data form an organic whole at the semantic level, and cooperating with the dynamically generated data structure conversion strategy to convert data with different formats and structures into standardized spatio-temporal grid data, so that the data is also unified in structure, thereby avoiding the analysis error problem caused by semantic deviation and data structure differences and improving the accuracy of the analysis results; in addition, type feature analysis is performed on the structurally unified standardized spatio-temporal grid data, and the optimal data processing mode is automatically selected based on the analysis results, giving full play to the advantages of different processing modes, improving the calculation efficiency and further improving the accuracy at the same time.

[0122] Please refer to Figure 3 , Figure 3 which is the embodiment diagram of the electronic device provided by the embodiments of the present invention. As Figure 3 shown, the embodiments of the present invention provide an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored on the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, the following steps are implemented:

[0123] Extract and identify the multi-source heterogeneous data of natural resources obtained to obtain structural features and semantic features;

[0124] Based on the structural features and semantic features, construct a multi-dimensional metadata mapping relationship, and perform association mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph;

[0125] Perform semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantically unified data;

[0126] Perform semantic information and structural information analysis on the semantically unified data respectively, determine the data structure conversion strategy, and convert the multi-source heterogeneous data into spatio-temporal grid data based on the data structure conversion strategy;

[0127] Analyze the type features of the target spatio-temporal grid data to obtain analysis results, and determine the optimal processing mode based on the analysis results.

[0128] Please refer to Figure 4 , Figure 4This is an embodiment diagram of the computer-readable storage medium provided by the embodiments of the present invention. As Figure 4 shown, this embodiment provides a computer-readable storage medium 400, on which a computer program 311 is stored. When the computer program 311 is executed by a processor, the following steps are implemented:

[0129] Extract and identify the multi-source heterogeneous data of natural resources obtained, and obtain structural features and semantic features;

[0130] Based on the structural features and semantic features, construct a multi-dimensional metadata mapping relationship, and perform an association mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph;

[0131] Perform semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantically unified data;

[0132] Analyze the semantic information and structural information of the semantically unified data respectively, determine a data structure conversion strategy, and convert the multi-source heterogeneous data into spatio-temporal grid data based on the data structure conversion strategy;

[0133] Analyze the type characteristics of the target spatio-temporal grid data to obtain an analysis result, and determine an optimal processing mode based on the analysis result.

[0134] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-source heterogeneous data analysis method provided by each of the above methods. The method includes:

[0135] Extract and identify the multi-source heterogeneous data of natural resources obtained, and obtain structural features and semantic features;

[0136] Based on the structural features and semantic features, construct a multi-dimensional metadata mapping relationship, and perform an association mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph;

[0137] Perform semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantically unified data;

[0138] Analyze the semantic information and structural information of the semantically unified data respectively, determine a data structure conversion strategy, and convert the multi-source heterogeneous data into spatio-temporal grid data based on the data structure conversion strategy;

[0139] Analyze the type characteristics of the target spatio-temporal grid data to obtain an analysis result, and determine an optimal processing mode based on the analysis result.

[0140] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for analyzing multi-source heterogeneous data of natural resources, characterized in that Including: Extract and identify the multi-source heterogeneous data of natural resources obtained, and obtain structural features and semantic features; Based on the structural features and the semantic features, construct a multi-dimensional metadata mapping relationship, and perform an association mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph; Perform semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantically unified data; Analyze the semantic information and structural information of the semantically unified data respectively, determine a data structure conversion strategy, and convert the multi-source heterogeneous data into target spatio-temporal grid data based on the data structure conversion strategy; Analyze the type characteristics of the target spatio-temporal grid data to obtain an analysis result, and determine an optimal processing mode based on the analysis result.

2. The method for analyzing multi-source heterogeneous data of natural resources according to claim 1, wherein The constructing a multi-dimensional metadata mapping relationship based on the structural features and the semantic features, and performing an association mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph includes: Construct a structural feature topology graph with the structural features as nodes and the dependency relationships and hierarchical relationships between the structural features as edges; Classify the semantic features and hierarchically arrange the classification results to obtain a semantic feature hierarchy tree; Starting from the root node of the structural feature topology graph, associate the root node with the top-level node of the semantic feature hierarchy tree to obtain an initial mapping relationship framework; Based on the initial mapping relationship framework, perform association expansion along the edges of the structural feature topology graph and the branches of the semantic feature hierarchy tree to determine a multi-dimensional metadata mapping relationship; Perform an association mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain the metadata knowledge graph.

3. The method for analyzing multi-source heterogeneous data of natural resources according to claim 2, wherein The performing an association mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph includes: Extract context information from the multi-source heterogeneous data to obtain context data; Perform a similarity process on each association mapping relationship in the metadata mapping relationship based on the context data and a preset similarity function to obtain a context similarity; Filter and sort the association mapping relationships in the multi-dimensional data mapping relationship based on the context similarity to obtain candidate association mapping relationships; Perform an association mapping on the multi-source heterogeneous data based on the candidate association mapping relationships to obtain an initial knowledge graph, and label the attributes of the nodes and edges of the initial knowledge graph based on the context data to obtain the metadata knowledge graph.

4. The method for analyzing multi-source heterogeneous data of natural resources according to claim 1, wherein The performing semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantically unified data includes: For each data source in the multi-source heterogeneous data, extract nodes from the metadata knowledge graph based on the features involved in each data source to obtain feature nodes; Construct a feature topology structure for the relationships between the feature nodes in each data source; Match the feature nodes between the feature topology structures of two different data sources based on a preset structure similarity function to obtain a matching similarity; If the matching similarity is greater than or equal to a preset threshold, map the two feature nodes corresponding to the matching similarity to obtain a node mapping result; Integrate the multi-source heterogeneous data based on the node mapping result to obtain the semantically unified data.

5. The method for analyzing multi-source heterogeneous data of natural resources according to claim 1, wherein The analysis of semantic information and structural information of the semantically unified data respectively to determine the data structure conversion strategy includes: Perform semantic element analysis on the semantically unified data to obtain a semantic element analysis result, and construct a semantic association network based on the semantic element analysis result; Perform structural feature information analysis on the semantically unified data to obtain structural feature categories, and perform time and space dimension analysis on the semantically unified data to determine spatio-temporal dimension information; Map the nodes in the semantic association network to the structural feature categories to obtain an association mapping result, and determine an initial structure conversion strategy based on the association mapping result; Update the initial structure conversion strategy based on the spatio-temporal dimension information to obtain a data structure conversion strategy.

6. The method for analyzing multi-source heterogeneous data of natural resources according to claim 5, wherein The multi-source heterogeneous data includes structured data, semi-structured data, and unstructured data; Converting the multi-source heterogeneous data into target spatio-temporal grid data based on the data structure conversion strategy includes: Preprocess the structural features and semantic features of the structured data, and perform data conversion on the preprocessed data based on the data structure conversion strategy to obtain the first spatio-temporal grid data; Parse and extract the structural features and semantic features of the semi-structured data, and perform data conversion on the parsed data based on the data structure conversion strategy to obtain the second spatio-temporal grid data; Perform association processing on the structural features and semantic features of the unstructured data, and perform data conversion on the processed data based on the data structure conversion strategy to obtain the third spatio-temporal grid data; Fuse the first spatio-temporal grid data, the second spatio-temporal grid data, and the third spatio-temporal grid data to obtain the target spatio-temporal grid data.

7. The method for analyzing multi-source heterogeneous data of natural resources according to claim 1, wherein The analysis of the type features of the target spatio-temporal grid data to obtain an analysis result, and determine the optimal processing mode based on the analysis result includes: Divide the target spatio-temporal grid data by time sequence to obtain multiple time slices; Analyze the type features of the data in each time slice to obtain a feature analysis result; Construct a data type feature evolution trajectory based on the feature analysis result, and perform identification based on the data type feature evolution trajectory to determine the feature evolution mode; Match the feature evolution mode with the features of different processing modes in the preset processing mode library to determine the optimal processing mode.

8. A multi-source heterogeneous data analysis system for natural resources, characterized in that Applied to the natural resource multi-source heterogeneous data analysis method according to any one of claims 1 to 7; The natural resource multi-source heterogeneous data analysis system includes: An extraction and recognition module for extracting and recognizing the obtained natural resource multi-source heterogeneous data to obtain structural features and semantic features; A knowledge graph construction module, configured to construct a multi-dimensional metadata mapping relationship based on the structural features and the semantic features, and perform an association mapping on the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph; A semantic alignment module, configured to perform semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantically unified data; A data conversion module, configured to respectively perform semantic information and structural information analysis on the semantically unified data, determine a data structure conversion strategy, and convert the multi-source heterogeneous data into spatio-temporal grid data based on the data structure conversion strategy; An analysis module, configured to analyze the type characteristics of the target spatio-temporal grid data to obtain an analysis result, and determine an optimal processing mode based on the analysis result.

9. An electronic device, comprising: A memory, configured to store a computer software program; A processor, configured to read and execute the computer software program, wherein when the processor executes the computer software program, the natural resource multi-source heterogeneous data analysis method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium storing a computer software program, characterized in that, When the computer software program is executed by the processor, the natural resource multi-source heterogeneous data analysis method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Cooperative disambiguation method based on deep semantic neighbor and multivariate entity association

    CN112883199A

  • Spatial relation knowledge graph data model construction method and device and query method

    CN116108205A

  • Supply chain multidimensional data mining and intelligent recommendation decision-making method based on knowledge graph

    CN119741038A

  • Multi-source heterogeneous data fusion and processing method based on big data

    CN119783037A

  • Data federation method, system, electronic device, and storage medium

    WO2025043726A1

Cited By

  • Natural resource element identification method based on spatial correlation memory

    CN121095778A

  • Natural resource asset assessment method and system based on multi-source data

    CN122045796A

  • A Method and System for Natural Resource Asset Valuation Based on Multi-Source Data

    CN122045796B