A natural resource multi-source heterogeneous data analysis method and system thereof

By constructing multi-dimensional metadata mapping relationships and semantic alignment, and converting them into spatiotemporal grid data, the problems of lag and error in multi-source heterogeneous data analysis are solved, and efficient and accurate natural resource data analysis is achieved.

CN120372249BActive Publication Date: 2025-12-09GUANGDONG JINGDI PLANNING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510485726.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-12-09
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Existing methods for analyzing multi-source heterogeneous data of natural resources are time-consuming in data preprocessing and computation, and the results are delayed and erroneous due to semantic and structural differences.

Method used

By extracting structural and semantic features from multi-source heterogeneous data, a multi-dimensional metadata mapping relationship is constructed, a metadata knowledge graph is generated, and semantic alignment is performed. The data is then converted into spatiotemporal grid data, and the optimal processing mode is selected for analysis.

Benefits of technology

It improves computation speed, reduces data preprocessing time, eliminates analytical errors caused by semantic bias and structural differences, and enhances the accuracy and efficiency of analytical results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372249B_ABST
    Figure CN120372249B_ABST
Patent Text Reader

Abstract

The application provides a natural resource multi-source heterogeneous data analysis method and system, which comprises the following steps: extracting and identifying the obtained natural resource multi-source heterogeneous data to obtain structural features and semantic features; constructing a multi-dimensional metadata mapping relationship based on the structural features and the semantic features, and correlating and mapping the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph; performing semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantic unified data; analyzing the semantic information and the structural information of the semantic unified data respectively, determining a data structure conversion strategy, and converting the multi-source heterogeneous data into target space-time grid data based on the data structure conversion strategy; analyzing the type features of the target space-time grid data to obtain an analysis result, and determining an optimal processing mode based on the analysis result. The application improves the calculation speed and the accuracy of the analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of information technology, in particular to a natural resource multi-source heterogeneous data analysis method and system thereof. BACKGROUND

[0002] In the current era of rapid development of big data and information technology, the analysis of multi-source heterogeneous data in the field of natural resource management and research is increasingly valued. The analysis of multi-source heterogeneous data not only provides rich and comprehensive basis for natural resource management decision-making, but also helps to uncover the characteristics and rules of natural resources in the spatial and temporal distribution, and clarifies the internal relationship between different resources, thereby improving the efficiency of resource utilization.

[0003] However, the current multi-source heterogeneous data analysis method is not suitable for actual situations. Due to the extremely diverse natural resource data sources, not only are there many types, but also there are great differences in data structure and format. In the face of massive data, the current method needs to invest a lot of time and resources in data preprocessing and subsequent calculation process, resulting in lagging analysis results, which is difficult to provide timely and strong support for decision-making. In addition, the data generated by different data sources differ in semantics and structure, which makes it easy to deviate when integrating data for unified analysis, thereby affecting the accuracy of the analysis results. SUMMARY

[0004] The present application provides a natural resource multi-source heterogeneous data analysis method and system to improve the calculation speed, effectively solve the problem of lagging analysis results of the existing method, and avoid the analysis error problem caused by semantic deviation and data structure difference.

[0005] In a first aspect, the present application provides a natural resource multi-source heterogeneous data analysis method, comprising:

[0006] extracting and identifying the obtained natural resource multi-source heterogeneous data to obtain structural features and semantic features;

[0007] based on the structural features and the semantic features, constructing a multi-dimensional metadata mapping relationship, and based on the metadata mapping relationship, correlating and mapping the multi-source heterogeneous data to obtain a metadata knowledge graph;

[0008] based on the metadata knowledge graph, performing semantic alignment on the multi-source heterogeneous data to obtain semantic unified data;

[0009] performing semantic information and structural information analysis on the semantic unified data respectively, determining a data structure conversion strategy, and converting the multi-source heterogeneous data into spatio-temporal grid data based on the data structure conversion strategy;

[0010] Analyze the type characteristics of the target spatio-temporal grid data to obtain an analysis result, and determine an optimal processing mode based on the analysis result.

[0011] In a second aspect, the present application further provides a natural resource multi-source heterogeneous data analysis system, which is applied to the natural resource multi-source heterogeneous data analysis method as described in the first aspect. The natural resource multi-source heterogeneous data analysis system comprises:

[0012] An extraction and recognition module is configured to extract and recognize the obtained natural resource multi-source heterogeneous data to obtain structural characteristics and semantic characteristics.

[0013] A knowledge graph construction module is configured to construct a multi-dimensional metadata mapping relationship based on the structural characteristics and the semantic characteristics, and to associate and map the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph.

[0014] A semantic alignment module is configured to perform semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantic unified data.

[0015] A data conversion module is configured to analyze semantic information and structural information of the semantic unified data respectively, to determine a data structure conversion strategy, and to convert the multi-source heterogeneous data into spatio-temporal grid data based on the data structure conversion strategy.

[0016] An analysis module is configured to analyze the type characteristics of the target spatio-temporal grid data to obtain an analysis result, and to determine an optimal processing mode based on the analysis result.

[0017] In a third aspect, the present application further provides an electronic device, which comprises a memory configured to store a computer software program, and a processor configured to read and execute the computer software program to realize the natural resource multi-source heterogeneous data analysis method as described in any one of the above aspects.

[0018] In a fourth aspect, the present application further provides a non-transitory computer readable storage medium, which stores a computer software program. When the computer software program is executed by a processor, the natural resource multi-source heterogeneous data analysis method as described in any one of the above aspects is realized.

[0019] In a fifth aspect, the present application further provides a computer program product, which comprises a computer program. When the computer program is executed by a processor, the natural resource multi-source heterogeneous data analysis method as described in any one of the above aspects is realized.

[0020] The natural resource multi-source heterogeneous data analysis method provided by the embodiment of the present application integrates the scattered data source information together by constructing the multi-dimensional data mapping relationship through the structural features and semantic features in the multi-source heterogeneous data, and correlates the different data in the multi-source heterogeneous data according to the multi-dimensional data mapping relationship, forms a unified and dynamic metadata knowledge graph, ensures that the data of different data sources can be accurately corresponded in time and space, reduces the time and workload of data preprocessing, improves the calculation speed, and effectively solves the problem of lagging analysis result of the existing method; further, the multi-source heterogeneous data is semantically aligned according to the metadata knowledge graph, the semantic uniform data with unified semantic hierarchy is determined, the data forms an organic whole on the semantic level, and the data structure conversion strategy generated dynamically is used to convert the data of different formats and structures into standardized space-time grid data, so that the data also forms a unity in structure, and the analysis error problem caused by semantic deviation and data structure difference is avoided, and the accuracy of the analysis result is improved; in addition, the type feature analysis is performed on the structure-unified standardized space-time grid data, and the optimal data processing mode is automatically selected based on the analysis result, the advantages of different processing modes are fully utilized, the calculation efficiency is improved, and the accuracy is further improved. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 is a flowchart of the natural resource multi-source heterogeneous data analysis method provided by the embodiment of the present application;

[0022] Figure 2 is a structural schematic diagram of the natural resource multi-source heterogeneous data analysis system provided by the embodiment of the present application;

[0023] Figure 3 is an embodiment diagram of the electronic device provided by the embodiment of the present application;

[0024] Figure 4 is an embodiment diagram of the computer readable storage medium provided by the embodiment of the present application. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0026] In the description of the present application, the terms "first", "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise specifically limited.

[0027] In the description of the present application, the term "as" is used to indicate "as an example, illustration or explanation". Any embodiment described as "as" in the present application is not necessarily interpreted as more preferred or more advantageous than other embodiments. The following description is given in order to enable any person skilled in the art to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that those skilled in the art can realize the present application without using these specific details. In other examples, well-known structures and processes will not be described in detail to avoid unnecessary details making the description of the present application obscure. Therefore, the present application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope in accordance with the principles and characteristics disclosed.

[0028] Referring to Figure 1 , Figure 1 is a flowchart of the natural resource multi-source heterogeneous data analysis method provided by the present application. In the embodiment of the present application, the execution subject of the natural resource multi-source heterogeneous data analysis method is a data analysis system, therefore, the natural resource multi-source heterogeneous data analysis method comprises:

[0029] Step 10, extracting and identifying the obtained natural resource multi-source heterogeneous data to obtain structural features and semantic features.

[0030] Optionally, the data analysis system extracts and identifies natural resource data of different sources and formats, such as satellite remote sensing image data (e.g., in.tif format), geographic information system (GIS) vector data (e.g., in.shp format), and meteorological monitoring station text data, etc. For image data, the OpenCV library of Python can be used to read image files, and the structural features of the image, such as the resolution of the image, can be extracted by analyzing the pixel arrangement rules and band information of the image. For semantic feature extraction, a convolutional neural network (CNN) model in deep learning can be used to train the image and identify the natural resource types in the image, such as forests and water areas, to obtain semantic features. For text data, natural language processing (NLP) techniques, such as a word vector model (e.g., Word2Vec), can be used to map words in the text to a low-dimensional vector space, and semantic features can be extracted by analyzing the vector relationships between words, i.e., determining the natural resource-related concepts represented by the data, such as land use types, mineral resource types, and meteorological elements, as well as the relationships between these concepts. For structured data such as database tables, the field names, types, and constraints can be obtained by querying the database metadata to obtain structural features, and the semantic features can be determined based on the data dictionary and business rules. For semi-structured data (e.g., XML, JSON), the structural features can be obtained by parsing the tags and key-value pairs, and the semantic features can be determined based on the content and context. By comprehensively and accurately extracting the structural and semantic features of the data, a solid foundation is provided for subsequent metadata mapping and semantic alignment, which can effectively reduce information loss and misunderstanding in the data processing process and improve the reliability of the analysis results. In addition, presenting different types of data in a unified feature form also facilitates subsequent data processing and analysis, improving the relevance and accuracy of data processing.

[0031] In step 20, based on the structural features and semantic features, a multi-dimensional metadata mapping relationship is constructed, and the multi-source heterogeneous data is associated and mapped based on the metadata mapping relationship to obtain a metadata knowledge graph.

[0032] Optionally, the data analysis system extracts the structural features and semantic features obtained in step 10, and constructs a multi-dimensional metadata mapping relationship based on the association relationship between the structural features and the semantic features, as well as the natural resource data in different dimensions such as space and semantics. For data items representing the same or related concepts in different data sources, they are associated through the mapping relationship, and finally these mapping relationships are integrated into a knowledge graph to form a unified and dynamic metadata knowledge graph, which provides comprehensive semantic and structural information for subsequent data processing, as described in steps 201-205.

[0033] Further, the data analysis system can comprehensively integrate multi-source heterogeneous data by constructing multi-dimensional metadata mapping relationship, so that the originally isolated data are organically connected. Moreover, the complex relationship between data is intuitively displayed in the form of a knowledge graph, which facilitates subsequent data query and analysis.

[0034] In step 30, semantic alignment is performed on the multi-source heterogeneous data based on the metadata knowledge graph, and semantic unified data is obtained.

[0035] Optionally, according to the obtained metadata knowledge graph, the data analysis system can use alignment technology (such as ontology alignment technology) to perform semantic alignment on the same concepts and the relationship between the concepts in different data sources, and obtain semantic unified data, as described in steps 301-305. Through semantic alignment, the multi-source data has consistency at the semantic level, eliminating the understanding and processing obstacles caused by semantic expression differences, and improving the usability of the data.

[0036] Further, in an embodiment, the "cultivated land area" in the land resource data source and the "land scale for cultivation" in the agricultural statistical data source are conceptually matched.

[0037] In step 40, semantic information and structural information of the semantic unified data are analyzed, a data structure conversion strategy is determined, and the multi-source heterogeneous data is converted into spatio-temporal grid data based on the data structure conversion strategy.

[0038] Optionally, the data analysis system analyzes the semantic information and structural information of the received semantic unified data again, determines the data structure conversion strategy of different data sources in combination with the information in the spatio-temporal dimension, as described in steps 4011-4014. Then, the data analysis system converts the multi-source heterogeneous data (i.e., structured data, unstructured data, and semi-structured data) into standardized spatio-temporal grid data according to the determined data structure conversion strategy, as described in steps 4021-4024.

[0039] Further, the standardized spatio-temporal grid is a unified data representation method, which divides time and space into regular grid cells, and each grid cell contains specific attribute values. For example, for satellite remote sensing image data, it is gridized according to a certain spatio-temporal resolution, and the information of each pixel is mapped into the corresponding grid cell; for meteorological monitoring data, the monitoring values of different places and times are interpolated into the grid cells. Through this mapping, data of different formats and structures are unified into standardized spatio-temporal grid data, so as to facilitate subsequent calculation and analysis.

[0040] In step 50, the type characteristics of the target spatio-temporal grid data are analyzed, and an analysis result is obtained, and an optimal processing mode is determined based on the analysis result.

[0041] Optionally, the data analysis system conducts comprehensive data type feature analysis according to the spatio-temporal grid data. On one hand, the overall type is analyzed to determine the characteristics of the data in the spatio-temporal dimension, and on the other hand, the internal data type is analyzed in depth, and the optimal processing mode is automatically selected according to the analysis result, as described in steps 501-504.

[0042] Further, in an embodiment, for data with complex network relationships, such as the mutual relationship data between species in an ecological system, a graph computing mode is selected to analyze the connection and propagation law between nodes through a graph algorithm; for high-dimensional data, such as multi-band satellite remote sensing image data, a tensor decomposition mode is selected to decompose the high-dimensional data into low-dimensional tensors to reduce the complexity of the data; for the case where there are missing values in the spatio-temporal data, a spatio-temporal interpolation mode is selected to predict the value of the missing point according to the value of the known data point. By automatically selecting the optimal computing mode, the advantages of different computing modes are fully utilized to improve the computing efficiency and accuracy, and the flexibility and adaptability of the system are improved, so that it can better cope with complex and variable natural resource multi-source heterogeneous data.

[0043] The embodiment of the present application constructs a multi-dimensional data mapping relationship through the structural features and semantic features in the multi-source heterogeneous data, integrates the scattered data source information together, and associates different data in the multi-source heterogeneous data according to the multi-dimensional data mapping relationship to form a unified and dynamic metadata knowledge graph, so that the data of different data sources can be accurately corresponded in space and time, the time and workload of data preprocessing are reduced, the computing speed is improved, and the problem of lagging analysis result of the existing method is effectively solved; further, the multi-source heterogeneous data is semantically aligned according to the metadata knowledge graph to determine semantic unified data with unified semantic levels, so that the data forms an organic whole at the semantic level, and different formats and structures of data are converted into standardized spatio-temporal grid data by cooperating with the dynamically generated data structure conversion strategy, so that the data also forms a unity in structure, thereby avoiding the analysis error problem caused by semantic deviation and data structure difference, and improving the accuracy of the analysis result; in addition, type feature analysis is conducted on the structure-unified standardized spatio-temporal grid data, and the optimal data processing mode is automatically selected based on the analysis result, so that the advantages of different processing modes are fully utilized to improve the computing efficiency while further improving the accuracy.

[0044] In an embodiment, steps 201-205 are described as follows:

[0045] Step 201: construct a structural feature topology graph with structural features as nodes and the dependency relationship and hierarchical relationship between structural features as edges.

[0046] Optionally, after obtaining the structural features of the multi-source heterogeneous data, the data analysis system sorts out the structural features, analyzes the dependency relationship and hierarchical relationship between the structural features, wherein the dependency relationship can be a causal relationship or an association relationship, such as in the association relationship, the structural features have fields such as land type, area, and ownership. The calculation of the area may depend on the boundary information of the land type, which constitutes a dependency relationship, such as a causal relationship, the change of air pressure in meteorological data may cause the change of wind direction. The hierarchical relationship includes the inclusion relationship, such as in a large geographical area containing multiple small sub-areas; in the land use table, multiple sub-area land use tables are included. According to the dependency relationship and hierarchical relationship between the structural features, a structural feature topology graph is constructed using the graph theory method, each structural feature as a node, and the relationship as an edge. The structural feature topology graph directly shows the organization and association mode of the data structure.

[0047] Further, in an embodiment, taking a forest resource monitoring system as an example, the data sources include a tree species database, a forest area statistical document, and a pest and disease record table. The tree species field in the tree species database has a dependency relationship with the coverage area of different tree species in the forest area statistical document, that is, the area of the corresponding tree species can be calculated through the tree species information. Taking the structural features such as tree species and different tree species coverage area as nodes and the dependency relationship as edges, a structural feature topology graph is constructed.

[0048] Step 202, classifying the semantic features and hierarchically arranging the classification results to obtain a semantic feature hierarchical tree.

[0049] Optionally, the data analysis system classifies the extracted semantic features using natural language processing (NLP) technology. A pre-trained word vector model (such as BERT) can be used to represent the semantic features as vectors, and then a clustering algorithm (such as K-means clustering) can be used to cluster similar semantic features into a class. For example, forests, grasslands, and wetlands can be classified into the ecological environment class; coal mines, iron mines, and oil can be classified into the mineral resources class. After classification, the semantic features are hierarchically arranged according to the degree of abstraction. The class with high degree of abstraction is located at the top layer of the hierarchical tree, and the specific semantic features are located at the bottom layer. The hierarchical tree can be represented using a tree data structure, each node contains a class or semantic feature, and there is a hierarchical relationship between the child node and the parent node.

[0050] Step 203, starting from the root node of the structural feature topology graph, associating the root node with the top node of the semantic feature hierarchical tree to obtain an initial mapping relationship framework.

[0051] Optionally, after determining the structure feature topology graph and the semantic feature hierarchical tree, the data analysis system first determines the root node of the structure feature topology graph, which is usually the most top-level and most comprehensive structure feature, and then determines the top-level node of the semantic feature hierarchical tree, associates the root node with the top-level node, and obtains an initial mapping relationship framework. For example, in a comprehensive natural resource management data system, the root node of the structure feature topology graph can be a natural resource basic information table, and the top-level node of the semantic feature hierarchical tree can be a resource type and a resource state. The natural resource basic information table and the top-level nodes such as the resource type and the resource state are associated based on a corresponding relationship between the business logic or the domain knowledge, so as to form the initial mapping relationship framework.

[0052] In step 204, the initial mapping relationship framework is used to perform associated expansion along edges of the structure feature topology graph and branches of the semantic feature hierarchical tree, so as to determine a multi-dimensional metadata mapping relationship.

[0053] Optionally, the data analysis system starts from the root node of the structure feature topology graph, traverses the topology graph along the direction of the edges, and starts from the top-level node of the semantic feature hierarchical tree, and expands downward along the branches. When a node of the structure feature topology graph is traversed, a corresponding semantic node in the semantic feature hierarchical tree is found based on the relationship between the node and the root node and the established mapping, and a new mapping relationship is established. For example, a land use sub-table node in the structure feature topology graph is connected to a land use table root node through an edge, and based on the association between the land use table and a land use semantic category in the initial mapping, a specific semantic branch (such as farmland use details) corresponding to the land use sub-table content is found under the land use semantic category to perform association. This process is repeatedly performed until all nodes of the structure feature topology graph and all branches of the semantic feature hierarchical tree are traversed, and a multi-dimensional metadata mapping relationship is obtained.

[0054] In step 205, multi-source heterogeneous data are associated and mapped based on the metadata mapping relationship, so as to obtain a metadata knowledge graph.

[0055] Optionally, the data analysis system associates and maps the multi-source heterogeneous data based on the obtained multi-dimensional metadata mapping relationship, more accurately judges the association relationship between the data according to context information of the multi-source heterogeneous data, and thus constructs a metadata knowledge graph with better context awareness. Details are described in steps 2051-2054.

[0056] The embodiment of the application can organically combine the structure features and the semantic features in the multi-source heterogeneous data, and mine deep-level associations between the data. The integration degree and the usability of the data are improved, and strong support is provided for analysis and application of natural resource data.

[0057] In an embodiment, steps 2051-2054 are described as follows:

[0058] Step 2051, context information extraction is performed on the multi-source heterogeneous data to obtain context data.

[0059] Optionally, the data analysis system extracts context information from multi-source heterogeneous data of different sources and formats using data mining and natural language processing techniques. For structured data (such as table data in a database), information such as the association between fields, the timestamp and geographic location of the data is analyzed, such as, for meteorological data, in addition to basic data such as temperature and humidity, the collection time and collection location can also be extracted as context information. For semi-structured data (such as XML documents), the hierarchical structure of the tags and the overall theme of the document constitute the context information. For unstructured data (such as geological exploration reports), these context information is extracted from the text through topic modeling and named entity recognition methods in natural language processing techniques.

[0060] Step 2052, similarity processing is performed on each associated mapping relationship in the metadata mapping relationship based on the context data and the preset similarity function to obtain a context similarity.

[0061] Optionally, the data analysis system extracts the time information contained in the two data associated with each associated mapping relationship in the metadata mapping relationship, such as the collection time and update time of the data, and converts it into a timestamp form. The difference between the two timestamps is calculated, and the difference is converted into a value in years to obtain a time distance value S time , which is between 0 and 1. The smaller the difference, the closer the distance between the two data in the time dimension, and the stronger the correlation in the time aspect. Then, the spatial information in the associated data, i.e. the geographic coordinates (such as latitude and longitude), is extracted. And with the help of the spherical distance formula (such as using the parameters such as the radius of the earth to calculate the shortest distance between two points on the sphere), the interval between the two data in the actual space is calculated to obtain S space . The value is between 0 and 1. The smaller the spatial interval, the closer the connection between the two in the spatial dimension. Finally, the source of the associated data is evaluated, considering factors such as the nature of the data source (such as government departments, research institutions, and enterprise statistics), authority (whether peer-reviewed, widely cited, etc.). Set a certain evaluation standard, if the similarity degree of the source nature and authority is high, the value of S source is high, which is between 0 and 1. Finally, the preset similarity function is obtained, where S xsdThe context similarity value is represented as a value between 0 and 1, and α, β and γ are weight coefficients, and α+β+γ=1, which are set according to specific business requirements and the importance of different dimensions.

[0062] In step 2053, the associated mapping relationship in the multi-dimensional data mapping relationship is screened and sorted based on the context similarity, to obtain a candidate associated mapping relationship.

[0063] Optionally, the data analysis system screens the associated mapping relationship with a context similarity higher than a similarity threshold value according to the obtained context similarity value, indicating that the screened associated mapping relationship belongs to a relatively reliable and close association under the current context environment. Then, the screened associated mapping relationship is sorted from high to low according to the context similarity, to obtain a candidate associated mapping relationship. Through screening and sorting, low-quality and irrelevant associated mapping relationships can be removed, and noise and redundancy in the knowledge graph construction process can be reduced. The setting of the similarity threshold value is set according to the actual situation, and the screened associations can also be manually judged to adjust the similarity threshold value.

[0064] In step 2054, the multi-source heterogeneous data is associated and mapped based on the candidate associated mapping relationship, to obtain an initial knowledge graph, and the attributes of the nodes and edges of the initial knowledge graph are labeled based on the context data, to obtain a metadata knowledge graph.

[0065] Optionally, the data analysis system associates each data element in the multi-source heterogeneous data according to the screened and sorted candidate associated mapping relationship, to construct an initial knowledge graph. In the initial knowledge graph, the data elements (such as database fields, key concepts in the text, etc.) are nodes, and the associated mapping relationship is an edge. Then, the attributes of the nodes and edges in the initial knowledge graph are labeled using the context data. For the nodes, the data source, data type, context theme and other attributes are labeled. For the edges, the basis for association, context similarity value, description of the closeness of association and other attributes are labeled, to obtain a metadata knowledge graph.

[0066] Further, in an embodiment, taking processing multi-source heterogeneous data about natural resources in a certain region as an example, the data elements such as forests, mineral resources and rivers are associated according to the candidate associated mapping relationship, to construct an initial knowledge graph. Then, the attributes of the nodes and edges in the knowledge graph are labeled according to the context data (such as the time of forest fire, the history of mineral resource exploitation, etc.). For example, the XX forest node is labeled with the information that a fire occurred on XX month XX day in XX year, and the XX forest-ecological protection edge is labeled with the information that it was affected by the fire.

[0067] The embodiment of the present application can accurately screen out meaningful association mapping relationship in complex multi-source heterogeneous data by introducing context information and similarity processing, and construct a more accurate, rich and practical metadata knowledge graph.

[0068] In an embodiment, steps 301-305 are described as follows:

[0069] In step 301, for each data source in the multi-source heterogeneous data, node extraction is performed on the metadata knowledge graph based on the features involved in each data source, to obtain feature nodes.

[0070] Optionally, after receiving the multi-source heterogeneous data and determining the metadata knowledge graph, the data analysis system analyzes the data of each data source due to the diversified sources of the multi-source heterogeneous data, and analyzes the natural resource feature concept information contained therein. For example, in the forestry resource data source, the concept information reflecting the characteristics of forestry, such as forest coverage rate and tree species quantity, etc. In addition, the data analysis system pre-constructs a knowledge base of natural resource related concepts, and after extracting the natural resource feature concept information, obtains the standard definition, attribute information, etc. from the pre-constructed knowledge base of natural resource related concepts, such as the definition of forest coverage rate is the percentage of forest area to total land area, and the attributes include statistical time, statistical region, etc. Then, starting from each data source, the corresponding node in the metadata knowledge graph is matched according to the natural resource feature concept information (i.e. the features involved in the data source), as a feature node. Ensure that the extracted feature node has accurate and complete semantic information. Combined with the knowledge base information, concept misunderstanding can be avoided, and the accuracy and consistency of node semantics are ensured, laying a foundation for building reliable semantic unified data.

[0071] In step 302, the relationship between the feature nodes in each data source is constructed to form a feature topology structure.

[0072] Optionally, after obtaining the feature nodes of each data source, the data analysis system analyzes the relationship between the feature nodes, including causal relationship, hierarchical relationship, etc. For example, in the meteorological monitoring data source, there may be a certain correlation between temperature and humidity, temperature rise may lead to humidity decrease, and wind speed and wind direction are mutually related and parallel relationship; in the geographic information database, there may be a containing relationship between land use type and landform, such as mountain may correspond to forest land and other land use types, and the forest ecosystem contains the hierarchical relationship of tree species, soil quality and other sub-concepts, and then the feature nodes are taken as the vertices and the relationship between them is taken as the edges to construct a directed or undirected graph to represent the feature (concept) topology structure. The relationship between the feature nodes in each data source can be intuitively displayed, which helps to deeply understand the internal logic of the data source.

[0073] Step 303, matching the feature nodes between the feature topology structures of the two different data sources based on a preset structure similarity function to obtain a matching similarity.

[0074] Optionally, the data analysis system presets a structure similarity function for measuring the similarity of the feature topology structures of different data sources, and comprehensively considers the weight relationship of nodes and edges to match the feature nodes between the feature topology structures of the two different data sources.

[0075] Further, for two feature topology structures G1=(V1, E1, W1) and G2=(V2, E2, W2) (V represents a node set, E represents an edge set, and W represents a weight set), when calculating the structure similarity, the nodes are matched first. For nodes v i1 ∈V1 and v i2 ∈V2, if they represent similar concepts (judged by the semantic similarity before), they are considered as matched nodes, denoted as δ i1i2 =1, otherwise δ i1i2 =0. The matching of edges is similar, if the nodes connected by edges e j1 ∈E1 and e j2 ∈E2 are matched and the weight difference of the edges is within a certain range (set as Δw), the edges are considered to be matched. The matching similarity formula is:

[0076] Where n1 and n2 represent the number of nodes of G1 and G2 respectively; ω ij is the weight of the relationship between nodes i and j (if the nodes are not connected, ω ij =0); δ ij is a binary variable, used to represent whether nodes i and j are matched, if nodes i and j represent similar concepts, δ ij =1, otherwise δ ij =0; ω ij1 is the weight of the edge connecting nodes i and j in G1 (if nodes i and j are not connected, the value has no actual meaning); ω ij2 is the weight of the edge connecting the corresponding nodes (the nodes matched with nodes i and j in G1) in G2 (if the corresponding nodes are not connected, the value has no actual meaning), which are used to measure the influence of the weight difference of the edges on the structure similarity.

[0077] Further, in an embodiment, two data sources about natural resources have feature topology structures G1 and G2 respectively. In G1, there are forest area, tree species nodes and edges with a weight of 0.6 between them, and in G2, there are forest land range (semantically similar to forest area), tree species (semantically similar to tree species) nodes and edges with a weight of 0.5 between them. Through the structure similarity formula calculation, ω ij is the weight of the corresponding edge, and δij = 1 (node match), |ω ij1 - ω ij2 | = |0.6 - 0.5| = 0.1, the similarity contribution value of the two feature topologies in this part is calculated by substituting the formula, and the overall matching similarity is obtained by comprehending the conditions of other nodes and edges.

[0078] In step 304, if the matching similarity is greater than or equal to the preset threshold, the two feature nodes corresponding to the matching similarity are mapped to obtain a node mapping result.

[0079] Optionally, the data analysis system determines the preset threshold according to the standard determined according to the actual application requirement and the data characteristics, which is used to determine whether the feature nodes can be mapped. When the matching similarity of the feature topologies of the two data sources reaches or exceeds the threshold, it indicates that the corresponding feature nodes have high consistency in semantics and structure, and a mapping relationship is established. For example, in different regional water resource monitoring data sources, the matching similarity of drinking water source water quality nodes is higher than the threshold, which indicates that they are similar in semantics and can be mapped to the same semantic concept.

[0080] In step 305, the multi-source heterogeneous data is integrated based on the node mapping result to obtain semantic unified data.

[0081] Optionally, the data analysis system integrates the multi-source heterogeneous data according to the obtained node mapping result. First, the data format of different data sources mapped to the same concept is unified. For example, the unit of forest area in one data source is hectare, and the unit of forest area in another data source is square kilometer, which is unified to hectare. Then the data values are combined. If there are multiple data sources measuring the same concept, weighted average operation can be performed according to the reliability, timeliness and other factors of the data. Then the data content is combined, such as combining the monitoring data of different regions on precipitation according to the time sequence, and at the same time, the semantic description of the data is also unified and standardized, such as unifying different names of forest in different data sources to forest. After these processes, the differences between multi-source heterogeneous data are eliminated, and semantic unified data is obtained.

[0082] The embodiment of the application solves the problem of semantic inconsistency of multi-source heterogeneous data by fully utilizing the metadata knowledge graph and the internal structure information of the data source through the node extraction, topology structure construction, similarity calculation, node mapping and data integration process of the system. And a unique matching similarity function is designed, which can accurately identify the feature nodes with similar semantics and structure in different data sources, and provide a reliable basis for data integration.

[0083] In an embodiment, steps 4011-4014 are described as follows:

[0084] Step 4011, semantic element analysis is performed on the semantic unified data to obtain a semantic element analysis result, and a semantic correlation network is constructed based on the semantic element analysis result.

[0085] Optionally, the data analysis system deeply analyzes the semantic elements of the semantic unified data, and disassembles the semantic elements into specific semantic elements. For example, in natural resource data, forests, tree species, and water resource reserves are all semantic elements. The type and attribute of each semantic element can be determined mainly through natural language processing techniques such as part-of-speech tagging and named entity recognition. Then, according to the logical relationship between the semantic elements, the direct semantic correlation between the elements is determined. For example, forests and tree species have a direct semantic correlation because forests are composed of multiple tree species. The elements with direct semantic correlation are connected by edges to construct a semantic correlation network. In the semantic correlation network, nodes represent semantic elements, and edges represent semantic correlation relationships between elements.

[0086] Step 4012, structure feature information analysis is performed on the semantic unified data to obtain a structure feature category, and time and space dimension analysis is performed on the semantic unified data to determine spatio-temporal dimension information.

[0087] Optionally, the data analysis system analyzes the result features of the semantic unified data, and analyzes the structure features according to the representation form and storage method of the data. In the analysis process, if the data describes natural resources in the form of text paragraphs, it belongs to a text structure; if the data is a series of numerical records, such as natural resource index values monitored, it belongs to a numerical structure; if the data is recorded in chronological order, such as monthly monitoring data of a river flow, it belongs to a time series structure. At the same time, the data analysis system also analyzes the semantic unified data in time and space. In the time dimension, the time point, time period, etc. corresponding to the data are determined, such as the monitoring time in forest resource monitoring data; in the spatial dimension, the geographical location and regional range involved by the data are determined through geographical coordinates, place names, etc. to finally obtain spatio-temporal dimension information.

[0088] Step 4013, mapping is performed between the nodes in the semantic correlation network and the structure feature categories to obtain a correlation mapping result, and an initial structure conversion strategy is determined based on the correlation mapping result.

[0089] Optionally, the data analysis system maps each node (i.e., semantic element) in the semantic correlation network to a structure feature category. For example, the semantic node of forest area value is mapped to the numerical structure category because it involves a specific numerical value. Through this mapping method, the nodes in the semantic correlation network are traversed to establish a connection between semantics and structure, i.e., to obtain a correlation mapping result.

[0090] Further, after determining the association mapping result, the data analysis system formulates a corresponding initial structure conversion strategy. For example, if the semantic node belongs to a text structure and is related to a spatial position, the strategy is to extract the description of the spatial position in the text and convert it into grid coordinates. If the semantic node belongs to a numerical structure and is related to time, the data is sorted according to the time sequence. It should be noted that the initial structure conversion strategy provides preliminary guidance for subsequent data conversion, can process data with different semantics and structures in a targeted manner, and improves the efficiency and accuracy of data processing.

[0091] In step 4014, the initial structure conversion strategy is updated based on the spatio-temporal dimension information to obtain a data structure conversion strategy.

[0092] Optionally, after determining the initial structure conversion strategy, the data analysis system optimizes and updates the initial structure conversion strategy in combination with the spatio-temporal dimension information determined in step 4012. On the one hand, after considering the time factor, for spatial data that changes over time, such as urban land use change data at different time points, the conversion strategy can be refined to convert spatial data at different time points into grid data on the corresponding time slice in sequence according to the time sequence. On the other hand, after considering the spatial factor, the conversion strategy is adjusted according to the geographical location and regional range involved in the data, so that it is more in line with the actual spatial distribution. Therefore, by adding the spatio-temporal dimension information, the initial structure conversion strategy can be more in line with the requirements of the spatio-temporal grid data, and the converted data can accurately reflect the spatio-temporal characteristics of the data.

[0093] The embodiment of the present application gradually constructs a semantic association network, determines the structure feature category and spatio-temporal dimension information by performing semantic element analysis, structure feature analysis and spatio-temporal dimension analysis on the unified semantic data, and formulates and updates the data structure conversion strategy through semantic-structure association mapping. The semantic, structure and spatio-temporal characteristics of the data are comprehensively considered, the data connotation can be deeply understood, the data structure conversion strategy is more scientific and reasonable, the foundation is laid for subsequent processing such as converting multi-source heterogeneous data into spatio-temporal grid data, and the analysis and utilization efficiency of natural resource data is improved.

[0094] In an embodiment, the multi-source heterogeneous data includes structured data, semi-structured data and unstructured data. Steps 4021-4025 are described as follows:

[0095] In step 4021, the structure features and semantic features of the structured data are preprocessed, and the preprocessed data is converted based on the data structure conversion strategy to obtain first spatio-temporal grid data.

[0096] Optionally, the data analysis system pre-processes the structured data extracted in step 10 according to the structural features and semantic features of the structured data, such as for a database table recording land resource information, determining which fields are related to spatial location (such as land block coordinate field) and which are related to resource attributes (such as land type field) according to the semantic features of the fields, and if there are data missing or errors, filling or correcting them according to the semantic features and business rules of the data. After the pre-processing, the data analysis system maps the spatio-temporal related information in the structured data into the spatio-temporal grid according to the determined data structure conversion strategy, such as for the table with explicit geographic coordinate fields (such as longitude and latitude), the coordinates are mapped to the corresponding spatial grid cells according to the grid division rules. For time information (such as data recording time), it is mapped to the grid in the time dimension according to the time granularity (such as division by day, month). Other attributes (such as land type, area, etc.) between the spatio-temporal information in the structured data are filled into the corresponding spatio-temporal grid cells. Specifically, the storage mode and format of the attributes can be determined according to the semantic features of the data to ensure that their meanings can be accurately expressed in the spatio-temporal grid data. For example, the land type is stored in the attribute field of the corresponding grid cell in the form of text and the area is stored in the form of numerical value, and finally the first spatio-temporal grid data converted from the structured data is obtained.

[0097] In step 4022, the structural features and semantic features of the semi-structured data are analyzed and extracted, and the parsed data is converted based on the data structure conversion strategy to obtain the second spatio-temporal grid data.

[0098] Optionally, the data analysis system uses parsing tools to further analyze and process the structural features (such as tag hierarchy of XML, key-value pair structure of JSON) and semantic features (such as the meaning represented by the tags or key-value pairs) of the semi-structured data extracted in step 10, such as for a JSON file describing forest resources, determining the organization form of the data according to the structural features of the key-value pairs, and determining the forest resource information (such as tree species, tree height) represented by each key-value pair according to the semantic features. Then, the spatio-temporal related information is extracted from the parsed data, such as the geographic location of the forest (represented by the longitude and latitude key-value pairs) and the monitoring time (such as the monitoring date key-value pair). After the analysis and extraction, the data analysis system maps the spatio-temporal information to the corresponding grid cells according to the division rules of the spatio-temporal grid according to the data structure conversion strategy. For example, the geographic location coordinates of the forest are converted to spatial grid coordinates, and the monitoring time is mapped to the time grid. And according to the semantic features, other attribute information (such as tree species name, tree height value, etc.) in the semi-structured data is mapped to the corresponding attribute field of the spatio-temporal grid cell. Ensuring that the mapping of the attributes accurately reflects the semantics of the data, such as mapping the tree species name to the vegetation type attribute field and the tree height value to the height attribute field, to obtain the second spatio-temporal grid data.

[0099] At step 4023, the structure features and semantic features of the unstructured data are associated, and the processed data is converted based on a data structure conversion strategy to obtain third spatio-temporal grid data.

[0100] Optionally, the data analysis system performs associated processing on the structure features (such as keyword position in text, object distribution in image, etc.) and semantic features (such as text theme, image semantics) of the unstructured data extracted in step 10. For example, for text data, natural language processing techniques can be used to further extract key information, such as extracting key semantic information such as forest area change, protection measures, etc. from forest protection report text; for image data, image recognition algorithms are used to determine the category and position of natural resource objects in the image. Then the semantic features are associated with the structure features, so that the semantics and structure complement each other. For example, in text, keywords are associated with themes; in images, feature points are associated with detected target semantics. Finally, the processed unstructured data is converted into third spatio-temporal grid data according to the data structure conversion strategy, such as for text data, if it contains time and space related descriptions, extract spatio-temporal information and map to spatio-temporal grid, and semantic and structure features as attributes of grid cells. For image data, determine its spatial position by geolocation technology, and determine the time dimension by combining the shooting time, and the semantic and structure features of the image as attribute values.

[0101] At step 4024, the first spatio-temporal grid data, the second spatio-temporal grid data, and the third spatio-temporal grid data are fused to obtain target spatio-temporal grid data.

[0102] Optionally, the data analysis system takes the spatio-temporal coordinates as the reference, matches the grid cells with the same spatio-temporal coordinates in the first spatio-temporal grid data, the second spatio-temporal grid data, and the third spatio-temporal grid data, and fuses the attribute information of the matched grid cells. If the attribute is numerical, weighted average, median, etc. can be used for fusion. If the attribute is text or category, merging or voting, etc. can be used for fusion. For example, for the land use type attribute, if the three data sets are respectively cultivated land, farmland, and planting land, they can be merged into cultivated land (farmland, planting land) to form complete spatio-temporal grid data. The fused data is checked for quality to ensure accuracy, consistency and completeness, and the data with problems is corrected or supplemented.

[0103] The embodiment of the present application is aimed at multi-source heterogeneous structured, semi-structured and unstructured data, and respectively pre-processes, analyzes or correlates from the structural features and semantic features, and then converts into the spatio-temporal grid data based on the data structure conversion strategy, and finally performs data fusion. The characteristics of different types of data are fully considered, the structure and semantic information of the data are fully mined, the effective conversion and fusion of the multi-source heterogeneous data into the spatio-temporal grid data are realized, and the quality and integrity of the target spatio-temporal grid data are improved.

[0104] In an embodiment, steps 501-504 are described as follows:

[0105] Step 501, the target spatio-temporal grid data is divided in time sequence to obtain a plurality of time slices.

[0106] Optionally, after the data analysis system determines the target spatio-temporal grid data, the target spatio-temporal grid data is divided according to the time sequence, and the time interval of the division can be determined according to the data characteristics and analysis requirements. For example, for high-frequency monitoring meteorological data, the time interval can be divided by hours; for annual statistical natural resource data, the time interval can be divided by years. The continuous time sequence is divided into a plurality of discrete time periods, and each time period corresponds to a time slice. Each time slice contains spatial data and other related data characteristics in the time period, such as when analyzing the spatio-temporal grid data of forest resources in a certain area, the time slice is divided by quarter, and each slice contains the area of the forest in the quarter, the vegetation type and other related monitoring index data in the area. By dividing the time slice, the data can be decomposed according to the time granularity, and the data characteristics can be mined from a more micro time scale.

[0107] Step 502, the type characteristics of the data in each time slice are analyzed to obtain a feature analysis result.

[0108] Optionally, the data analysis system analyzes the type characteristics of the data in each time slice from multiple dimensions. In terms of dimensions, it is determined whether the data is one-dimensional, two-dimensional or multi-dimensional, such as temperature in meteorological data which can be one-dimensional data, and soil moisture data containing latitude and longitude information which is two-dimensional data. In terms of spatial distribution, the uniformity index of the data in space is calculated, which can be measured by calculating the variance of the data in the spatial grid. In terms of data correlation, the correlation coefficient between different variables is calculated, such as the Pearson correlation coefficient between vegetation coverage and soil moisture in ecological environment spatio-temporal grid data, to judge the linear correlation degree of the two. Through these analyses, the feature analysis result of the data in each time slice is obtained. Comprehensive analysis of the type characteristics of the data in the time slice can deeply understand the characteristics of the data in different time periods, and provide detailed and accurate information for constructing the feature evolution track and determining the processing mode.

[0109] At step 503, the data type feature evolution track is constructed based on the feature analysis result, and the feature evolution mode is determined based on the data type feature evolution track.

[0110] Optionally, the data analysis system draws a curve of the change of the feature over time according to the feature analysis result of each time slice, with time as the horizontal axis and the value or index of each feature as the vertical axis, thereby constructing the evolution track of the data type feature. For example, a curve of the change of the spatial distribution uniformity index over time, a curve of the change of the correlation coefficient between different variables over time, etc. are drawn. By observing the shape and trend of these curves, the evolution mode is identified. The evolution mode includes periodic change (e.g. the yield of some natural resources presents periodic fluctuation in a specific season of each year), trend change (e.g. the forest coverage rate of a certain region presents a gradual upward or downward trend over time), mutation (e.g. a sudden natural disaster causes the ecological index of a certain region to suddenly deteriorate), etc. Constructing the feature evolution track and identifying the evolution mode can intuitively show the dynamic change trend of the data type feature over time and grasp the change rule of the data from a macro perspective.

[0111] At step 504, the optimal processing mode is determined by matching the feature evolution mode with the features of different processing modes in the preset processing mode library.

[0112] Optionally, the data analysis system is connected with the constructed preset processing mode library, and the preset processing mode library contains multiple processing modes, such as spatio-temporal interpolation mode, graph calculation mode, tensor decomposition mode, etc. Each mode has its applicable data feature and application scenario. The identified feature evolution mode is compared and matched with the features of each mode in the preset processing mode library. If the data presents periodic change and involves spatio-temporal relationship, the spatio-temporal interpolation mode is more appropriate, and the periodic characteristics can be used to interpolate and predict the missing data; for data with complex network relationship, such as the relationship between pixel points in an image, the graph calculation mode can be used to analyze the connection and propagation law between nodes through graph algorithm; for high-dimensional data, the tensor decomposition mode can be used to decompose it into low-dimensional tensors, reducing the data complexity. It should be noted that in the matching process, the stability of the processing effect of the processing mode on different time slices also needs to be considered to ensure the optimal overall processing effect. By matching with the preset processing mode library, the advantages of different processing modes can be fully utilized, the most suitable processing method can be selected according to the feature evolution mode of the data, and the calculation efficiency and accuracy are improved.

[0113] The embodiment of the present application divides the space-time grid data by time slicing, analyzes the features, constructs the evolution track and mode recognition, and finally matches with the preset processing mode library to determine the optimal processing mode. The change characteristics of the data in the time dimension are analyzed, the evolution law of the data can be accurately grasped, and thus the processing mode most suitable for the data characteristics is selected. The advantages of different processing modes are fully utilized, the efficiency and accuracy of data processing are improved, the ability of the system to deal with complex and changeable natural resource multi-source heterogeneous data is enhanced, and strong support is provided for subsequent data analysis and decision-making.

[0114] Further, the natural resource multi-source heterogeneous data analysis system provided by the present application is described below, and the natural resource multi-source heterogeneous data analysis system described below can be correspondingly referred to the natural resource multi-source heterogeneous data analysis method described above.

[0115] Optionally, referring to Figure 2 , Figure 2 is a structural schematic diagram of the natural resource multi-source heterogeneous data analysis system provided by the present application, and the natural resource multi-source heterogeneous data analysis system comprises.

[0116] The extraction and recognition module is configured to extract and recognize the obtained natural resource multi-source heterogeneous data to obtain structural features and semantic features.

[0117] The knowledge graph construction module is configured to construct a multi-dimensional metadata mapping relationship based on the structural features and the semantic features, and to associate and map the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph.

[0118] The semantic alignment module is configured to perform semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantic unified data.

[0119] The data conversion module is configured to analyze semantic information and structural information of the semantic unified data respectively, to determine a data structure conversion strategy, and to convert the multi-source heterogeneous data into space-time grid data based on the data structure conversion strategy.

[0120] The analysis module is configured to analyze the type features of the target space-time grid data to obtain an analysis result, and to determine an optimal processing mode based on the analysis result.

[0121] The embodiment of the present application constructs a multi-dimensional data mapping relationship through the structural features and semantic features in the multi-source heterogeneous data, integrates the scattered data source information together, and correlates different data in the multi-source heterogeneous data according to the multi-dimensional data mapping relationship, forms a unified and dynamic metadata knowledge graph, ensures that the data of different data sources can be accurately corresponded in time and space, reduces the time and workload of data preprocessing, improves the calculation speed, and effectively solves the problem of lagging analysis result of the existing method; further, the multi-source heterogeneous data is semantically aligned according to the metadata knowledge graph, the semantic uniform data of the semantic level is determined, the data forms an organic whole on the semantic level, and the data of different formats and structures is converted into standardized space-time grid data by cooperating with the dynamically generated data structure conversion strategy, so that the data also forms unity on the structure, and the analysis error problem caused by semantic deviation and data structure difference is avoided, and the accuracy of the analysis result is improved; in addition, the type feature analysis is performed on the structure-unified standardized space-time grid data, and the optimal data processing mode is automatically selected based on the analysis result, the advantages of different processing modes are fully utilized, the calculation efficiency is improved, and the accuracy is further improved.

[0122] Please refer to Figure 3 , Figure 3 The embodiment of the electronic device provided by the present application is shown in the figure. As shown in Figure 3 The present embodiment provides an electronic device 300, which includes a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, the following steps are implemented:

[0123] The obtained natural resource multi-source heterogeneous data is extracted and identified to obtain structural features and semantic features;

[0124] Based on the structural features and semantic features, a multi-dimensional metadata mapping relationship is constructed, and the multi-source heterogeneous data is associated and mapped based on the metadata mapping relationship to obtain a metadata knowledge graph;

[0125] Based on the metadata knowledge graph, the multi-source heterogeneous data is semantically aligned to obtain semantic uniform data;

[0126] The semantic information and structural information of the semantic uniform data are analyzed respectively to determine a data structure conversion strategy, and the multi-source heterogeneous data is converted into space-time grid data based on the data structure conversion strategy;

[0127] The type features of the target space-time grid data are analyzed to obtain an analysis result, and the optimal processing mode is determined based on the analysis result.

[0128] Please refer to Figure 4 , Figure 4An embodiment of a computer readable storage medium provided for an embodiment of the present application is shown in the figure. Figure 4 As shown in the figure, the present embodiment provides a computer readable storage medium 400, which stores a computer program 311, and the computer program 311 is executed by a processor to implement the following steps:

[0129] The obtained natural resource multi-source heterogeneous data is extracted and identified to obtain structural features and semantic features;

[0130] Based on the structural features and the semantic features, a multi-dimensional metadata mapping relationship is constructed, and the multi-source heterogeneous data is associated and mapped based on the metadata mapping relationship to obtain a metadata knowledge graph;

[0131] Based on the metadata knowledge graph, the multi-source heterogeneous data is semantically aligned to obtain semantic unified data;

[0132] The semantic unified data is analyzed for semantic information and structural information respectively, a data structure conversion strategy is determined, and the multi-source heterogeneous data is converted into spatio-temporal grid data based on the data structure conversion strategy;

[0133] The type features of the target spatio-temporal grid data are analyzed to obtain an analysis result, and an optimal processing mode is determined based on the analysis result.

[0134] On the other hand, the present application also provides a computer program product, which includes a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program is executed by a processor, and the computer can execute the natural resource multi-source heterogeneous data analysis method provided by the above-mentioned methods, the method includes:

[0135] The obtained natural resource multi-source heterogeneous data is extracted and identified to obtain structural features and semantic features;

[0136] Based on the structural features and the semantic features, a multi-dimensional metadata mapping relationship is constructed, and the multi-source heterogeneous data is associated and mapped based on the metadata mapping relationship to obtain a metadata knowledge graph;

[0137] Based on the metadata knowledge graph, the multi-source heterogeneous data is semantically aligned to obtain semantic unified data;

[0138] The semantic unified data is analyzed for semantic information and structural information respectively, a data structure conversion strategy is determined, and the multi-source heterogeneous data is converted into spatio-temporal grid data based on the data structure conversion strategy;

[0139] The type features of the target spatio-temporal grid data are analyzed to obtain an analysis result, and an optimal processing mode is determined based on the analysis result.

[0140] The system embodiments described above are merely illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and necessary universal hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and include a plurality of instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods of the embodiments or some parts of the embodiments.

[0142] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A natural resource multi-source heterogeneous data analysis method, characterized in that, The method comprises the following steps: extracting and identifying the obtained natural resource multi-source heterogeneous data to obtain structural features and semantic features; the natural resource multi-source heterogeneous data comprises satellite remote sensing image data, geographic information system vector data, and meteorological monitoring station text data; based on the structural features and the semantic features, constructing a multi-dimensional metadata mapping relationship, and based on the metadata mapping relationship, correlating and mapping the multi-source heterogeneous data to obtain a metadata knowledge graph; based on the metadata knowledge graph, performing semantic alignment on the multi-source heterogeneous data to obtain semantic unified data; performing semantic information and structural information analysis on the semantic unified data, respectively, determining a data structure conversion strategy, and based on the data structure conversion strategy, converting the multi-source heterogeneous data into target spatio-temporal grid data; analyzing the type features of the target spatio-temporal grid data to obtain an analysis result, and determining an optimal processing mode based on the analysis result; based on the structural features and the semantic features, constructing a multi-dimensional metadata mapping relationship, and based on the metadata mapping relationship, correlating and mapping the multi-source heterogeneous data to obtain a metadata knowledge graph, comprising: constructing a structural feature topology graph by taking the structural features as nodes and the dependency relationships and hierarchical relationships between the structural features as edges; classifying the semantic features and hierarchically arranging the classification results to obtain a semantic feature hierarchical tree; starting from the root node of the structural feature topology graph, correlating the root node with the top node of the semantic feature hierarchical tree to obtain an initial mapping relationship framework; based on the initial mapping relationship framework, correlating and expanding along the edges of the structural feature topology graph and the branches of the semantic feature hierarchical tree to determine a multi-dimensional metadata mapping relationship; based on the metadata mapping relationship, correlating and mapping the multi-source heterogeneous data to obtain the metadata knowledge graph; based on the metadata mapping relationship, correlating and mapping the multi-source heterogeneous data to obtain the metadata knowledge graph, comprising: extracting context information from the multi-source heterogeneous data to obtain context data; based on the context data and a preset similarity function, performing similarity processing on each correlation mapping relationship in the metadata mapping relationship to obtain a context similarity; based on the context similarity, screening and sorting the correlation mapping relationships in the multi-dimensional data mapping relationship to obtain candidate correlation mapping relationships; based on the candidate correlation mapping relationships, correlating and mapping the multi-source heterogeneous data to obtain an initial knowledge graph, and based on the context data, labeling the attributes of the nodes and edges of the initial knowledge graph to obtain the metadata knowledge graph.

2. The natural resource multi-source heterogeneous data analysis method according to claim 1, characterized in that, based on the metadata knowledge graph, performing semantic alignment on the multi-source heterogeneous data to obtain semantic unified data, comprising: for each data source in the multi-source heterogeneous data, extracting nodes from the metadata knowledge graph based on the features involved in each data source to obtain feature nodes; for the relationships between the feature nodes in each data source, constructing a feature topology structure; The feature nodes between the feature topology structures of two different data sources are matched based on a preset structure similarity function to obtain a matching similarity; If the matching similarity is greater than or equal to a preset threshold, two feature nodes corresponding to the matching similarity are mapped to obtain a node mapping result; The multi-source heterogeneous data is integrated based on the node mapping result to obtain the semantic unified data.

3. The natural resource multi-source heterogeneous data analysis method according to claim 1, characterized in that, The semantic unified data is analyzed for semantic information and structure information respectively to determine a data structure conversion strategy, including: The semantic unified data is analyzed for semantic elements to obtain a semantic element analysis result, and a semantic correlation network is constructed based on the semantic element analysis result; The semantic unified data is analyzed for structure feature information to obtain a structure feature category, and the semantic unified data is analyzed for time and space dimensions to determine time and space dimension information; The nodes in the semantic correlation network are mapped with the structure feature category to obtain a correlation mapping result, and an initial structure conversion strategy is determined based on the correlation mapping result; The initial structure conversion strategy is updated based on the time and space dimension information to obtain the data structure conversion strategy.

4. The natural resource multi-source heterogeneous data analysis method according to claim 3, characterized in that, The multi-source heterogeneous data includes structured data, semi-structured data and unstructured data; The multi-source heterogeneous data is converted into target spatio-temporal grid data based on the data structure conversion strategy, including: The structure features and semantic features of the structured data are preprocessed, and the preprocessed data is converted based on the data structure conversion strategy to obtain first spatio-temporal grid data; The structure features and semantic features of the semi-structured data are parsed and extracted, and the parsed data is converted based on the data structure conversion strategy to obtain second spatio-temporal grid data; The structure features and semantic features of the unstructured data are associated and processed, and the processed data is converted based on the data structure conversion strategy to obtain third spatio-temporal grid data; The first spatio-temporal grid data, the second spatio-temporal grid data and the third spatio-temporal grid data are fused to obtain the target spatio-temporal grid data.

5. The natural resource multi-source heterogeneous data analysis method according to claim 1, characterized in that, The type features of the target spatio-temporal grid data are analyzed to obtain an analysis result, and an optimal processing mode is determined based on the analysis result, including: The target spatio-temporal grid data is divided in time sequence to obtain a plurality of time slices; The type features of the data in each time slice are analyzed to obtain a feature analysis result; A data type feature evolution trajectory is constructed based on the feature analysis result, and a feature evolution mode is determined based on the data type feature evolution trajectory; The feature evolution mode is matched with the features of different processing modes in a preset processing mode library to determine an optimal processing mode.

6. A natural resource multi-source heterogeneous data analysis system, characterized in that, The natural resource multi-source heterogeneous data analysis method of any one of claims 1 to 5; The natural resource multi-source heterogeneous data analysis system includes: An extraction and recognition module is configured to extract and recognize the obtained natural resource multi-source heterogeneous data to obtain structural features and semantic features; the natural resource multi-source heterogeneous data includes satellite remote sensing image data, geographic information system vector data, and meteorological monitoring station text data; A knowledge graph construction module is configured to construct a multi-dimensional metadata mapping relationship based on the structural features and the semantic features, and to associate and map the multi-source heterogeneous data based on the metadata mapping relationship to obtain a metadata knowledge graph; A semantic alignment module is configured to perform semantic alignment on the multi-source heterogeneous data based on the metadata knowledge graph to obtain semantic unified data; A data conversion module is configured to analyze semantic information and structural information of the semantic unified data respectively, to determine a data structure conversion strategy, and to convert the multi-source heterogeneous data into spatio-temporal grid data based on the data structure conversion strategy; An analysis module is configured to analyze type features of the target spatio-temporal grid data to obtain an analysis result, and to determine an optimal processing mode based on the analysis result.

7. An electronic device, comprising: A memory is configured to store a computer software program; A processor is configured to read and execute the computer software program, and when the processor executes the computer software program, the natural resource multi-source heterogeneous data analysis method of any one of claims 1 to 5 is implemented.

8. A non-transitory computer readable storage medium having stored therein a computer software program, characterized in that, When the computer software program is executed by the processor, the natural resource multi-source heterogeneous data analysis method of any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Supply chain multidimensional data mining and intelligent recommendation decision-making method based on knowledge graph

    CN119741038A

  • Multi-source heterogeneous data fusion and processing method based on big data

    CN119783037A