A geographic information data processing system based on cloud computing

By designing a geographic information data processing system based on cloud computing, using iterative region division and fusion, difference detection, correlation analysis and recurrent neural network prediction models, the problems of low data quality and computing efficiency in multi-source heterogeneous geographic information data processing are solved, and efficient and intelligent data fusion and prediction are achieved.

CN119829693BActive Publication Date: 2025-07-01SHANDONG INST OF GEOLOGICAL SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510307706.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-01
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

When integrating multi-source heterogeneous data, existing geographic information data processing technologies face problems such as uneven data quality, significant spatial heterogeneity and insufficient correlation mining, resulting in low data fusion accuracy and low computing efficiency.

Method used

A geographic information data processing system based on cloud computing is designed, including data processing module, division and fusion module, difference detection module, association analysis module, feature prediction module and data selection module. The system realizes the automated and intelligent fusion of multi-source heterogeneous data through iterative region division and fusion, differential detection, correlation analysis and prediction model based on recurrent neural networks.

Benefits of technology

It effectively improves the quality and processing efficiency of geographic information data, realizes the automation and intelligent integration of data, can accurately identify and locate data conflict areas, fully explore the intrinsic relationships of data, and is suitable for large-scale geographic information data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119829693B_ABST
    Figure CN119829693B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of geographic information data processing, and particularly to a geographic information data processing system based on cloud computing. The system includes a data processing module, a division and fusion module, a difference detection module, an association analysis module, a feature prediction module, and a data selection module. The data processing module preprocesses the received geographic information data; the division and fusion module performs iterative regional division and fusion; the difference detection module detects the difference regions; the association analysis module obtains the association relationships between the difference regions and other regions; the feature prediction module obtains the predicted geographic features of the difference regions; and the data selection module selects the final regional information data. The present invention identifies the data conflict regions through difference detection, excavates the complex associations inherent in the geographic information data, predicts the geographic features of the difference regions, and selects the most appropriate geographic information data as the final data based on similarity calculation, realizing the intelligence and automation of the data processing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geographic information data processing, and in particular to a geographic information data processing system based on cloud computing. Background Art

[0002] Geographic information data, also often referred to as spatial data, refers to data that describes various geographical elements on the earth's surface and their interrelationships. It is a type of data that contains location information and can be used to describe natural or human phenomena on the earth.

[0003] When integrating multi-source heterogeneous data, existing geographic information data processing technologies usually face problems such as uneven data quality, significant spatial heterogeneity, and insufficient mining of associations. Specifically, traditional methods often lack refined processing of geographic information data characteristics in the data preprocessing stage, resulting in low accuracy of subsequent analysis results; in the process of data fusion, simple and rough regional division and fusion strategies are difficult to ensure the integrity and accuracy of data, and are prone to introduce noise and errors; in view of the differences between different data sources in the same area, existing technologies often use simple superposition or averaging methods, ignoring the quality and reliability of data sources, and lack effective difference detection and processing mechanisms, and cannot accurately identify and locate data conflict areas; in addition, traditional methods do not adequately mine the complex associations inherent in geographic information data, resulting in the inability to fully release data value and limited prediction accuracy; more importantly, in the face of massive geographic information data, traditional methods have bottlenecks in computing efficiency and scalability, and are difficult to meet the needs of large-scale geographic information data processing. Therefore, how to effectively integrate multi-source heterogeneous geographic information data, improve data quality and fusion accuracy, fully mine the inherent associations of data, and improve data processing efficiency and scalability are key issues that need to be urgently solved in existing technologies. Summary of the invention

[0004] The present invention provides a geographic information data processing system based on cloud computing, which is used to solve the defects in the prior art.

[0005] The present invention provides a geographic information data processing system based on cloud computing, comprising:

[0006] The data processing module is used to receive geographic information data from different data sources and pre-process each geographic information data to obtain pre-processed data.

[0007] The division and fusion module is used to iteratively divide and fuse the pre-processed data to generate preliminary complete geographic information data.

[0008] A difference detection module is used to extract the geographical features of each region, detect the difference values of the geographical information data from different data sources in the same region based on the geographical features, set a difference threshold, and take the regions with difference values higher than the difference threshold as difference regions.

[0009] An association analysis module is used to obtain the association relationships between the difference regions and other regions according to the geographical features. The association relationships include spatial association degree, attribute association degree, and temporal association degree.

[0010] A feature prediction module is used to construct a prediction model based on a recurrent neural network, input the geographical features of the previous moment of the difference region, the region with the largest spatial association degree of the difference region, the region with the largest attribute association degree of the difference region, and the region with the largest temporal association degree of the difference region, and output the predicted geographical features of the difference region.

[0011] A data selection module is used to calculate the similarity between the geographical features of the difference region and the predicted geographical features, and take the information data of the geographical information data with the highest similarity between the difference region and the predicted geographical features in this region as the final regional information data.

[0012] According to a geographical information data processing system based on cloud computing provided by the present invention, the data processing module includes a multi-source data access unit, a data cleaning unit, and a data conversion unit. The multi-source data access unit is used to receive geographical information data from different data sources. The data cleaning unit is used to perform cleaning processing on the received geographical information data to obtain cleaned data. The cleaning processing includes missing value processing and outlier processing. The data conversion unit is used to convert the cleaned data into a unified data type as preprocessed data.

[0013] According to a geographical information data processing system based on cloud computing provided by the present invention, the process of iterative regional division and fusion of the preprocessed data includes:

[0014] Initial regional division: The preprocessed data is initially divided by using a grid division method to obtain N initial regions;

[0015] Regional data fusion: For each initial region, the Kalman filtering algorithm is used to fuse the data from different data sources to obtain a fusion result;

[0016] Fusion effect evaluation: Evaluate the fusion quality of the fusion result. The fusion quality includes data integrity and data consistency;

[0017] Repeat the initial regional division, regional data fusion, and fusion effect evaluation, and set an iteration stop criterion. The iteration stop criterion includes that the fusion quality reaches a preset threshold.

[0018] According to a geographic information data processing system based on cloud computing provided by the present invention, the difference detection module includes a geometric difference detection unit and an attribute difference detection unit. The geometric difference detection unit is used to detect the differences in geometric shapes of the same area in different data sources. The attribute difference detection unit is used to detect the differences in attribute values of the same area in different data sources.

[0019] According to a geographic information data processing system based on cloud computing provided by the present invention, the difference detection module further includes a spatial relationship feature extraction unit and a statistical feature extraction unit. The spatial relationship feature extraction unit is used to extract the spatial relationship features between each area and its adjacent areas. The statistical feature extraction unit is used to extract the statistical features of the attribute values within each area.

[0020] According to a geographic information data processing system based on cloud computing provided by the present invention, the spatial association degree includes proximity and connectivity. Proximity represents the physical distance between the difference area and other areas calculated through a spatial weight matrix. Connectivity represents the road connection between the difference area and other areas judged through graph theory algorithms.

[0021] According to a geographic information data processing system based on cloud computing provided by the present invention, the attribute association degree includes attribute similarity and correlation. Attribute similarity represents the degree of similarity between the difference area and other areas calculated through the Jaccard similarity coefficient. Correlation represents the attribute statistical dependence relationship between the difference area and other areas calculated through chi-square test and mutual information.

[0022] According to a geographic information data processing system based on cloud computing provided by the present invention, the time association degree includes time series similarity and causal relationship. Time series similarity represents the degree of similarity of the time series data between the difference area and other areas calculated through dynamic time warping and cross-correlation coefficient. Causal relationship represents the mutual influence relationship between the difference area and other areas analyzed through Granger causality test.

[0023] According to a geographic information data processing system based on cloud computing provided by the present invention, the process of constructing a prediction model based on a recurrent neural network includes:

[0024] Collect the historical moment geographic features of the difference area, the area with the largest spatial association degree of the difference area, the area with the largest attribute association degree of the difference area, and the area with the largest time association degree of the difference area, and label the next moment geographic features corresponding to the historical moment geographic features of the difference area.

[0025] Construct a basic recurrent neural network model, take the historical geographical features of the differential region, the region with the highest spatial correlation degree in the differential region, the region with the highest attribute correlation degree in the differential region, and the region with the highest temporal correlation degree in the differential region as inputs, and take the geographical features at the next moment corresponding to the historical geographical features of the differential region as outputs, and train the basic recurrent neural network model; retain the model parameters that meet the test accuracy to obtain a prediction model.

[0026] According to a geographic information data processing system based on cloud computing provided by the present invention, the process of calculating the similarity between the geographical features of the differential region and the predicted geographical features includes:

[0027] Decompose the geographical features of the differential region and the predicted geographical features into a numerical feature set and a categorical feature set respectively.

[0028] Calculate the numerical similarity of the numerical feature set in the geographical features of the differential region and the numerical feature set in the predicted geographical features through the Euclidean distance algorithm.

[0029] Calculate the classification similarity of the categorical feature set in the geographical features of the differential region and the categorical feature set in the predicted geographical features through the Jaccard similarity coefficient.

[0030] Adopt the weighted summation method to calculate the similarity between the geographical features of the differential region and the predicted geographical features.

[0031] A geographic information data processing system based on cloud computing provided by the present invention realizes the automatic and intelligent fusion of multi-source heterogeneous geographic information data through the coordinated action of each module, effectively improves the data quality and processing efficiency. It can be compatible with multiple data formats, breaks the data island effect, and realizes the unified access of multi-source heterogeneous data. Adopting an iterative regional division and fusion strategy can dynamically adjust the regional boundary according to the actual distribution of data, so as to achieve a more refined regional division. Through the difference detection module, the data conflict area can be accurately identified, and the complex internal associations of geographic information data can be mined through the association analysis module, providing a decision-making basis for data selection and fusion. By constructing a prediction model based on a recurrent neural network, the geographical features of the differential region can be effectively predicted, and the most suitable geographic information data can be selected based on the similarity calculation as the final data, realizing the intelligence and automation of the data processing process. Based on the cloud computing platform, it has good scalability and high concurrency processing capabilities, and can meet the needs of large-scale geographic information data processing. It has a higher degree of automation, higher processing accuracy, stronger intelligent decision-making ability and better scalability. Therefore, the system can be widely applied to fields such as urban planning, resource management, environmental protection, disaster management, transportation, etc., providing users with higher-quality geographic information services and providing strong support for scientific research and decision-making in related fields. Brief Description of the Drawings

[0032] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0033] Figure 1 It is a schematic structural diagram of a geographic information data processing system based on cloud computing provided by an embodiment of the present invention. Detailed Embodiments

[0034] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0035] The following will describe Figure 1 a geographic information data processing system based on cloud computing of the present invention.

[0036] Figure 1 It is a schematic structural diagram of a geographic information data processing system based on cloud computing provided by an embodiment of the present invention.

[0037] As Figure 1 shown, a geographic information data processing system based on cloud computing provided by an embodiment of the present invention includes a data processing module, a division and fusion module, a difference detection module, an association analysis module, a feature prediction module, and a data selection module.

[0038] The data processing module is used to receive geographic information data from different data sources and preprocess each piece of geographic information data to obtain preprocessed data.

[0039] The data processing module includes a multi-source data access unit, a data cleaning unit, and a data conversion unit. The multi-source data access unit is used to receive geographic information data from different data sources. The data cleaning unit is used to clean the received geographic information data to obtain cleaned data. The cleaning process includes missing value processing and outlier processing. The data conversion unit is used to convert the cleaned data into a unified data type as the preprocessed data.

[0040] In this embodiment, the multi-source data access unit needs to identify the channels from which the data comes, such as commercial geographic information system databases, remote sensing image data, crowd-sourced geographic data, sensor network data, etc. For different data sources, corresponding connection channels need to be established, and different technologies and protocols are used, including API interfaces, database connections, and file reading. Extract the required geographic information data from the data sources, including spatial range data, attribute field selection, and time range filtering. The data cleaning unit receives the geographic information data transmitted by the multi-source data access unit and processes the errors, incompleteness, and inconsistencies in the data to improve the data quality. For different missing situations, if the proportion of missing values is very small and has little impact on the analysis, the records containing missing values can be directly deleted. Or filling methods such as mean and median filling are used to fill the missing values. Statistical methods are used to detect whether there are data that significantly deviate from the normal range in the data as outliers. Statistical methods can include the standard deviation method and the box plot method. If it is confirmed that the outlier is incorrect data, it is directly deleted. If the correct value of the outlier can be determined, it is corrected. Filtering methods are used to remove unnecessary fluctuations in the data and highlight the main trends and features of the data. Check and remove data records that are completely duplicate or duplicate in key attributes to ensure the uniqueness of the data. Through standardization processing, the data format is unified to solve the inconsistencies caused by issues such as inconsistent units and naming conflicts in the data set. The data conversion unit converts the cleaned data into a unified data type and format for subsequent regional division and fusion, including data type conversion, coordinate system conversion, data format conversion, and data standardization.

[0041] By receiving the geographic information data from different data sources and performing preprocessing, the quality of subsequent data fusion and analysis is significantly improved. It can be compatible with multiple data formats and achieve the unified access of multi-source heterogeneous data. According to the characteristics of geographic information data, the preprocessing process can effectively clean the data, thereby improving the accuracy and reliability of the data. Through data conversion and standardization, data of different scales and units are unified under the same standard, laying a foundation for subsequent regional division and fusion.

[0042] The division and fusion module is used to perform iterative regional division and fusion on the preprocessed data to generate preliminary complete geographic information data.

[0043] The process of performing iterative regional division and fusion on the preprocessed data includes:

[0044] Initial regional division: The preprocessed data is initially divided using the grid division method to obtain N initial regions;

[0045] Regional data fusion: For each initial region, the Kalman filter algorithm is used to fuse the data from different data sources to obtain the fusion result;

[0046] Fusion effect evaluation: Evaluate the fusion quality of the fusion result, where the fusion quality includes data integrity and data consistency;

[0047] Repeat the initial area division, area data fusion, and fusion effect evaluation, and set the iteration stop criterion, where the iteration stop criterion includes that the fusion quality reaches a preset threshold.

[0048] In this embodiment, the goal of iterative area division and fusion is to continuously adjust the area division method and fuse data from different data sources, and finally obtain a geographic information dataset with complete data and high consistency. Divide the entire research area into regular grids, such as square or rectangular grids. Select the grid size according to the density and spatial distribution characteristics of the data. Obtain N initial areas, and each area contains geographic information data from different data sources. For each initial area, use the Kalman filter algorithm to fuse the data from different data sources to generate a unified dataset. The process includes:

[0049] Define the state variables to be estimated, including the height, position, and attribute values of the ground objects.

[0050] Establish an observation model to describe the relationship between the observed values of different data sources and the state variables.

[0051] Establish a process model to describe the variation law of the state variables over time.

[0052] Initialize the state variables and the error covariance matrix.

[0053] Use the Kalman filter algorithm for iterative update to continuously correct the estimated values of the state variables.

[0054] Based on the process model, predict the state variables at the next moment.

[0055] Compare the predicted values with the actual observed values, calculate the Kalman gain, and update the state variables and the error covariance matrix.

[0056] For each initial area, obtain a fused dataset that contains the optimal estimated state of the area.

[0057] Evaluate the quality of the fusion result to determine whether the expected fusion effect is achieved. Evaluate the missing situation of each attribute in the fused dataset, including the proportion of missing values and the geographical range covered by the data. Evaluate whether there are contradictions or conflicts between different attributes and between different data sources in the fused dataset, including attribute consistency: whether the attributes such as the type, name, and elevation of the same ground object are consistent. Spatial consistency: whether the geometric shape and position of the same ground object are consistent. Conduct statistical analysis on the fused data and calculate indicators such as the missing rate and consistency.

[0058] Set the thresholds for data integrity and data consistency. When the fusion quality reaches the threshold, stop the iteration. For data integrity, check if there are enough data points in each region, set a minimum data point threshold, and regions below this threshold are marked as having insufficient data integrity. For data consistency, calculate the standard deviation between different data sources within the same region. The smaller the standard deviation, the more consistent the data. Set a maximum standard deviation threshold, and regions exceeding this value need to be re-divided. According to the quality assessment results, for regions with low data integrity or poor data consistency, use a smaller grid for division to more precisely capture data features. Based on the quality assessment results of each iteration, set a feedback mechanism to determine whether to continue the iteration or adjust the grid size. If the fusion effect does not meet the iteration stop criteria, repeat the initial region division, regional data fusion, and fusion effect assessment. During each iteration, adjust the grid size or Kalman filter parameters to optimize the fusion effect.

[0059] Through iterative region division and fusion, the problem of difficult-to-determine region division granularity in traditional methods is solved, and the integrity and accuracy of data fusion are improved. Traditional static region division methods often struggle to adapt to the complex and variable characteristics of geographic information data, easily resulting in inaccurate region boundaries or excessive data heterogeneity within regions. The adopted iterative division and fusion strategy can dynamically adjust region boundaries according to the actual distribution of data, thereby achieving more refined region division. During the division process, this module comprehensively considers factors such as the spatial distribution and attribute characteristics of the data to ensure the homogeneity of data within regions. Effectively handle the stitching problem of region boundaries to ensure data continuity and integrity. Through iterative division and fusion, preliminary complete geographic information data can be generated, providing a reliable data basis for subsequent difference detection and correlation analysis, and also greatly enhancing the effects of data visualization and application analysis.

[0060] The difference detection module is used to extract the geographic features of each region and detect the difference values of the geographic information data from different data sources in the same region according to the geographic features. Set a difference threshold, and regions with difference values higher than the difference threshold are regarded as difference regions.

[0061] The difference detection module includes a geometric difference detection unit and an attribute difference detection unit. The geometric difference detection unit is used to detect the differences in geometric shapes of the same region in different data sources. The attribute difference detection unit is used to detect the differences in attribute values of the same region in different data sources. The difference detection module also includes a spatial relationship feature extraction unit and a statistical feature extraction unit. The spatial relationship feature extraction unit is used to extract the spatial relationship features between each region and its adjacent regions. The statistical feature extraction unit is used to extract the statistical features of the attribute values within each region.

[0062] In this embodiment, the process of the geometric difference detection unit detecting the geometric shape differences of the same area in different data sources includes: aligning the geometric data of different data sources to ensure they are in the same coordinate system. If the coordinate systems of the data sources are different, coordinate transformation is required. The Douglas-Peucker algorithm is used to reduce the number of geometric vertices and simplify the geometric data. Select geometric difference measurement methods, including: area difference, calculating the area difference of the same area in different data sources. Perimeter difference, calculating the perimeter difference of the same area in different data sources. Shape index, calculating the shape index difference of the same area in different data sources. The shape index includes: compactness index, describing the compactness of the area shape. Fractal dimension, describing the complexity of the area shape. Compare the calculated geometric difference value with a preset geometric difference threshold. If the difference value is higher than the threshold, it is considered that there is a geometric difference in this area, and it is marked as a difference area.

[0063] The process of the attribute difference detection unit detecting the attribute value differences of the same area in different data sources includes: matching the attribute names and semantic understanding to determine the attributes to be compared in different data sources. Convert the attribute values of different data sources into a unified data type. Select attribute difference measurement methods, including numerical attributes, categorical attributes, and text attributes.

[0064] Numerical attributes include: absolute difference, calculating the absolute difference of the same attribute in different data sources. Relative difference, calculating the relative difference of the same attribute in different data sources. Standardized difference, standardizing the attribute values and then calculating the difference. Categorical attributes include: consistency ratio, calculating the ratio of the same attribute having the same value in different data sources. Kappa coefficient, measuring the classification consistency of different data sources in categorical attributes. Text attributes include: edit distance, calculating the edit distance between two text strings, reflecting their similarity. Semantic similarity, using natural language processing technology to calculate the semantic similarity between two text strings. Compare the calculated attribute difference value with a preset attribute difference threshold. If the difference value is higher than the threshold, it is considered that there is an attribute difference in this area, and it is marked as a difference area.

[0065] Perform data analysis on the historical geographical information data of the target area, and use the distribution of difference values of each area in the historical data to calculate the mean and standard deviation of the difference values. Set the difference threshold to the mean plus K times the standard deviation, where K ranges from 1.5 to 2, to filter out normal fluctuations. During the processing, dynamically adjust the difference threshold according to the changes in real-time data. For example, if the difference values generally increase within a certain time period, appropriately increase the threshold.

[0066] The spatial relationship feature extraction unit extracts the spatial relationship features between each region and its adjacent regions, including: determining adjacent regions based on topological relationships, including adjacency, containment, and intersection. Extract the spatial relationship features between each region and its adjacent regions. The features include a spatial weight matrix and a spatial autocorrelation coefficient. The spatial weight matrix describes the strength of the spatial relationship between regions. It includes: an adjacency matrix, where the weight is 1 if two regions are adjacent and 0 otherwise. A distance weight matrix, where the weight is inversely proportional to the distance between regions. The spatial autocorrelation coefficient measures the degree of aggregation of regional attribute values in space.

[0067] The statistical feature extraction unit extracts the statistical features of the attribute values within each region, including: selecting the attributes for which statistical features need to be extracted and calculating the statistical features of the attribute values within each region. These include: the mean, median, standard deviation, variance, minimum value, maximum value of the attribute values, and quantiles of the attribute values, such as the 25% quantile and 75% quantile. And create a histogram of the attribute values, and extract the peak, skewness, and kurtosis features of the histogram to reflect the distribution of the attribute values.

[0068] By detecting the difference values of different data sources in the same region and extracting the geographical features of each region, data conflict regions are effectively identified and located, providing a decision-making basis for subsequent data selection and fusion. Due to differences in collection time, collection methods, data accuracy, etc., there are often data conflict phenomena in the geographical information data of different data sources. For example, the positions, attributes, or shapes of the same geographical feature are inconsistent in different data sources. Traditional simple overlay or averaging methods cannot effectively solve these data conflicts and are prone to lead to incorrect analysis results. However, this module can accurately identify the regions where the difference values are higher than the threshold by setting a difference threshold and conduct key analysis on these regions as difference regions. In addition, this module can also extract the geographical features of each region, such as terrain, landform, vegetation, water system, transportation, etc., providing basic data for subsequent correlation analysis. By comprehensively analyzing the geographical features and difference values of the difference regions, this module can effectively judge the reasons for data conflicts and provide a decision-making basis for subsequent data selection and fusion. For example, if the difference region is located in a high-incidence area of geological disasters, it may be necessary to preferentially select the latest data or data sources with higher reliability. This difference detection mechanism can effectively improve the accuracy of data fusion and provide users with more reliable geographical information data.

[0069] The correlation analysis module is used to obtain the correlation relationships between the difference regions and other regions based on geographical features. The correlation relationships include spatial correlation degree, attribute correlation degree, and temporal correlation degree.

[0070] Spatial correlation includes proximity and connectivity. Proximity represents the physical distance between the differential region and other regions calculated through a spatial weight matrix. Connectivity represents the road connection between the differential region and other regions judged by graph theory algorithms.

[0071] In this embodiment, spatial correlation analysis is used to determine the degree of spatial correlation between the differential region and other regions, including calculating proximity, calculating connectivity, and comprehensive spatial correlation. Among them, calculating proximity constructs a spatial weight matrix to describe the proximity relationship between regions. It includes: Adjacency weight matrix: If two regions are adjacent, the weight is 1, otherwise it is 0. Distance weight matrix: The weight is inversely proportional to the distance between regions. K-nearest neighbor weight matrix: For each region, select K nearest regions as its neighbors, the weight is 1, and the rest are 0. Based on the spatial weight matrix, use the Euclidean distance or Manhattan distance metric method to calculate the physical distance between the differential region and other regions. Use the reciprocal of the distance as the proximity to convert the physical distance into proximity.

[0072] The process of calculating connectivity includes: constructing the road data of the study area into a graph, where nodes represent road intersections and edges represent roads. Use graph theory algorithms to judge the road connection between the differential region and other regions, including: Shortest path algorithm, calculate the shortest path length between the differential region and other regions. Connected component analysis, judge whether the differential region and other regions belong to the same connected component.

[0073] Use the weighted average method to comprehensively combine proximity and connectivity to obtain the final spatial correlation.

[0074] Attribute correlation includes attribute similarity and correlation. Attribute similarity represents the degree of similarity between the differential region and other regions calculated through the Jaccard similarity coefficient. Correlation represents the attribute statistical dependence relationship between the differential region and other regions calculated through chi-square test and mutual information.

[0075] In this embodiment, attribute correlation analysis is used to determine the degree of attribute correlation between the differential region and other regions, including calculating attribute similarity, calculating correlation, and comprehensive attribute correlation. Calculating attribute similarity includes using the Jaccard similarity coefficient to represent the land use type of the region as a set and calculating the land use type similarity between two regions. The Jaccard similarity coefficient is defined as the size of the intersection of two sets divided by the size of the union. It can also include cosine similarity, representing the attributes of the region as a vector and using cosine similarity to calculate the attribute vector similarity between two regions.

[0076] Calculating correlation includes chi-square test to examine whether there is a significant association between two categorical variables. Mutual information measures the degree of mutual dependence between two variables. The greater the mutual information, the higher the degree of association between the two variables. Analyze the association between categorical and numerical attributes. At the same time, use the information gain ratio to correct the preference of mutual information for high-cardinality attributes.

[0077] Use the weighted average method to synthesize the attribute similarity and correlation to obtain the final attribute association degree.

[0078] Temporal association degree includes time series similarity and causal relationship. Time series similarity represents the degree of similarity of time series data between the difference region and other regions calculated through dynamic time warping and cross-correlation coefficient. Causal relationship represents the mutual influence relationship between the difference region and other regions analyzed through Granger causality test.

[0079] In this embodiment, temporal association degree analysis is used to determine the temporal association degree between the difference region and other regions, including calculating time series similarity, analyzing causal relationship, and synthesizing temporal association degree. Time series similarity represents the degree of similarity of time series data between the difference region and other regions calculated through dynamic time warping and cross-correlation coefficient. Dynamic time warping can measure time series similarity, allowing time series to bend on the time axis to find the best matching path, which is applicable to time series of different lengths. Cross-correlation coefficient can measure the correlation degree between two time series at different time lags and can be used to discover the lag relationship between time series.

[0080] Causal relationship represents the mutual influence relationship between the difference region and other regions analyzed through Granger causality test. Use Granger causality test to judge whether one time series can be used to predict another time series. If one time series X has Granger causality on another time series Y, it means that the past values of X can be used to predict the current value of Y.

[0081] Use the weighted average method to synthesize time series similarity and causal relationship to obtain the final temporal association degree.

[0082] By obtaining the spatial, attribute, and temporal correlation relationships between the differential region and other regions, the complex associations inherent in geographic information data are effectively mined, providing an important basis for model prediction and data selection. Geographic information data often exhibits complex spatial dependence, attribute correlation, and temporal evolution. For example, the topographies and landforms of adjacent regions tend to be similar, there may be mutual influences between different types of geographic elements, and the geographic elements in the same region may change at different times. Traditional analysis methods often overlook these internal associations, resulting in limitations in the accuracy and reliability of the analysis results. This module can accurately obtain these correlation relationships by comprehensively analyzing the spatial distance, attribute similarity, and temporal correlation between the differential region and other regions. For example, through spatial correlation analysis, the associations between the differential region and the topographies, landforms, vegetation cover, etc. of its surrounding regions can be discovered; through attribute correlation analysis, the associations between the differential region and the land use types, population densities, etc. of the surrounding areas can be discovered; through temporal correlation analysis, the evolution laws of the differential region at different times can be discovered. These correlation relationships can provide important input features for model prediction, improving the accuracy and reliability of prediction.

[0083] The feature prediction module is used to construct a prediction model based on a recurrent neural network. It inputs the previous moment's geographic features of the differential region, the region with the highest spatial correlation degree to the differential region, the region with the highest attribute correlation degree to the differential region, and the region with the highest temporal correlation degree to the differential region, and outputs the predicted geographic features of the differential region.

[0084] The process of constructing a prediction model based on a recurrent neural network includes:

[0085] Collect historical geographic feature data of the differential region. Geographic features include: land use type, vegetation coverage, population density, economic indicators, meteorological data, and POI data.

[0086] Collect historical geographic feature data of the region with the highest spatial correlation degree, the highest attribute correlation degree, and the highest temporal correlation degree to the differential region. And divide the collected data into a training set and a test set.

[0087] Set the time span and time granularity. The time span should be long enough to capture the long-term change trends of geographic features. The time granularity should be selected according to the data availability and prediction requirements.

[0088] Perform missing value processing, outlier processing, data standardization, and time series processing on the collected data.

[0089] Combine the historical geographical features of the differential region, the region with the highest spatial correlation degree of the differential region, the region with the highest attribute correlation degree of the differential region, and the region with the highest temporal correlation degree of the differential region into the input features of the model. Assuming the time granularity is monthly, the geographical feature data for the past 12 months can be used as the input of the model. Use the geographical features at the next moment corresponding to the historical geographical features of the differential region as the output label of the model. For example, if the input is the geographical feature data for the past 12 months, the output is the geographical feature data for the 13th month.

[0090] Design the structure of the basic recurrent neural network model, including an input layer to receive the input features. It contains multiple recurrent layers. The number of recurrent layers and the number of recurrent units in each layer are hyperparameters that need to be adjusted. The output layer outputs the prediction result.

[0091] Select the tanh activation function for the recurrent layer and the linear activation function for the output layer.

[0092] Train the basic recurrent neural network model on the training set, use the validation set to evaluate the performance of the basic recurrent neural network model during the training process, and use the test set to test the generalization ability of the trained basic recurrent neural network model.

[0093] Retain the model parameters that meet the test accuracy to obtain the prediction model.

[0094] Deploy the trained prediction model to the cloud computing platform, input the geographical features at the previous moment of the differential region, the region with the highest spatial correlation degree of the differential region, the region with the highest attribute correlation degree of the differential region, and the region with the highest temporal correlation degree of the differential region, and output the predicted geographical features of the differential region.

[0095] In this embodiment, by constructing a prediction model based on a recurrent neural network and inputting the geographical features of the previous moment of the differential region and its associated regions, the geographical features of the differential region can be effectively predicted, providing an important reference basis for the data selection module. The recurrent neural network is particularly good at processing time series data and can capture the laws of geographical feature changes over time. In the processing of geographical information data, this module makes full use of this advantage of the RNN and can effectively predict the geographical features of the differential region, such as land use type, vegetation coverage, population density, etc. By inputting the geographical features of the previous moment of the differential region and the region with the highest spatial, attribute, and time correlation, this model can comprehensively consider the influence of various factors, thereby improving the accuracy and reliability of the prediction. The prediction results can provide an important reference basis for the data selection module. For example, geographical information data with the highest similarity to the predicted geographical features can be selected as the final data for this region. This model prediction mechanism can effectively improve the rationality of data selection, reduce the need for manual intervention, and thus improve the automation degree and efficiency of data processing. At the same time, this model can also be adaptively adjusted according to the actual situation. For example, the prediction accuracy can be improved by adjusting the network structure, optimizing training parameters, etc.

[0096] The data selection module is used to calculate the similarity between the geographical features of the differential region and the predicted geographical features, and use the geographical information data of the region with the highest similarity between the differential region and the predicted geographical features as the final regional information data.

[0097] The process of calculating the similarity between the geographical features of the differential region and the predicted geographical features includes:

[0098] Decompose the geographical features of the differential region and the predicted geographical features into a numerical feature set and a categorical feature set respectively.

[0099] Calculate the numerical similarity of the numerical feature set in the geographical features of the differential region and the numerical feature set in the predicted geographical features through the Euclidean distance algorithm.

[0100] Calculate the classification similarity of the categorical feature set in the geographical features of the differential region and the categorical feature set in the predicted geographical features through the Jaccard similarity coefficient.

[0101] Adopt the weighted summation method to calculate the similarity between the geographical features of the differential region and the predicted geographical features.

[0102] Use the geographical information of the geographical information data with the highest similarity to the predicted geographical features in this differential region as the final data, so as to obtain the complete geographical information data of the target region with the highest accuracy.

[0103] In this embodiment, by calculating the similarity between the geographical features of the difference region and the predicted geographical features, and selecting the geographical information data with the highest similarity as the final regional information data, the intelligence and automation of the data processing process are realized, and the data quality and analysis accuracy are improved. Traditional geographical information data fusion often relies on manual judgment and selection, is easily affected by subjective factors, and has low efficiency. However, by introducing similarity calculation, this module can objectively evaluate the quality and reliability of different data sources, and select the data closest to the prediction result as the final data. This data selection strategy can effectively reduce the need for manual intervention and improve the degree of automation of data processing. In addition, this module can also select different similarity calculation methods according to different application scenarios. For example, Euclidean distance, cosine similarity or correlation coefficient can be selected to adapt to different data characteristics and analysis requirements. By selecting the most suitable geographical information data as the final data, this module can effectively improve the accuracy of data fusion and provide a more reliable data basis for subsequent analysis and applications. This intelligent data selection mechanism can effectively improve the efficiency and quality of data processing and provide users with better geographical information services.

[0104] In summary, this embodiment provides a geographical information data processing system based on cloud computing. Through the collaborative action of each module, the automatic and intelligent fusion of multi-source heterogeneous geographical information data is realized, effectively improving the data quality and processing efficiency. It can be compatible with a variety of data formats, breaking the data island effect and realizing the unified access of multi-source heterogeneous data. By adopting an iterative regional division and fusion strategy, it can dynamically adjust the regional boundary according to the actual distribution of data, so as to achieve a more refined regional division. Through the difference detection module, data conflict regions can be accurately identified, and the complex internal associations of geographical information data can be mined through the association analysis module, providing a decision-making basis for data selection and fusion. By constructing a prediction model based on a recurrent neural network, the geographical features of the difference region can be effectively predicted, and the most suitable geographical information data can be selected based on similarity calculation as the final data, realizing the intelligence and automation of the data processing process. Based on the cloud computing platform, it has good scalability and high concurrency processing capabilities, and can meet the needs of large-scale geographical information data processing. It has a higher degree of automation, higher processing accuracy, stronger intelligent decision-making ability and better scalability. Therefore, this system can be widely applied in fields such as urban planning, resource management, environmental protection, disaster management, and transportation, providing users with better geographical information services and providing strong support for scientific research and decision-making in related fields.

[0105] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A geographic information data processing system based on cloud computing, characterized in that: include: A data processing module, used for receiving geographic information data from different data sources, and preprocessing each of the geographic information data to obtain preprocessed data; A division and fusion module, used for iteratively dividing and fusion of the pre-processed data to generate preliminary complete geographic information data; A difference detection module is used to extract the geographical features of each area, and detect the difference values ​​of geographical information data from different data sources in the same area according to the geographical features, set a difference threshold, and regard the area with a difference value higher than the difference threshold as a difference area; An association analysis module, used for obtaining the association relationship between the difference area and other areas according to the geographical features, wherein the association relationship includes spatial association, attribute association and time association; A feature prediction module is used to construct a prediction model based on a recurrent neural network, input the geographical features of the difference area, the area with the maximum spatial correlation of the difference area, the area with the maximum attribute correlation of the difference area, and the area with the maximum temporal correlation of the difference area at the previous moment, and output the predicted geographical features of the difference area; The data selection module is used to calculate the similarity between the geographical features of the difference area and the predicted geographical features, and use the information data in the area where the geographical information data with the highest similarity between the difference area and the predicted geographical features is the final area information data.

2. A geographic information data processing system based on cloud computing according to claim 1, characterized in that: The data processing module includes a multi-source data access unit, a data cleaning unit and a data conversion unit; the multi-source data access unit is used to receive geographic information data from different data sources; the data cleaning unit is used to perform cleaning processing on the received geographic information data to obtain cleaned data, and the cleaning processing includes missing value processing and outlier processing; the data conversion unit is used to convert the cleaned data into a unified data type as preprocessing data.

3. The geographic information data processing system based on cloud computing according to claim 1, characterized in that: The process of iteratively dividing and fusing the preprocessed data includes: Initial region division: using a grid division method to perform initial division on the preprocessed data to obtain N initial regions; Regional data fusion: for each of the initial regions, a Kalman filter algorithm is used to fuse data from different data sources to obtain a fusion result; Fusion effect evaluation: evaluating the fusion quality of the fusion result, wherein the fusion quality includes data integrity and data consistency; The initial region division, regional data fusion and fusion effect evaluation are repeated, and an iteration stop criterion is set, wherein the iteration stop criterion includes that the fusion quality reaches a preset threshold.

4. The geographic information data processing system based on cloud computing according to claim 1, characterized in that: The difference detection module includes a geometric difference detection unit and an attribute difference detection unit; the geometric difference detection unit is used to detect the difference in geometric shapes of the same area in different data sources; the attribute difference detection unit is used to detect the difference in attribute values ​​of the same area in different data sources.

5. The geographic information data processing system based on cloud computing according to claim 1, characterized in that: The difference detection module also includes a spatial relationship feature extraction unit and a statistical feature extraction unit; The spatial relationship feature extraction unit is used to extract the spatial relationship features between each region and its adjacent regions; the statistical feature extraction unit is used to extract the statistical features of the internal attribute values ​​of each region.

6. A geographic information data processing system based on cloud computing according to claim 4, characterized in that: The spatial association includes proximity and connectivity; the proximity represents the physical distance between the difference area and other areas calculated by a spatial weight matrix; the connectivity represents the road connection between the difference area and other areas determined by a graph theory algorithm.

7. The geographic information data processing system based on cloud computing according to claim 1, characterized in that: The attribute association degree includes attribute similarity and correlation; the attribute similarity represents the similarity between the difference area and other areas calculated by the Jaccard similarity coefficient; the correlation represents the attribute statistical dependency relationship between the difference area and other areas calculated by the chi-square test and mutual information.

8. The geographic information data processing system based on cloud computing according to claim 1, characterized in that: The temporal correlation includes time series similarity and causal relationship; the time series similarity represents the similarity of the time series data between the difference area and other areas calculated by dynamic time warping and mutual correlation coefficient; the causal relationship represents the mutual influence relationship between the difference area and other areas analyzed by Granger causality test.

9. The geographic information data processing system based on cloud computing according to claim 1, characterized in that: The process of building a prediction model based on a recurrent neural network includes: Collect the geographical features of the difference area, the area with the largest spatial correlation in the difference area, the area with the largest attribute correlation in the difference area, and the area with the largest temporal correlation in the difference area at the historical moment, and mark the geographical features of the difference area at the next moment corresponding to the geographical features at the historical moment; A recurrent neural network basic model is constructed, and the geographical features of the difference area, the area with the maximum spatial correlation of the difference area, the area with the maximum attribute correlation of the difference area, and the area with the maximum time correlation of the difference area at historical moments are taken as input, and the geographical features of the difference area at the next moment corresponding to the geographical features at the historical moments are taken as output, and the recurrent neural network basic model is trained; the model parameters that meet the test accuracy are retained to obtain a prediction model.

10. The geographic information data processing system based on cloud computing according to claim 1, characterized in that: The process of calculating the similarity between the geographical features of the difference area and the predicted geographical features includes: Decomposing the geographical features of the difference area and the predicted geographical features into a numerical feature set and a categorical feature set respectively; Calculating the numerical similarity between the numerical feature set in the geographical features of the difference area and the numerical feature set in the predicted geographical features by using the Euclidean distance algorithm; Calculate the classification similarity between the classification feature set in the geographical features of the difference area and the classification feature set in the predicted geographical features by using the Jaccard similarity coefficient; The weighted sum method is used to calculate the similarity between the geographical features of the difference area and the predicted geographical features.

Citation Information

Patent Citations

  • Construction engineering quality data management method based on artificial intelligence

    CN119377209A

  • Geographic Dataset Preparation and Analytics Systems

    US20230036170A1