Geographic information data fusion system
Through the geographic information data fusion system, the Geo-SDNN model and semantic hierarchical fusion algorithm are used to solve the problem of inaccurate semantic relationship capture in traditional systems, and the efficient, accurate fusion and personalized visualization of geographic information data are achieved, which meets the needs of modern geographic information processing.
Patent Information
- Application Number
- CN202510435917.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional geographic information data fusion systems are difficult to accurately capture the complex semantic relationships between geographical entities, and cannot maintain high accuracy and consistency during the fusion process, and cannot meet the needs of modern geographic information data processing.
The geographic information data fusion system is adopted, including data acquisition and preprocessing units, deep learning semantic analysis model construction units, semantic-level data fusion units and data application and visualization units. The Geo-SDNN model is used for semantic relationship inference, combined with multi-source data acquisition and preprocessing technology, semantic feature extraction and relationship inference are performed through a bidirectional long and short-term memory network of multi-scale spatial convolutional neural network and attention mechanism, and semantic hierarchical fusion algorithm and hierarchical visual rendering strategy are used to achieve efficient and accurate fusion of data.
It significantly improves the efficiency and accuracy of geographic information data fusion, enhances the environmental adaptability and visualization effects of data fusion, meets users' personalized visualization needs, and provides solid data foundation support.
Smart Images

Figure CN120354350A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of geographic information data processing, and particularly to a geographic information data fusion system. Background Art
[0002] With the rapid development of geographic information technology, the demands for the acquisition, processing, and application of geographic information are increasing day by day. However, the sources of geographic information data are diverse, the formats are inconsistent, and there are often redundancies, contradictions, or inconsistencies among the data, which poses challenges to the effective utilization of geographic information data.
[0003] There are still some obvious drawbacks in traditional technologies. Traditional models often have difficulty accurately capturing the complex semantic relationships between geographic entities. In addition, high accuracy and consistency cannot be maintained during the fusion process.
[0004] In summary, the traditional geographic information data fusion system has obvious deficiencies in the construction unit of the deep learning semantic analysis model and the semantic-level data fusion unit, and it is difficult to meet the requirements of modern geographic information data processing. Therefore, it is particularly important to develop a geographic information data fusion system. Summary of the Invention
[0005] The purpose of the present invention is to make up for the deficiencies of the existing technology and provide a geographic information data fusion system. It can achieve the efficient and accurate fusion of geographic information data by introducing advanced deep learning models and semantic fusion algorithms, combined with multi-source data acquisition and preprocessing technologies, providing a solid data foundation for geographic information applications.
[0006] To solve the above technical problems, the present invention provides the following technical solution: A geographic information data fusion system, which includes the following components: a data acquisition and preprocessing unit, a deep learning semantic analysis model construction unit, a semantic-level data fusion unit, and a data application and visualization unit;
[0007] The data acquisition and preprocessing unit: A multi-source data acquisition interface that can establish connections with various geographical information data sources, including but not limited to high-resolution satellite image data sources, ground geographical sensor network data sources, and various geographical information databases. For satellite image data sources, through specific orbit prediction and data capture algorithms, based on satellite orbit parameters and the Earth's rotation model, the satellite overpass time and data reception window are accurately calculated. For ground geographical sensor network data sources, according to the sensor network topology structure and communication protocol, a distributed data acquisition strategy is adopted. The sensor network topology structure is determined through the analysis of the topography and geomorphology of the geographical monitoring area and the study of the distribution characteristics of monitoring targets. For geographical information databases, data extraction is carried out according to the database access interface specification and data indexing algorithm. The data indexing algorithm is constructed based on the classification system and spatial indexing technology of geographical data. The spatial indexing technology adopts a hybrid indexing structure combining quadtree and R-tree. The segmentation threshold and hierarchical structure of the quadtree and R-tree are determined through the analysis of the spatial distribution characteristics of a large amount of geographical data;
[0008] Using a noise filtering algorithm based on the analysis of the distribution characteristics and spatial autocorrelation of geographical data. First, calculate the Moran's I index of geographical data in the spatial dimension. The formula for Moran's I index is: where n is the number of geographical data points, x i and x j are the attribute values of different data points, is the mean of the attribute values, w ij is the spatial weight matrix, representing the spatial correlation degree between data points i and j. The spatial weight matrix is constructed based on the reciprocal relationship of the distances between geographical data points. The distance threshold d t is determined through the density analysis of the spatial distribution of geographical data. When Moran's I index is lower than the set threshold M1, it is determined that there is noise abnormality in the data and it is excluded. Then, data in different formats are uniformly converted into a structured data representation form based on geographical entity objects. Each geographical entity object includes a unique identifier, spatial geometric information, attribute information, and timestamp information. The spatial geometric information is represented by a sequence of geographical coordinate points and geometric shape descriptors. The attribute information covers various physical and human attributes of geographical entities. The timestamp information records the data acquisition or update time;
[0009] The deep learning semantic analysis model construction unit: The model architecture adopts the Geo-SDNN (Geographical Semantic Deep Neural Network). The input layer of the network receives the structured data of geographical entity objects output by the data collection and preprocessing unit. The first layer is the geographical entity encoding layer, which uses a hash encoding algorithm based on the types and attributes of geographical entities to map the type information and attribute values of geographical entities into hash codes of fixed length. The parameters of the hash function are determined through the analysis of the geographical entity classification system and the attribute value range. The second layer is the spatial semantic feature extraction layer, which uses a multi-scale spatial convolutional neural network containing multiple groups of convolutional kernels of different scales. The size of the convolutional kernels at each scale is determined according to the spatial resolution levels of geographical data. Let the lowest resolution of geographical data be R min , and the highest resolution be R max . A total of K resolution levels are divided, then the size of the convolutional kernels at the k-th scale is . Through convolutional operations of different scales, semantic feature vectors of geographical entities at different spatial scales are extracted. The third layer is the semantic relationship reasoning layer, which adopts the Att-BiLSTM (Attention-based Bidirectional Long Short-Term Memory Network). The weight calculation of the attention mechanism is based on the semantic similarity and spatial proximity between the semantic feature vectors of geographical entities. The semantic similarity is calculated using the cosine similarity. Let the semantic feature vectors be v i and v j , and their semantic similarity . The spatial proximity is calculated according to the spatial coordinate distance d ij of geographical entities . The attention weight A ij = αS ij +(1 - α)P ij , where α is the semantic similarity weight coefficient, which is determined through the analysis of the semantic relationships of a large number of geographical data samples. The Att-BiLSTM performs weighted processing on the semantic feature vectors according to the attention weights, and then conducts bidirectional sequence modeling to infer the semantic relationships between geographical entities, constructing a semantic relationship graph. The Geo-SDNN model is trained with a large number of sample data marked with the semantic relationships of geographical entities. The marked data comes from geographical expert knowledge, the in-depth analysis of historical geographical evolution cases, and the collation of field geographical survey data;
[0010] The semantic-level data fusion unit: The semantic relationship mapping module, based on the semantic relationship graph output by the Geo-SDNN model, maps geographical information data into the semantic space and constructs a semantic connection matrix. For the determination of semantic connections, a semantic association intensity threshold T1 and a spatial influence range threshold R s are set. The semantic association intensity is represented by the edge weights in the semantic relationship graph, and the edge weights are determined by the semantic relationship scores output by the Att-BiLSTM network. When the semantic association intensity is higher than T1 and the spatial distance between geographical entities is less than R sWhen establishing a semantic connection, the matrix element values represent the semantic connection strength;
[0011] For the strong connection data in the semantic connection matrix, a semantic hierarchy fusion algorithm is adopted. First, the geographical data is divided into a basic layer, an intermediate layer, and a target layer according to the semantic hierarchy. The division basis is the functional hierarchy and data abstraction degree of the geographical data in the geographical information system. Then, the fusion order is determined according to the semantic hierarchy structure, and the fusion starts from the basic layer and goes upward. The fusion formula is: Where F i is the fusion result of the i-th layer, D i is the original data of the i-th layer, w i is the semantic weight of the i-th layer data. The semantic weight is determined according to the importance of the geographical data in a specific analysis task, and a fused geographical information data set is generated;
[0012] The data application and visualization unit: constructs a general application interface that follows the geographical information data interaction standard, adopts an asynchronous data transmission mechanism based on a message queue. The capacity Q of the message queue c is determined according to the data traffic peak and system processing capacity. The capacity size is determined through statistical analysis of historical data traffic and system performance testing, enabling the fused geographical information data to be efficiently docked with third-party application systems such as geographical analysis software, urban planning systems, and traffic management systems. The interface protocol supports multiple data transmission formats;
[0013] According to the semantic category information and semantic connection strength in the semantic connection matrix, a hierarchical visualization rendering strategy is adopted. For semantic categories, unique colors, symbols, and texture styles are defined for each category. For semantic connection strength, it is represented by the line thickness and transparency. The higher the connection strength, the thicker the line and the lower the transparency. The adjustment range r of the line thickness t is determined according to visualization effect testing and user visual perception research. The transparency is calculated based on the normalization processing of the connection strength value, mapping the connection strength value to the [0, 1] interval as the transparency value to intuitively display the semantic relationship and interaction strength between geographical entities.
[0014] Furthermore, the multi-source data acquisition interface of the data acquisition and preprocessing unit also includes a data quality assessment module. When collecting data, the data quality is evaluated based on the accuracy, integrity, and timeliness of the data. The accuracy assessment uses a method of comparing with known standard geographical data. For terrain data, it is compared with high-precision terrain measurement data, and the mean elevation error μ e and standard deviation σ e are calculated. When μ e exceeds the set threshold E1 or σ eWhen it exceeds E2, it is determined that the data accuracy does not meet the standard. The integrity assessment passes by checking whether the data lacks key geographical attribute information. If the land use data lacks the land type attribute, it is determined that the data is incomplete. The timeliness assessment is based on the time difference Δt between the data collection time and the current time, combined with the data update cycle T u , when Δt>T u , it is determined that the data timeliness is insufficient. According to the quality assessment results, the data is marked and screened. High-quality data is preferentially collected to improve the quality of the basic data for data fusion.
[0015] Furthermore, in the training sample optimization module in the deep learning semantic analysis model construction unit, sample screening and synthesis techniques are adopted. Sample screening is based on the principles of data diversity and representativeness, and the spatial distribution entropy H of the sample data is calculated s and the attribute value distribution entropy H a . The calculation formula for the spatial distribution entropy is: where p i is the probability of the sample data in the i-th spatial region. The attribute value distribution entropy is similar. The entropy value is determined by analyzing the spatial and attribute distributions of the sample data, and the samples with entropy values higher than the set threshold H min are retained to ensure the diversity of the samples. Sample synthesis uses a method based on geographical data transformation rules to perform small-range random translation and rotation operations on the spatial positions of geographical entities. The translation ranges Δx and Δy are determined according to the spatial scale and accuracy of the geographical data, and the rotation angle range Δθ is determined according to the geometric shape characteristics of the geographical entities, which is obtained through the analysis of the geometric shapes of a large number of geographical entities. At the same time, linear interpolation transformation is performed on the attribute values to generate new training samples to expand the number of training samples and improve the generalization ability of the Geo-SDNN model.
[0016] Furthermore, when constructing the semantic connection matrix in the semantic relationship mapping module in the semantic-level data fusion unit, a geographical environment dynamic factor correction term is introduced. In earthquake-active regions, for geographical entities related to geological structures, according to the earthquake activity frequency f and magnitude m in the earthquake monitoring data, a correction coefficient C f =βf + γm is constructed, where β and γ are coefficients determined by analyzing the influence degree of earthquakes on the semantic relationships of geographical entities, which is obtained through the study of the relationship between historical earthquake data and geographical entity changes, to correct the semantic connection strength and make the fusion result more conform to the dynamic changes of the geographical environment, enhancing the environmental adaptability of data fusion.
[0017] Furthermore, in the fusion strategy formulation and execution module of the semantic-level data fusion unit, in the semantic-level fusion algorithm, the adaptive adjustment mechanism of semantic weights is based on the real-time monitoring feedback of geographical data. When it is detected that the land use type in a certain area changes rapidly and farmland is converted into construction land, the semantic weight w of land use data in the urban expansion analysis task lu increases by Δw lu . The adjustment amplitude Δw lu =δr lu , where r lu is the land use change rate, and δ is a coefficient determined based on the urban expansion model and the analysis of the impact of land use change. It is obtained through the study of a large number of urban development cases and land use change data to reflect the impact of geographical data changes on data fusion in real time and ensure the timeliness and adaptability of the fused data.
[0018] Furthermore, when the visualization module of the data application and visualization unit renders geographical entities, in addition to relying on the semantic connection matrix, it also performs dynamic visualization adjustment according to user interaction instructions. Users can specify the geographical area or geographical entity type of interest through the interaction interface. The system adjusts the visualization range and display details according to the user's instructions. When the user selects to focus on the business district of a certain city, the system focuses the visualization range on the business district and increases the detailed display of geographical entities inside the business district. The extraction and display rules of these detailed information are determined based on the storage structure and data association relationship of business geographical information data and are obtained through in-depth analysis and sorting of business geographical information data to meet the personalized visualization needs of users and assist users in accurate geographical analysis.
[0019] Furthermore, the system also includes a data storage and management unit, which adopts a distributed storage architecture and a data caching mechanism. The distributed storage architecture is based on the regional division and data type division of geographical data. Geographical data is stored in different storage node clusters. The regional division is based on the geographical region division standard of the earth, and the data type division includes image data, vector data, and attribute data. Each storage node cluster adopts a redundant storage strategy, and the redundancy r d is determined according to the importance and volatility of the data. The redundancy size is determined through the value evaluation of geographical data and the analysis of data loss risk to improve the reliability of data storage. The data caching mechanism establishes a data cache area in the memory. The size of the cache area is determined according to the system running memory capacity and data access frequency and is determined through system performance testing and analysis of historical data access patterns. It caches the data that is frequently accessed recently to improve data access speed and ensure the efficiency of data storage and management and the smoothness of data flow.
[0020] Furthermore, the system also has a data fusion process monitoring unit, which monitors the fusion process based on the status monitoring of data processing flow nodes and data flow tracking. Status monitoring indicators are set for each process node in data acquisition, preprocessing, model construction, and data fusion. The acquisition rate r of the data acquisition node a , the data volume q d , the processing time t of the preprocessing node p , the data conversion success rate s c . The thresholds of the status monitoring indicators are determined according to system performance requirements and historical data processing experience. When the monitoring indicator of a certain node exceeds the threshold, a warning message is issued. At the same time, through data flow tracking technology, the transfer path and processing time of data between each process node are recorded. Data flow tracking uses blockchain-based distributed ledger technology, and each data processing operation is recorded in the blockchain ledger, including operation time, operation content, and operator information, so as to trace back the data processing process and troubleshoot problems when problems occur, ensuring the quality and reliability of the fused data and providing a solid data foundation for geographic information applications.
[0021] Furthermore, the Geo-SDNN model in the deep learning semantic analysis model construction unit adopts model fusion and parameter optimization strategies during the training process. Model fusion uses the ensemble learning method to perform weighted combination on multiple trained Geo-SDNN models. The weight w m is determined according to the performance evaluation results of each model on the validation set. The performance evaluation indicators include accuracy, recall rate, and F1 value. The weight is determined by analyzing the performance of different models in the geographical semantic relationship inference task. Parameter optimization uses a global optimization method based on genetic algorithms. The parameters of the Geo-SDNN model are encoded as chromosomes, and the fitness function is defined as the loss function value of the model on the training set. Through the selection, crossover, and mutation operations of genetic algorithms, the optimal parameter combination is found. The population size p of the genetic algorithm s , the crossover probability c p , and the mutation probability m p parameters are determined according to the model parameter search space and optimization difficulty. The parameter values are determined through the analysis of the model parameter space and multiple optimization experiments to improve the accuracy and stability of the model and ensure the high-performance performance of the Geo-SDNN model in complex geographic information data processing.
[0022] Compared with the prior art, the geographical information data fusion system has the following beneficial effects:
[0023] 1. The system constructs a deep learning semantic analysis model unit, uses the Geo-SDNN model to infer semantic relationships of geographical entities, and can accurately capture complex semantic relationships between geographical entities. At the same time, the semantic-level data fusion unit adopts a fusion algorithm based on the semantic hierarchy to ensure high accuracy and consistency of data during the fusion process. In addition, the multi-source data acquisition interface and data quality assessment module of the data acquisition and preprocessing unit effectively improve the quality of basic data, further enhancing the accuracy of data fusion. The overall system design optimizes the data processing flow, significantly improving the efficiency of geographical information data fusion.
[0024] 2. By introducing a correction term for dynamic geographical environment factors, the system can correct the semantic connection strength according to changes in the actual geographical environment, making the fusion result more in line with the actual environmental dynamics and enhancing the environmental adaptability of data fusion. At the same time, the data application and visualization unit adopts a hierarchical visualization rendering strategy, dynamically adjusting the visualization according to the semantic connection matrix and user interaction instructions, intuitively displaying the semantic relationships and interaction strengths between geographical entities, meeting the personalized visualization needs of users, and assisting users in more accurate geographical analysis. This enhanced visualization effect provides more intuitive and rich data support for geographical information applications.
[0025] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be learned from the practice of the present invention. Brief Description of the Drawings
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0027] Figure 1 It is a flowchart operation diagram of a geographical information data fusion system. Detailed Embodiments
[0028] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention objective, the following, in combination with the accompanying drawings and preferred embodiments, details the specific embodiments, structures, features, and their effects of the present invention as follows.
[0029] Embodiment 1
[0030] This embodiment describes the need to integrate multi-source geographic information data in a planning project for a large city to provide decision support for the reasonable layout, functional zoning and infrastructure construction of the city.
[0031] Through the multi-source data acquisition interface of the data acquisition and preprocessing unit, the high-resolution satellite image data source is connected to obtain urban topographic information, and the satellite overpass time and data receiving window are accurately calculated based on the satellite orbit parameters and the earth's rotation model to ensure the integrity of the image data.
[0032] Collect traffic flow and environmental monitoring real-time data from ground geographic sensor network data sources. According to the sensor network topology and communication protocol, adopt distributed data collection strategy to efficiently collect scattered sensor data.
[0033] Extract historical data on land use and population distribution from the geographic information database, and use the data indexing algorithm built based on the geographic data classification system and spatial indexing technology to quickly and accurately obtain the required data.
[0034] Using the noise filtering algorithm, the Moran index of geographic data is calculated, and the formula is: Where n is the number of geographic data points, x i and x j are the attribute values of different data points, is the mean value of the attribute, w ij It is a spatial weight matrix, which indicates the degree of spatial correlation between data points i and j. Noise and abnormal data are eliminated according to the spatial weight matrix, and data in different formats are uniformly converted into structured data, including the unique identification, spatial geometry information, attribute information and timestamp information of geographic entities, such as the coordinates, height, purpose and construction time information of buildings.
[0035] The Geo-SDNN model is constructed. The first layer is the geographic entity encoding layer. A hash encoding algorithm based on the type and attribute of geographic entities is used to map the type information and attribute values of geographic entities into a hash code of fixed length. The parameters of the hash function are determined by analyzing the classification system of geographic entities and the attribute value domain. The second layer is the spatial semantic feature extraction layer. A multi-scale spatial convolutional neural network is used, which contains a convolution kernel group of multiple scales. The size of the convolution kernel of each scale is determined according to the spatial resolution level of geographic data. The minimum resolution of geographic data is set to R min , the highest resolution is R max , which is divided into K resolution levels, then the size of the k-th scale convolution kernel is Extract the semantic feature vectors of geographical entities at different spatial scales through convolutional operations of different scales. The third layer is the semantic relationship reasoning layer, which uses a bidirectional long short-term memory network with attention mechanism, Att-BiLSTM. The weight calculation of the attention mechanism is based on the semantic similarity and spatial proximity between the semantic feature vectors of geographical entities. The semantic similarity is calculated using cosine similarity. Let the semantic feature vectors v i and v j , and their semantic similarity The spatial proximity is calculated according to the spatial coordinate distance d ij of geographical entities Calculate the spatial proximity, and the attention weight A ij =αS ij +(1 - α)P ij , where α is the semantic similarity weight coefficient, which is determined by analyzing the semantic relationships of a large number of geographical data samples. Att-BiLSTM performs weighted processing on the semantic feature vectors according to the attention weights, and then conducts bidirectional sequence modeling to infer the semantic relationships between geographical entities, construct a semantic relationship graph, and train the Geo-SDNN model through a large number of sample data labeled with the semantic relationships of geographical entities. The labeled data comes from geographical expert knowledge, in-depth analysis of historical geographical evolution cases, and collation of field geographical survey data. The input layer receives the preprocessed structured data of geographical entity objects. The geographical entity encoding layer uses a hash encoding algorithm to map to hash codes according to the geographical entity types (such as roads, parks) and attributes (such as length, area).
[0036] The spatial semantic feature extraction layer uses a multi-scale spatial convolutional neural network to determine the convolutional kernel size according to different resolution levels of urban geographical data, and extracts the semantic feature vectors of roads at different scales, such as the connectivity features of main roads at the macroscopic scale and the number of lanes features at the microscopic scale.
[0037] The semantic relationship reasoning layer uses the Att-BiLSTM network to calculate the semantic similarity and spatial proximity based on the attention mechanism, determine the weights, infer the semantic relationships between geographical entities, such as analyzing the service relationship between a park and the surrounding residential areas, construct a semantic relationship graph, and train the model through sample data labeled with the semantic relationships of geographical entities (such as the reasonable layout relationship between a park and a residential area in a historical urban planning case).
[0038] Based on the semantic relationship graph output by the Geo-SDNN model, the semantic relationship mapping module sets the semantic association strength threshold and the spatial influence range threshold, maps the geographical information data to the semantic space to construct a semantic connection matrix, and determines the semantic connections between geographical entities such as parks and surrounding roads and residential areas. The matrix element values represent the connection strength.
[0039] For strongly connected data, use the semantic hierarchical fusion algorithm, and the formula is: where F i is the fusion result of the i-th layer, D i is the original data of the i-th layer, w i is the semantic weight of the data of the i-th layer. The geographical data is divided into a basic layer (such as terrain data), an intermediate layer (such as land use data), and a target layer (such as urban functional zoning planning data) according to the functional hierarchy and abstraction level. The fusion is performed upward from the basic layer according to the fusion order, and the semantic weight is determined according to the importance of the geographical data in the urban planning task, such as the importance of terrain data for the site selection of infrastructure construction, to generate a fused geographical information dataset.
[0040] The data application and visualization unit constructs a general application interface and adopts an asynchronous data transmission mechanism to efficiently dock the fused geographical information data with the urban planning system, support multiple data transmission formats, meet the data interaction requirements between systems. According to the semantic connection matrix, a hierarchical visualization rendering strategy is adopted to define unique styles for different semantic categories (such as commercial areas, residential areas), and the semantic connection strength is represented by the line thickness and transparency, intuitively showing the relationships and interaction strengths between urban geographical entities, and assisting planners in making decisions.
[0041] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or variations equivalent to the equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modifications, equivalent variations, and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A geographic information data fusion system, characterized in that, The system includes the following components: a data acquisition and preprocessing unit, a deep learning semantic analysis model construction unit, a semantic-level data fusion unit, and a data application and visualization unit; The data acquisition and preprocessing unit: a multi-source data acquisition interface that can establish connections with various geographic information data sources, including but not limited to high-resolution satellite image data sources, ground geographic sensor network data sources, and various geographic information databases. For satellite image data sources, through specific orbit prediction and data capture algorithms, based on satellite orbit parameters and the Earth's rotation model, the satellite overpass time and data reception window are accurately calculated. For ground geographic sensor network data sources, according to the sensor network topology structure and communication protocol, a distributed data acquisition strategy is adopted. The sensor network topology structure is determined through terrain and geomorphology analysis of the geographic monitoring area and research on the distribution characteristics of monitoring targets. The geographic information database extracts data according to the database access interface specification and data indexing algorithm. The data indexing algorithm is constructed based on the classification system and spatial indexing technology of geographic data. The spatial indexing technology adopts a hybrid indexing structure combining quadtree and R-tree, and the segmentation threshold and hierarchical structure of the quadtree and R-tree are determined through analysis of the spatial distribution characteristics of a large amount of geographic data; Using a noise filtering algorithm based on the analysis of the distribution characteristics and spatial autocorrelation of geographical data, first, calculate the Moran's Index of geographical data in the spatial dimension. The formula for Moran's Index is as follows: where n is the number of geographical data points, x i and x j are the attribute values of different data points, is the mean of the attribute values, w ij is the spatial weight matrix, representing the spatial correlation degree between data points i and j. The spatial weight matrix is constructed based on the reciprocal relationship of the distances between geographical data points. The distance threshold d t is determined by the density analysis of the spatial distribution of geographical data. When Moran's Index is lower than the set threshold M1, it is determined that there is noise abnormality in the data and the data is excluded. Then, data in different formats are uniformly converted into a structured data representation form based on geographical entity objects. Each geographical entity object contains a unique identifier, spatial geometric information, attribute information, and timestamp information. The spatial geometric information is represented by a sequence of geographical coordinate points and geometric shape descriptors. The attribute information covers various physical and human attributes of geographical entities. The timestamp information records the data collection or update time; The deep learning semantic analysis model construction unit: The model architecture adopts the Geo-Spatial Deep Neural Network (Geo-SDNN). The input layer of the network receives the structured data of geographical entity objects output by the data collection and preprocessing unit. The first layer is the geographical entity encoding layer, which uses a hash encoding algorithm based on the types and attributes of geographical entities to map the type information and attribute values of geographical entities into hash codes of fixed length. The parameters of the hash function are determined by analyzing the geographical entity classification system and the attribute value range. The second layer is the spatial semantic feature extraction layer, which uses a multi-scale spatial convolutional neural network containing multiple sets of convolutional kernels at different scales. The size of the convolutional kernels at each scale is determined according to the spatial resolution levels of the geographical data. Let the lowest resolution of the geographical data be R min , and the highest resolution be R max . There are a total of K resolution levels. Then the size of the convolutional kernels at the k-th scale is . The semantic feature vectors of geographical entities at different spatial scales are extracted through convolutional operations at different scales. The third layer is the semantic relationship reasoning layer, which adopts the Attention-based Bidirectional Long Short-Term Memory Network (Att-BiLSTM). The weight calculation of the attention mechanism is based on the semantic similarity and spatial proximity between the semantic feature vectors of geographical entities. The semantic similarity is calculated using the cosine similarity. Let the semantic feature vectors be v i and v j . Their semantic similarity . The spatial proximity is calculated according to the spatial coordinate distance d ij of the geographical entities . The attention weight A ij = αS ij + (1 - α)P ij , where α is the semantic similarity weight coefficient, which is determined by analyzing the semantic relationships of a large number of geographical data samples. Att-BiLSTM performs weighted processing on the semantic feature vectors according to the attention weights and then conducts bidirectional sequence modeling to infer the semantic relationships between geographical entities, constructing a semantic relationship graph. The Geo-SDNN model is trained with a large number of sample data labeled with the semantic relationships of geographical entities. The labeled data comes from geographical expert knowledge, in-depth analysis of historical geographical evolution cases, and collation of on-site geographical survey data; The semantic-level data fusion unit: The semantic relationship mapping module maps the geographical information data to the semantic space based on the semantic relationship graph output by the Geo-SDNN model, constructs a semantic connection matrix, and sets a semantic association intensity threshold T1 and a spatial influence range threshold R for the determination of semantic connections. s , the semantic association intensity is represented by the edge weight in the semantic relationship graph, and the edge weight is determined by the semantic relationship score output by the Att-BiLSTM network. When the semantic association intensity is higher than T1 and the spatial distance between geographical entities is less than R s , a semantic connection is established, and the matrix element value represents the semantic connection intensity. For the strong connection data in the semantic connection matrix, a semantic hierarchical fusion algorithm is adopted. First, the geographical data is divided into a basic layer, an intermediate layer, and a target layer according to the semantic hierarchy. The division basis is the functional level and data abstraction degree of the geographical data in the geographical information system. Then, the fusion order is determined according to the semantic hierarchical structure, and the fusion starts from the basic layer and goes upward. The fusion formula is: Where F i is the fusion result of the i-th layer, D i is the original data of the i-th layer, w i is the semantic weight of the i-th layer data. The semantic weight is determined according to the importance of the geographical data in a specific analysis task, and a fused geographical information data set is generated; The data application and visualization unit: constructs a general application interface that follows the geographical information data interaction standard, adopts an asynchronous data transmission mechanism based on a message queue, and the capacity Q of the message queue c is determined according to the data traffic peak and system processing capacity, and the capacity size is determined through statistical analysis of historical data traffic and system performance testing, so that the fused geographical information data can be efficiently docked with third-party application systems such as geographical analysis software, urban planning systems, and traffic management systems, and the interface protocol supports multiple data transmission formats; According to the semantic category information and semantic connection strength in the semantic connection matrix, a hierarchical visualization rendering strategy is adopted. For semantic categories, unique colors, symbols, and texture styles are defined for each category. For semantic connection strength, it is represented by the line thickness and transparency. The higher the connection strength, the thicker the line and the lower the transparency. The adjustment range r of the line thickness t Determined according to the visualization effect test and user visual perception research, the calculation of transparency is based on the normalization of the connection strength value. The connection strength value is mapped to the interval [0, 1] as the transparency value to intuitively display the semantic relationship and interaction strength between geographical entities.
2. The geographic information data fusion system according to claim 1, characterized in that The multi-source data acquisition interface of the data acquisition and preprocessing unit further includes a data quality assessment module. During data acquisition, the quality of the data is evaluated based on its accuracy, integrity, and timeliness. For accuracy assessment, a method of comparing with known standard geographical data is used. For terrain data, it is compared with high-precision terrain measurement data, and the mean elevation error μ e and the standard deviation σ e are calculated. When μ e exceeds the set threshold E1 or σ e exceeds E2, it is determined that the data accuracy does not meet the standard. For integrity assessment, it is checked whether key geographical attribute information is missing from the data. If the land use type attribute is missing from the land use data, it is determined that the data is incomplete. For timeliness assessment, based on the time difference Δt between the data acquisition time and the current time, combined with the data update cycle T u , when Δt>T u , it is determined that the data timeliness is insufficient. Based on the quality assessment results, the data is marked and screened, and high-quality data is preferentially acquired.
3. A geographic information data fusion system according to claim 1, characterized in that, In the training sample optimization module of the deep learning semantic analysis model construction unit, sample screening and synthesis techniques are adopted. Sample screening is based on the principles of data diversity and representativeness, and the spatial distribution entropy H of the sample data is calculated. s And the attribute value distribution entropy H a . The calculation formula for the spatial distribution entropy is: Where p i is the probability of the sample data in the i-th spatial region. The attribute value distribution entropy is similar. The entropy value is determined by analyzing the spatial and attribute distributions of the sample data, and the samples with entropy values higher than the set threshold H min are retained. For sample synthesis, a method based on geographical data transformation rules is adopted. Small-range random translation and rotation operations are performed on the spatial positions of geographical entities. The translation ranges Δx and Δy are determined according to the spatial scale and accuracy of the geographical data, and the rotation angle range Δθ is determined according to the geometric shape characteristics of the geographical entities.
4. A geographic information data fusion system according to claim 1, characterized in that When constructing the semantic connection matrix, the semantic relationship mapping module in the semantic-level data fusion unit introduces a correction term for dynamic geographical environment factors. In seismically active regions, for geographical entities related to geological structures, a correction coefficient C is constructed based on the earthquake activity frequency f and magnitude m in the earthquake monitoring data. f = βf + γm, where β and γ are coefficients determined based on the analysis of the impact of earthquakes on the semantic relationships of geographical entities.
5. A geographic information data fusion system according to claim 1, characterized in that, In the fusion strategy formulation and execution module in the semantic-level data fusion unit, in the semantic-level fusion algorithm, the adaptive adjustment mechanism of semantic weights is based on the real-time monitoring feedback of geographical data. When it is detected that the land use type in a certain area changes rapidly and farmland is converted into construction land, the semantic weight w of the land use data in the urban expansion analysis task lu increases by Δw lu . The adjustment amplitude Δw lu =δr lu , where r lu is the land use change rate, and δ is a coefficient determined according to the urban expansion model and the analysis of the impact of land use change.
6. The geographic information data fusion system according to claim 1, wherein When the visualization module of the data application and visualization unit renders geographic entities, in addition to relying on the semantic connection matrix, it also performs dynamic visualization adjustment according to user interaction instructions. Users can specify the geographic area or type of geographic entity to be concerned through the interaction interface. The system adjusts the visualization range and display details according to the user's instructions. When the user selects to focus on the business district of a certain city, the system focuses the visualization range on the business district and increases the detail display of the geographic entities inside the business district. The extraction and display rules of these detail information are determined according to the storage structure and data association relationship of commercial geographic information data.
7. A geographic information data fusion system according to claim 1, characterized in that, The system further includes a data storage and management unit, which adopts a distributed storage architecture and a data caching mechanism. The distributed storage architecture is based on the regional division and data type division of geographical data. Geographical data is stored in different storage node clusters. The regional division is based on the geographical region division standard of the earth, and the data type division includes image data, vector data, and attribute data. Each storage node cluster adopts a redundant storage strategy, and the redundancy degree r d is determined according to the importance and volatility of the data. The redundancy degree is determined by evaluating the value of geographical data and analyzing the risk of data loss. The data caching mechanism establishes a data cache area in the memory, and the size of the cache area is determined according to the system running memory capacity and data access frequency.
8. A geographic information data fusion system according to claim 1, characterized in that, The system also has a data fusion process monitoring unit, which monitors the fusion process based on the status monitoring of data processing flow nodes and the tracking of data flow directions. Status monitoring indicators are set for each process node in data acquisition, preprocessing, model construction, and data fusion. The acquisition rate r a and the data volume q d of the data acquisition node, the processing time t p and the data conversion success rate s c of the preprocessing node. The thresholds of the status monitoring indicators are determined according to system performance requirements and historical data processing experience. When the monitoring indicator of a certain node exceeds the threshold, a warning message is issued. At the same time, through the data flow tracking technology, the transfer path and processing time of data between each process node are recorded. The data flow tracking adopts the distributed ledger technology based on blockchain, and each data processing operation is recorded in the blockchain ledger, including the operation time, operation content, and operator information.
9. A geographic information data fusion system according to claim 1, characterized in that, In the training process of the Geo-SDNN model in the deep learning semantic analysis model construction unit, a model fusion and parameter optimization strategy is adopted. The model fusion uses the ensemble learning method to perform weighted combination on multiple trained Geo-SDNN models, and the weight w m is determined according to the performance evaluation results of each model on the validation set. The performance evaluation metrics include accuracy, recall rate, and F1 value. The weight is determined by analyzing the performance of different models in the geographical semantic relationship inference task. The parameter optimization uses a global optimization method based on the genetic algorithm. The parameters of the Geo-SDNN model are encoded as chromosomes, and the fitness function is defined as the loss function value of the model on the training set. Through the selection, crossover, and mutation operations of the genetic algorithm, the optimal parameter combination is searched. The population size p s of the genetic algorithm, the crossover probability c p , and the mutation probability m p are determined according to the model parameter search space and optimization difficulty.
Citation Information
Cited By
Commercial opportunity recognition method and system based on space and semantic collaboration
CN120849718A
Historical place name space-time reconstruction method and device based on geographical context
CN121278638A
Visualization method and system based on land improvement and ecological restoration data
CN121747114A