Comprehensive traffic infrastructure spatial database construction method based on multi-source data
Through multi-source data fusion, library fragmentation, main and replica synchronization and elastic expansion mechanisms, various shortcomings of the existing transportation infrastructure database are solved, efficient management and retrieval are achieved, and the construction of smart transportation systems is supported.
Patent Information
- Application Number
- CN202510921817.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-07-04
AI Technical Summary
The existing transportation infrastructure database has obvious shortcomings in terms of heterogeneous data sources, inconsistent data standards, separate storage of spatial data and attribute data, inefficient retrieval efficiency, load imbalance, insufficient expansion capabilities, insufficient dynamic data management, etc., and cannot meet the needs of large-scale, multi-type, real-time updates of comprehensive transportation infrastructure databases.
Multi-source data fusion, library fragmentation, main and replica synchronization, unified search and elastic expansion mechanisms are adopted, and a relational database, document-based database and spatial database technology is combined to establish a full-text index and spatial index system, formulate data dictionary and metadata standards, and build a comprehensive transportation infrastructure spatial database, supporting multi-conditional semantic retrieval and spatial visual display.
It realizes efficient integration and management of multi-source data, improves retrieval efficiency and expansion capabilities, ensures the stability and availability of the system, and supports the construction and application of smart transportation systems.
Smart Images

Figure CN120429376A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic infrastructure data management, and particularly to a method for constructing an integrated traffic infrastructure spatial database based on multi-source data. Background Art
[0002] With the continuous acceleration of the urbanization process, the traffic infrastructure system has become increasingly complex, covering various transportation modes such as highways, waterways, and aviation. To effectively support the planning, construction, operation, and management of urban traffic, it is necessary to efficiently, accurately collect, manage, and apply traffic infrastructure data. Traditional traffic infrastructure data management systems mostly adopt a single data source and single data type storage method, usually based on relational databases such as Oracle or MySQL, and only store the structured attribute data of the infrastructure. This method can meet the basic needs in the early stage when the data volume is small and the data type is single, but with the exponential growth of the data scale and the continuous enrichment of data types, the traditional management method has exposed many limitations.
[0003] Currently, most traffic infrastructure databases face problems such as heterogeneous data sources, inconsistent data standards, separate storage of spatial data and attribute data, and low retrieval efficiency. Traffic infrastructure not only includes traditional attribute table data, but also involves a large amount of spatial vector data, satellite image data, and dynamic traffic flow data. Existing systems often cannot achieve unified management of structured data and unstructured data, resulting in serious data island phenomena, and it is difficult for data to be shared and linked between different departments and different subsystems. In addition, the existing systems have limited processing capabilities for spatial data, usually only supporting simple coordinate storage and static queries, lacking an efficient spatial index mechanism and dynamic expansion capabilities, and it is difficult to meet the management requirements of large-scale and highly real-time traffic data.
[0004] In terms of spatial data management, traditional traffic infrastructure databases generally have problems such as low spatial retrieval performance and load imbalance caused by uneven distribution of spatial data. Most systems adopt single-machine deployment when the amount of spatial data is small, and spatial retrieval depends on traditional B-tree or simple R-tree index structures. As the scale of spatial data expands, the retrieval latency of the system increases significantly, and the load balancing and expansion capabilities cannot meet the business requirements. At the same time, there is a lack of effective design for the sharding management and replica synchronization mechanisms for spatial data, resulting in difficulty in ensuring data consistency in a high-concurrency environment, and there are major potential risks in the overall stability and availability of the system.
[0005] At the retrieval level, existing traffic data systems mostly rely on single relational database indexes or simple full-text search modules, lacking composite search mechanisms designed for the multi-source data characteristics of transportation infrastructure. Traditional methods are particularly incapable of efficient joint retrieval when dealing with mixed queries of spatial and textual data. Existing systems lack optimized designs for the collaborative work of full-text and spatial indexes, resulting in limited retrieval performance and an inability to support rapid queries under complex combinations of conditions. Furthermore, faced with the continued growth of data volumes, traditional retrieval systems lack scalability, and their replica expansion and sharding management mechanisms are simplistically designed. This concentrates query pressure on a small number of nodes, leading to unstable query response times and limited system throughput.
[0006] At the computing level, with the increasing complexity of traffic data processing, the single-server model for data query, retrieval, and computation is no longer sufficient. Existing transportation infrastructure database systems generally lack sharded cluster-based computing scheduling mechanisms, and their inter-node load balancing and task allocation strategies are poorly designed. This can easily lead to node computing bottlenecks in high-concurrency scenarios, resulting in overall performance degradation. This is particularly true for large-scale spatial queries, joint retrieval, and dynamic data updates, as traditional systems lack elastic scalability, and processing capacity struggles to scale linearly with data volume.
[0007] In terms of data standardization, current transportation infrastructure data management systems often lack unified data dictionaries and metadata standards. Data formats, field definitions, and value standards vary across departments and systems, leading to difficulties in data integration and poor interface interoperability, severely impacting data sharing and application efficiency. Furthermore, an incomplete metadata management system, fragmented dataset descriptions, and missing or inconsistently defined meta-attributes such as spatial reference systems, timestamps, and access permissions further exacerbate data isolation issues between systems.
[0008] Given the dynamic nature of transportation infrastructure, current systems lack adequate management and update mechanisms for dynamic data such as traffic flow, air traffic flow, and aviation business volume. Traditional systems often rely on periodic manual batch updates, lacking real-time dynamic updates and version control systems. This results in poor data timeliness and an inability to promptly reflect changes in traffic status, impacting the data's value in decision-making. Furthermore, the lack of flexible version control and change traceability mechanisms makes version conflicts and data consistency issues more likely to arise when simultaneously updating multi-source data.
[0009] At the retrieval application level, traditional traffic infrastructure data management systems usually only support simple keyword retrieval or basic spatial range queries, lacking a semantic retrieval mechanism for multi-condition combinations of keywords, spatial ranges, timestamps, attribute tags, etc. This results in low retrieval efficiency and poor result relevance for users under complex query requirements. Especially when it comes to comprehensive retrieval across data types and data levels, existing systems are difficult to meet the requirements for in-depth and fine-grained analysis of traffic data, severely restricting the actual effects of application scenarios such as traffic planning, intelligent traffic dispatching, and emergency command.
[0010] In summary, existing traffic infrastructure data management systems have obvious deficiencies in aspects such as multi-source data fusion, spatial data sharding management, replica synchronization mechanisms, retrieval layer index optimization, computing layer elastic expansion, data standardization and unification, and dynamic data management, and cannot meet the construction requirements of large-scale, multi-type, and real-time updated comprehensive traffic infrastructure databases. Therefore, there is an urgent need for a new method for constructing a comprehensive traffic infrastructure spatial database that can achieve multi-source heterogeneous data fusion management, efficient spatial data sharding and retrieval, coordinated expansion of the storage, retrieval, and computing layers, standardized data governance, and dynamic data synchronization and update, so as to effectively improve the management efficiency, retrieval performance, and system scalability of traffic infrastructure data and support the development needs of future intelligent transportation systems.
[0011] Therefore, how to provide a method for constructing a comprehensive traffic infrastructure spatial database based on multi-source data is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0012] An object of the present invention is to propose a method for constructing a comprehensive traffic infrastructure spatial database based on multi-source data. The present invention fully integrates relational database, document database, and spatial database technologies, combines full-text indexing and spatial indexing systems, designs a database sharding, master-replica synchronization, unified retrieval, and elastic expansion mechanism, and details a spatial database construction plan for efficiently managing various traffic infrastructure data such as highways, waterways, and aviation, with the advantages of high data fusion degree, high retrieval efficiency, strong expansion ability, and high management standardization degree.
[0013] According to the method for constructing a comprehensive traffic infrastructure spatial database based on multi-source data of an embodiment of the present invention, the following steps are included: S1. Collect multi-source data of three types of traffic infrastructure, namely highways, waterways, and aviation; S2. Establish a data dictionary and metadata standard for the collected multi-source data according to industry specifications, and formulate a database sharding strategy and unified interface specification; S3. Perform standardization processing on infrastructure attribute data, satellite image data, and spatial vector data in the multi-source data; S4. Develop a spatial data sharding strategy based on geographical region division, construct a full-text index and a spatial index system, and optimize storage and query in a replication cluster environment; S5. Construct a comprehensive transportation infrastructure spatial database, and use MySQL, PostGIS, MongoDB, ElasticSearch, and K-V storage for collaborative management; S6. Design an extension mechanism at the storage layer, retrieval layer, and computing layer to adapt to the growth of transportation infrastructure data volume and changes in retrieval requirements; S7. Establish a unified retrieval platform based on the comprehensive transportation infrastructure spatial database, integrate raster image data and spatial vector data, and support spatial visualization display and semantic retrieval.
[0014] Optionally, the multi-source data includes infrastructure attribute data, traffic flow data, navigation flow data, air traffic volume data, origin-destination distribution data, administrative division data, economic data, population data, satellite image data, and spatial vector data.
[0015] Optionally, the standardization process includes: Perform field standardization on infrastructure attribute data; Slice satellite image data using the tile pyramid method and manage metadata in combination with the full-text index and the MongoDB document model; Standardize and express spatial vector data in GeoJSON format and store it in the PostGIS spatial database.
[0016] Optionally, the replication cluster environment includes sharding storage of spatial data, full-text index, and metadata, configuring a primary replica and at least one secondary replica for each data shard, using a replica synchronization mechanism to maintain data consistency, performing load balancing queries based on the replica set, and having the secondary replica take over the role of the primary replica in case of node failure.
[0017] Optionally, the specific steps of S2 include: S21. According to the national standards and industry technical specifications in the field of geographic information, establish a data dictionary for the collected multi-source data. The data dictionary defines the data field name, field data type, field value range, and field unit type; S22. Based on the attribute information of the multi-source data, establish a metadata standard. The metadata standard includes data set name, data source, spatial reference system, data update time, and data access permission fields; S23. Develop a sub-database strategy to store infrastructure attribute data, traffic flow data, navigation flow data, air traffic volume data, origin-destination distribution data, administrative division data, economic data, and population data in a MySQL relational database, satellite image data in a MongoDB document database, and spatial vector data in a PostGIS spatial database; S24. Develop a unified interface specification, which includes data access interfaces, data retrieval interfaces, and data exchange format types.
[0018] Optionally, the specific steps of S3 are as follows: S31. Perform field standardization processing on infrastructure attribute data to unify data field names, field data types, field value ranges, and field unit types; S32. Perform tile pyramid hierarchical slicing processing on satellite image data: ; Among them, is the number of slicing levels, is the original image spatial resolution, is the target tile slicing resolution, represents the ceiling operation; S33. Generate a slicing service for the sliced satellite image data through Geoserver, establish a full-text index in combination with ElasticSearch, and manage the metadata of image data based on the MongoDB document model; S34. Standardize the representation of spatial vector data in GeoJSON format, store it in the PostGIS spatial database, and establish a spatial index based on the R-tree structure and an auxiliary index based on the hash structure for spatial vector features to optimize the performance of spatial data retrieval.
[0019] Optionally, the specific steps of S4 are as follows: S41. Divide spatial data based on geographical regions and develop a sharding strategy: ; Among them, is the spatial data sharding number, and are the longitude and latitude coordinates of the spatial data, and are the minimum longitude and latitude of the spatial region respectively, and are the spatial sharding division intervals respectively, is the number of shards in the longitude direction, represents the floor operation; S42. Within the spatial data shard, establish an R-tree-based spatial index for the spatial elements, record the minimum bounding rectangle parameters of the spatial objects, and use a hash-assisted index to accelerate attribute retrieval; S43. Establish a full-text index for the metadata associated with the spatial data using the ElasticSearch indexing engine. The index fields include the dataset identifier, spatial range description, timestamp, and key information tags. S44. In a replication cluster environment, spatial data shards and full-text index replicas are distributed and stored, and query optimization is performed based on the replica node load status: ; in, For the The query load of replica nodes, For the The query request frequency of each node, For the The performance weight of each node, For the The number of replicas maintained by a node.
[0020] Optionally, the S5 specifically includes: S51. Use a MySQL relational database to store structured data, including infrastructure attribute data, traffic flow data, air traffic flow data, aviation business volume data, travel origin-destination distribution data, administrative division data, economic data, and population data; S52. Use PostGIS spatial database to manage spatial vector data and associated attributes, establish spatial index based on R-tree structure, and use auxiliary hash index to accelerate attribute retrieval; S53. Use MongoDB document-based database to manage metadata, including dataset description information, spatial reference system parameters, data update time, and data access permission identifier; S54. Use the ElasticSearch search engine to manage full-text index data, where the full-text index data includes a dataset identifier, a spatial range description, a time tag, and a description of key attributes; S55. Use the KV storage system to manage frequently accessed cache data, including query result cache, index metadata cache, and replica synchronization status cache. Calculate the cache space requirements: ; in, is the total cache space required, For the The space occupied by a single item in the class cache, For the Class cache access frequency ratio, is the total number of cache categories.
[0021] Optionally, S6 specifically includes: S61. The storage layer consists of MySQL, MongoDB, and PostGIS. A data expansion mechanism is designed in the storage layer, and the data capacity is expanded using the partition and replica expansion methods: ; Among them, is the total data capacity, is the number of partitions, is the storage capacity of a single partition, is the replica redundancy ratio; S62. The retrieval layer consists of ElasticSearch. An index expansion mechanism is designed in the retrieval layer, and the retrieval load is expanded based on the full-text index sharding and replica sharding strategies: ; Among them, is the number of index replicas, is the current query request rate, is the maximum query rate supported by a single replica, represents the ceiling operation; S63. The computing layer consists of a sharded cluster. A task expansion mechanism is designed in the computing layer, and the task load sharding and node concurrency expansion strategies are adopted: ; Among them, is the load level of the th computing node, is the amount of tasks assigned to the th node, is the computing performance coefficient of the th node.
[0022] Optionally, S7 specifically includes: S71. Based on the integrated transportation infrastructure spatial database, build a unified retrieval platform, and the platform accesses structured data, spatial data, metadata, and full-text index data; S72. Integrate raster image data and spatial vector data to build a multi-level spatial visualization model, and support the overlay rendering of spatial elements and raster layers; S73. Establish a semantic retrieval mechanism based on full-text index, and support the combined query of keywords, spatial range, timestamp, and attribute tags; S74. According to the retrieval results, dynamically load spatial vector elements and corresponding raster image layers for visual display and view update.
[0023] The beneficial effects of the present invention are as follows: By integrating multi-source data of three types of transportation infrastructure, namely roads, waterways, and aviation, the present invention establishes a unified data dictionary and metadata standard, and formulates standardized data access and management specifications, solving the problems of chaotic data formats, inconsistent standards, and difficult sharing in existing systems, and realizing data fusion and interconnection between transportation subsystems. By classifying and storing different types of data, and using MySQL relational database, PostGIS spatial database, MongoDB document database, and ElasticSearch full-text search engine for collaborative management, a hybrid storage architecture integrating structured and unstructured data is constructed, significantly improving the flexibility of data storage and management efficiency.
[0024] The present invention proposes a spatial data sharding strategy based on geographical area division, constructs a multi-level retrieval system by combining spatial index and full-text index, and realizes distributed storage and load balancing optimization of spatial data and index copies in a replication cluster environment, solving the problems of spatial data retrieval performance bottleneck and storage load imbalance in existing technologies, and greatly improving the query efficiency and system stability of large-scale spatial data. At the same time, the present invention designs elastic expansion mechanisms at the storage layer, retrieval layer, and computing layer respectively, supporting dynamic expansion of data capacity, number of replicas, and number of computing nodes, ensuring that the system still has good availability and scalability when facing massive data growth and access peaks.
[0025] In view of the dynamic change characteristics of transportation infrastructure, the present invention establishes a dynamic update mechanism and version control system for traffic flow, navigation flow, and air traffic volume data, supporting data incremental update and historical version traceability, and improving the timeliness and consistency of data management. In addition, through the construction of a unified retrieval platform, the present invention realizes multi-condition semantic retrieval based on keywords, spatial range, timestamp, and attribute tags, and conducts spatial visualization display by integrating raster image data and spatial vector data, breaking through the problems of single retrieval method and slow response in traditional systems, and improving the comprehensive retrieval and analysis ability of transportation infrastructure data.
[0026] In summary, through multi-source heterogeneous data fusion, standardized governance, spatial data sharding optimization, three-layer expansion mechanism, and dynamic update system design, the present invention systematically solves the deficiencies of existing transportation infrastructure databases in terms of data standardization, storage and retrieval performance, system scalability, and dynamic management ability, significantly improving the management efficiency, retrieval performance, and application flexibility of comprehensive transportation infrastructure data, and providing strong data support for the construction and application of intelligent transportation systems. Brief Description of the Drawings
[0027] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings: Figure 1 is a flowchart of a method for constructing a spatial database of integrated transportation infrastructure based on multi-source data proposed by the present invention; Figure 2 is a schematic diagram of a spatial data sharding strategy for a method for constructing a spatial database of integrated transportation infrastructure based on multi-source data proposed by the present invention; Figure 3 is a schematic diagram of the functional structure of a unified retrieval platform for a method for constructing a spatial database of integrated transportation infrastructure based on multi-source data proposed by the present invention. Detailed implementation manners
[0028] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0029] Refer to Figures 1-3 , a method for constructing a spatial database of integrated transportation infrastructure based on multi-source data, includes the following steps: S1. Collect multi-source data of three types of transportation infrastructure, namely highways, waterways, and aviation; S2. According to industry specifications, establish a data dictionary and metadata standard for the collected multi-source data, and formulate a sub-database strategy and a unified interface specification; S3. Standardize the infrastructure attribute data, satellite image data, and spatial vector data in the multi-source data; S4. Formulate a spatial data sharding strategy based on geographical region division, construct a full-text index and a spatial index system, and optimize storage and query in a replication cluster environment; S5. Construct a spatial database of integrated transportation infrastructure, and jointly manage it using MySQL, PostGIS, MongoDB, ElasticSearch, and K-V storage; S6. Design an extension mechanism at the storage layer, retrieval layer, and computing layer to adapt to the growth of transportation infrastructure data volume and the change of retrieval requirements; S7. Based on the spatial database of integrated transportation infrastructure, establish a unified retrieval platform, integrate raster image data and spatial vector data, and support spatial visualization display and semantic retrieval.
[0030] The present invention collects multi-source data of three types of transportation infrastructure, namely highway, waterway, and aviation, constructs a unified and integrated data foundation, realizes the standardized access and comprehensive management of heterogeneous data between different transportation subsystems, and significantly improves the integration degree of transportation infrastructure data and the ability to support the construction of intelligent transportation systems.
[0031] In this embodiment, the multi-source data includes infrastructure attribute data, traffic flow data, navigation flow data, aviation business volume data, origin-destination distribution data, administrative division data, economic data, population data, satellite image data, and spatial vector data.
[0032] The present invention clarifies the types and scopes of multi-source data, covers various static and dynamic data required for the operation and management of transportation infrastructure, ensures the comprehensiveness and adaptability of the data in the integrated transportation database, and improves the integrity and practicality of the unified integration of multi-source heterogeneous data.
[0033] In this embodiment, the standardization processing includes: Performing field standardization processing on the infrastructure attribute data; Slicing the satellite image data in the form of a tile pyramid and managing the metadata in combination with full-text indexing and the MongoDB document model; Standardizing and expressing the spatial vector data in the GeoJSON format and storing it in the PostGIS spatial database.
[0034] The present invention performs standardization processing on the infrastructure attribute data, satellite image data, and spatial vector data, eliminates the differences in data structures and inconsistent descriptions, and improves the compatibility, sharing ability, and sustainable management ability of multi-source data in a unified database environment.
[0035] In this embodiment, the replication cluster environment includes sharding the storage of spatial data, full-text indexing, and metadata, configuring a primary replica and at least one secondary replica for each data shard, adopting a replica synchronization mechanism to maintain data consistency, performing load-balanced queries based on the replica set, and having the secondary replica take over the role of the primary replica in case of node failure.
[0036] The present invention realizes high availability and load-balanced queries by introducing the sharding storage of spatial data and index replicas and the primary-secondary synchronization mechanism in the replication cluster environment, and enhances the stability, reliability, and fault tolerance of the system in the management of large-scale transportation spatial data.
[0037] In this embodiment, the specific steps of S2 include: S21. According to the national standards and industry technical specifications in the field of geographic information, establish a data dictionary for the collected multi-source data. The data dictionary defines the data field names, field data types, field value ranges, and field unit types. S22. Establish a metadata standard based on multi-source data attribute information. The metadata standard includes data set name, data source, spatial reference system, data update time, and data access permission fields. S23. Develop a sub-library strategy to store infrastructure attribute data, traffic flow data, navigation flow data, air traffic volume data, origin-destination distribution data, administrative division data, economic data, and population data in a MySQL relational database, satellite image data in a MongoDB document database, and spatial vector data in a PostGIS spatial database. S24. Develop a unified interface specification, which includes data access interfaces, data retrieval interfaces, and data exchange format types.
[0038] By establishing a data dictionary and a metadata standard, developing a sub-library strategy and a unified interface specification, the present invention effectively standardizes the traffic infrastructure data management system, improves data access consistency, system interoperability, and the flexibility of multi-source data expansion, and supports the standardized construction of the traffic information system.
[0039] In this embodiment, S3 specifically includes: S31. Perform field standardization processing on infrastructure attribute data to unify data field names, field data types, field value ranges, and field unit types. S32. Perform tile pyramid hierarchical slicing processing on satellite image data: ; where is the number of slicing levels, is the original image spatial resolution, is the target tile slicing resolution, represents the ceiling operation; S33. Generate a slicing service for the sliced satellite image data through Geoserver, establish a full-text index in combination with ElasticSearch, and manage the meta-information of the image data based on the MongoDB document model. S34. Standardize the representation of spatial vector data in GeoJSON format, store it in the PostGIS spatial database, and establish a spatial index based on the R-tree structure and an auxiliary index based on the hash structure for spatial vector features to optimize the spatial data retrieval performance.
[0040] By standardizing the attribute data field processing, tile slicing and full-text index construction of satellite image data, standardizing the storage of spatial vector data, and optimizing the index, the present invention systematically improves the retrieval efficiency, storage optimization degree, and data access convenience of traffic infrastructure spatial data.
[0041] In this embodiment, step S4 specifically includes: S41. Divide the spatial data based on geographical regions and formulate a slicing strategy: ; Among them, is the slicing number of the spatial data, is associated with is the longitude and latitude coordinates of the spatial data, [[ID=1,6]]is associated with are respectively the minimum longitude and latitude of the spatial region, is associated with are respectively the slicing intervals of the spatial slicing, is the number of slices in the longitude direction, represents the floor operation; S42. Within the spatial data slices, establish a spatial index based on the R-tree for spatial elements, record the minimum bounding rectangle parameters of spatial objects, and use a hash-assisted index to accelerate attribute retrieval; S43. Establish a full-text index for the metadata associated with the spatial data, use the ElasticSearch index engine, and the index fields include dataset identification, spatial range description, timestamp, and key information tags; S44. In a replicated cluster environment, distribute and store the spatial data slices and full-text index replicas, and optimize the query based on the load status of replica nodes: ; Among them, is the query load of the th replica node, is the query request frequency of the th node, is the performance weight of the th node, is the number of replicas maintained by the th node.
[0042] Through the design of the spatial data slicing strategy based on geographical region division, combined with spatial index and full-text index to optimize the query load, the present invention improves the retrieval performance, load balancing, and system scalability of traffic spatial data in a distributed environment, and supports high-concurrency and large-scale traffic data application scenarios.
[0043] In this embodiment, step S5 specifically includes: S51. Use a MySQL relational database to store structured data, including infrastructure attribute data, traffic flow data, navigation flow data, air traffic volume data, origin-destination distribution data, administrative division data, economic data, and population data; S52. Manage spatial vector data and associated attributes using a PostGIS spatial database, establish a spatial index based on the R-tree structure, and use an auxiliary hash index to accelerate attribute retrieval; S53. Manage metadata using a MongoDB document database. The metadata includes dataset description information, spatial reference system parameters, data update time, and data access permission identifiers; S54. Manage full-text index data using an ElasticSearch search engine. The full-text index data includes dataset identifiers, spatial range descriptions, time tags, and key attribute description content; S55. Manage high-frequency access cache data using a K-V storage system. The cache data includes query result cache, index metadata cache, and replica synchronization status cache, and calculate the cache space requirements: ; where is the total cache requirement space, is the space occupied by the th type of cache item, is the access frequency ratio of the th type of cache, is the total number of cache categories.
[0044] By classifying different types of traffic data according to the characteristics of structured data, spatial data, metadata, and index data, and using multiple databases to collaboratively build a fusion storage system, the present invention significantly improves the storage efficiency, query performance, and system response speed of traffic infrastructure data.
[0045] In this embodiment, the S6 specifically includes: The storage layer consists of MySQL, MongoDB, and PostGIS. Design a data expansion mechanism in the storage layer and use partitioning and replica expansion methods to expand the data capacity: ; where is the total data capacity, is the number of partitions, is the storage capacity of a single partition, is the replica redundancy ratio; The retrieval layer consists of ElasticSearch. Design an index expansion mechanism in the retrieval layer and expand the retrieval load based on the full-text index sharding and replica sharding strategies: ; where is the number of index replicas, is the current query request rate, is the maximum query rate supported by single-copy, represents the ceiling operation; S63. The computing layer consists of a sharding cluster. A task expansion mechanism is designed in the computing layer, and a task load sharding and node concurrent expansion strategy is adopted: ; Among them, is the load level of the th computing node, is the amount of tasks allocated to the th node, is the computing performance coefficient of the th node.
[0046] Through designing expansion mechanisms in the storage layer, retrieval layer and computing layer respectively, and adopting partition expansion, replica synchronization and distributed computing load balancing strategies, the present invention ensures the high availability and high scalability of the transportation infrastructure database under the conditions of continuous growth of data volume and changing access pressure.
[0047] In this embodiment, the S7 specifically includes: S71. Based on the integrated transportation infrastructure spatial database, build a unified retrieval platform, and the platform accesses structured data, spatial data, metadata and full-text index data; S72. Integrate raster image data and spatial vector data, construct a multi-level spatial visualization model, and support the overlay rendering of spatial elements and raster layers; S73. Establish a semantic retrieval mechanism based on full-text index, and support the combined query of keywords, spatial range, timestamp and attribute tags; S74. According to the retrieval results, dynamically load spatial vector elements and corresponding raster image layers for visual display and view update.
[0048] By integrating raster and vector data display through a unified retrieval platform and establishing a semantic retrieval mechanism for the combination of keywords, spatial range, timestamp and attribute tags, the present invention improves the intelligent retrieval ability and spatial visualization analysis ability of integrated transportation data, and enhances the application breadth and depth of the integrated transportation information system.
[0049] Example 1: To verify the feasibility of the method for constructing an integrated transportation infrastructure spatial database based on multi-source data proposed by the present invention, the present invention is applied to the transportation informatization improvement project of the Anhui Provincial Comprehensive Transportation Management Center. The project background is to integrate the transportation infrastructure data of highways, waterways and aviation within the province, build a unified integrated transportation infrastructure database, and realize the interconnection, visual display and efficient retrieval of transportation infrastructure resources across the province.
[0050] In practical applications, traffic data from various business departments is first accessed through a data collection platform. The sources include the road attribute table and traffic flow records provided by the Highway Bureau, the channel cross-section map and port distribution data provided by the Waterway Bureau, and the airport facilities and flight network layout provided by the Civil Aviation Office. There are significant differences in data formats, including structured tables, vector maps, satellite images, and various dynamic logs. The overall initial data volume reaches approximately 86 TB.
[0051] To address the problems of multi-source heterogeneous data and inconsistent formats, the project team established an industry-standardized data dictionary and metadata specification in accordance with the method of this invention, unified the naming rules, data types, and unit systems. For example, the "route code" and "route name" in the road attribute table were standardized into unified field names, and the coding length and format were specified; for satellite image data, a unified spatial reference system information field was established, including EPSG code, resolution, etc. Through standardization processing, the data compatibility was greatly improved, and the field consistency rate increased from the original 61% to over 95%.
[0052] At the data storage level, according to the hybrid storage system proposed in this invention, different types of data are stored in appropriate databases respectively. Structured data is managed by MySQL, storing table information such as traffic flow and infrastructure attributes. Spatial vector data is managed by PostGIS, and spatial indexes and auxiliary hash indexes are constructed for road network nodes, channel networks, airport layouts, etc. Satellite image data is sliced through Geoserver in the form of a tile pyramid, and the metadata is stored in the MongoDB document database. At the same time, a full-text search index is established with the help of ElasticSearch. High-frequency query results and status caches are managed using Redis-like K-V storage.
[0053] In the process of spatial data organization, a spatial data sharding strategy was formulated according to geographical area division. Spatial sharding was carried out with the municipal administrative region as the basic unit and further refined into longitude and latitude grids. Taking the highway network as an example, originally containing 1.24 million road segment elements, about 520 sub-layers were generated after spatial sharding, and each sub-layer did not exceed 100 MB, ensuring that the target area could be quickly located during query, greatly reducing the system load. After testing, the time-consuming of a typical spatial query such as "retrieving the location of highway interchange nodes in northern Anhui" decreased from the original 3.2 seconds to 0.48 seconds, an improvement of about 6.67 times.
[0054] In building a data retrieval platform, a unified retrieval and visualization analysis platform is constructed based on an integrated transportation database. Users can quickly query information on facilities such as roads, waterways, airports, etc. and related business data through multi-dimensional conditions such as keywords, spatial scope, or timestamp. The platform supports real-time rendering of the overlay display of raster images and spatial vector data, enabling dynamic visualization analysis of highway networks, port layouts, and route networks. Taking the special application of "Spring Festival Transportation Situation Monitoring" carried out in January 2025 as an example, the dispatching and command center retrieves the traffic flow changes on the main sections of the provincial expressways in the past 30 days through the platform. The average retrieval response time is less than 2 seconds, and the data export and analysis efficiency has increased by more than 5 times.
[0055] In terms of dynamic data update, the present invention effectively solves the problem of spatio-temporal data consistency through an incremental update mechanism and a version control system. Taking the data update cycle in December 2024 as an example, the system has cumulatively accessed approximately 180 million traffic flow records, approximately 32,000 port throughput update records, and approximately 270,000 dynamically accessed flight operation records. The incremental update mechanism controls the growth of the storage space for newly added data to approximately 8%, effectively avoiding the disk waste and retrieval efficiency decline problems caused by full-scale updates.
[0056] After a six-month trial operation and evaluation, the integrated transportation infrastructure database platform has achieved a data coverage rate of 98.2%, a spatial retrieval accuracy rate of 97.5%, an average query response time controlled within 0.5 seconds, and the number of concurrent user accesses supported has increased to more than 5 times that of the original system. The application scope covers the dispatching centers of transportation departments, urban transportation bureaus, emergency management departments, and key port operation units, supporting traffic dispatching and emergency response during major periods such as the Spring Festival in 2025, demonstrating excellent stability, scalability, and application value.
[0057] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered within the protection scope of the present invention.
Claims
1. A method for constructing a comprehensive transportation infrastructure spatial database based on multi-source data, characterized by: The steps include: S1. Collect multi-source data on three types of transportation infrastructure: roads, water transport, and aviation; S2. Establish data dictionaries and metadata standards for collected multi-source data, formulate database partitioning strategies and unified interface specifications based on industry standards; S3, standardize the infrastructure attribute data, satellite image data and spatial vector data in multi-source data; S4. Develop a spatial data sharding strategy based on geographic regions, build a full-text index and spatial index system, and optimize storage and query in a replicated cluster environment. S5. Build a comprehensive transportation infrastructure spatial database, using MySQL, PostGIS, MongoDB, ElasticSearch, and KV storage for collaborative management; S6. Design expansion mechanisms at the storage, retrieval, and computing layers to accommodate the growth of transportation infrastructure data and changes in retrieval requirements. S7. Establish a unified retrieval platform based on the comprehensive transportation infrastructure spatial database, integrate raster image data and spatial vector data, and support spatial visualization and semantic retrieval.
2. The method for constructing a comprehensive transportation infrastructure spatial database based on multi-source data according to claim 1 is characterized in that: The multi-source data includes infrastructure attribute data, traffic flow data, navigation flow data, aviation business volume data, travel origin-destination distribution data, administrative division data, economic data, population data, satellite image data and spatial vector data.
3. The method for constructing a comprehensive transportation infrastructure spatial database based on multi-source data according to claim 1 is characterized in that: The standardization process includes: Standardize fields of infrastructure attribute data; Satellite image data is sliced using a tile pyramid approach and metadata is managed using a combination of full-text indexing and the MongoDB document model. Spatial vector data is expressed in a standardized format using GeoJSON and stored in the PostGIS spatial database.
4. The method for constructing a comprehensive transportation infrastructure spatial database based on multi-source data according to claim 1 is characterized in that: The replication cluster environment includes sharding storage of spatial data, full-text indexes, and metadata, configuring a master copy and at least one slave copy for each data shard, using a replica synchronization mechanism to maintain data consistency, performing load balancing queries based on replica sets, and having the slave copy take over the master copy role when a node fails.
5. The method for constructing a comprehensive transportation infrastructure spatial database based on multi-source data according to claim 1 is characterized in that: The S2 specifically includes: S21. Based on national standards and industry technical specifications in the field of geographic information, establish a data dictionary for the collected multi-source data. The data dictionary defines the data field name, field data type, field value range, and field unit type. S22. Establish metadata standards based on multi-source data attribute information. The metadata standards include fields such as dataset name, data source, spatial reference system, data update time, and data access rights. S23. Develop a database partitioning strategy to store infrastructure attribute data, traffic flow data, navigation flow data, air traffic volume data, travel origin-destination distribution data, administrative division data, economic data, and population data in a MySQL relational database; store satellite imagery data in a MongoDB document-based database; and store spatial vector data in a PostGIS spatial database. S24. Develop unified interface specifications, which include data access interface, data retrieval interface and data exchange format types.
6. The method for constructing a comprehensive transportation infrastructure spatial database based on multi-source data according to claim 1 is characterized in that: The S3 specifically includes: S31. Standardize the fields of infrastructure attribute data to unify data field names, field data types, field value ranges, and field unit types; S32. Perform tile pyramid layered slicing processing on satellite image data: ; in, is the number of slice levels, is the original image spatial resolution, is the target tile resolution, Indicates rounding up operation; S33. Use Geoserver to generate slice services for the sliced satellite image data, establish full-text indexes with ElasticSearch, and manage the metadata of the image data based on the MongoDB document model. S34. Spatial vector data is standardized in GeoJSON format and stored in the PostGIS spatial database. A spatial index based on the R-tree structure and an auxiliary index based on the hash structure are established for the spatial vector elements.
7. The method for constructing a comprehensive transportation infrastructure spatial database based on multi-source data according to claim 1 is characterized in that: The S4 specifically includes: S41. Divide spatial data based on geographical regions and formulate sharding strategies: ; in, Number the spatial data slices. and is the latitude and longitude coordinates of the spatial data, and are the minimum longitude and latitude of the spatial region, and Divide the intervals into spatial slices, is the number of slices in the longitude direction, Indicates a round-down operation; S42. Within the spatial data shard, establish an R-tree-based spatial index for the spatial elements, record the minimum bounding rectangle parameters of the spatial objects, and use a hash-assisted index to accelerate attribute retrieval; S43. Establish a full-text index for the metadata associated with the spatial data using the ElasticSearch indexing engine. The index fields include the dataset identifier, spatial range description, timestamp, and key information tags. S44. In a replication cluster environment, spatial data shards and full-text index replicas are distributed and stored, and query optimization is performed based on the load status of the replica nodes.
8. The method for constructing a comprehensive transportation infrastructure spatial database based on multi-source data according to claim 1 is characterized in that: The S5 specifically includes: S51. Use a MySQL relational database to store structured data, including infrastructure attribute data, traffic flow data, air traffic flow data, aviation business volume data, travel origin-destination distribution data, administrative division data, economic data, and population data; S52. Use PostGIS spatial database to manage spatial vector data and associated attributes, establish spatial index based on R-tree structure, and use auxiliary hash index to accelerate attribute retrieval; S53. Use MongoDB document-based database to manage metadata, including dataset description information, spatial reference system parameters, data update time, and data access permission identifier; S54. Use the ElasticSearch search engine to manage full-text index data, where the full-text index data includes a dataset identifier, a spatial range description, a time tag, and a description of key attributes; S55. Use the KV storage system to manage frequently accessed cache data, including query result cache, index metadata cache, and replica synchronization status cache.
9. The method for constructing a comprehensive transportation infrastructure spatial database based on multi-source data according to claim 1, characterized in that: The S6 specifically includes: S61. The storage layer consists of MySQL, MongoDB, and PostGIS. A data expansion mechanism is designed in the storage layer, and partitioning and replica expansion methods are used to expand data capacity. S62. The retrieval layer is composed of ElasticSearch. An index expansion mechanism is designed in the retrieval layer to expand the retrieval load based on full-text index sharding and replica sharding strategies. S63. The computing layer is composed of shard clusters. A task expansion mechanism is designed at the computing layer, and a task load sharding and node concurrency expansion strategy is adopted.
10. The method for constructing a comprehensive transportation infrastructure spatial database based on multi-source data according to claim 1, characterized in that: The S7 specifically includes: S71. Build a unified search platform based on the comprehensive transportation infrastructure spatial database, which accesses structured data, spatial data, metadata, and full-text index data; S72, integrate raster image data and spatial vector data to build a multi-level spatial visualization model, supporting overlay rendering of spatial elements and raster layers; S73. Establish a semantic retrieval mechanism based on full-text indexing, supporting combined queries based on keywords, spatial ranges, timestamps, and attribute tags; S74. Based on the search results, dynamically load the spatial vector elements and the corresponding raster image layers for visualization and view update.
Citation Information
Patent Citations
Space-time data management method
CN112328583A
Unconscious intelligent retrieval method based on remote sensing space big data
CN113987024A
Geographic information cloud disk for efficient retrieval and browsing based on cloud computing
CN114328779A
Space-time big data platform based on micro service
CN115238015A
Task scheduling method, device and storage medium for spatiotemporal big data
CN119739745A
Cited By
Multi-source disaster data fusion and standardization processing method and system
CN121501871A
Space-time information system and processing method based on multi-source remote sensing data
CN122086969A