A method and system for dynamically constructing high-quality datasets for transportation and logistics

By accessing multiple authoritative data sources, performing secure caching and standardization, identifying key topological feature points of the logistics network, constructing an associated vector model, and generating dynamic correction parameters, the problem of disconnect between fused data and the operational status of the logistics network in existing technologies is solved. This enables the dynamic construction of high-quality datasets, supporting accurate tracing of road network congestion and emergency material dispatch.

CN122066033BActive Publication Date: 2026-07-31RES INST OF HIGHWAY MINIST OF TRANSPORT
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RES INST OF HIGHWAY MINIST OF TRANSPORT
Filing Date
2026-02-09
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies lack dynamic data correction methods based on the evolution of logistics network topology, resulting in a disconnect between fused data and the actual operational status of the logistics network. This makes it impossible to accurately locate the correlation between congestion and node operation efficiency and channel connectivity, affecting the accurate decision-making for tracing the source of road network congestion and dispatching emergency supplies.

Method used

By accessing multiple authoritative data sources, performing secure caching and standardization, identifying key topological feature points of the logistics network, constructing an associated vector model and generating dynamic correction parameters, combining business rules and artificial intelligence for data annotation and repair, generating a hierarchical thematic database, and providing traceable data services, the system is optimized and evolved based on application feedback.

Benefits of technology

It enables the dynamic construction of high-quality datasets, accurately supporting precise decision-making in tracing the source of road network congestion and dispatching emergency supplies, thereby enhancing the foresight and efficiency of industry governance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122066033B_ABST
    Figure CN122066033B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for dynamically constructing high-quality transportation and logistics datasets, relating to the field of intelligent logistics information service technology. The method includes: accessing data from multiple authoritative data sources and securely caching it to obtain secure cache data; performing fusion and standardization processing on the secure cache data based on industry indicators to form spatiotemporally fused standardized data; analyzing the spatiotemporal distribution of the spatiotemporally fused standardized data to determine key topological feature points of the logistics network, where these key topological feature points correspond to core operating areas within logistics hubs, multimodal transport transshipment connection points, and bottleneck locations of major channels; constructing an association vector model based on the key topological feature points, and generating dynamic correction parameters reflecting the evolution of the network topology structure based on the association vector model. This invention is applicable to the integration, governance, quality control, and service output of multi-source transportation and logistics data, providing data support for precise industry governance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent logistics information service technology, and in particular to a method and system for dynamically constructing high-quality traffic and logistics datasets. Background Technology

[0002] Currently, my country's transportation and logistics industry is undergoing a continuous and in-depth digital transformation. Sensing equipment and business systems such as ETC gantries, freight platforms, and port and shipping management systems are constantly generating massive amounts of traffic operation and logistics activity data.

[0003] However, when supporting precise scenarios such as tracing the source of road network congestion and dispatching emergency supplies, existing technologies mostly suffer from the following shortcomings: They lack dynamic data correction and deep fusion mechanisms based on the evolution of logistics network topology. Existing data fusion is mostly superficial, failing to fully identify key topological features such as core operating areas of logistics hubs, multimodal transport transshipment points, and channel bottlenecks. They also fail to construct evolutionary models of the strength and direction of correlations between these features, and lack dynamic data correction methods based on topology evolution, resulting in a severe disconnect between fused data and the actual operational status of the logistics network. For example, at multimodal transport hub transshipment points, existing technologies simply overlay traffic flow and freight document data without dynamically correcting for changes in the operational efficiency of the transshipment points, failing to truly reflect the intrinsic correlation between transshipment efficiency and freight flow speed. Similarly, at bottleneck locations on major transport channels, only isolated statistics on congestion duration are presented, without linking freight trajectory flow patterns before and after the bottleneck, making it impossible to accurately analyze the causes of congestion and the impact of logistics delays.

[0004] This deficiency directly leads to the following: when tracing the source of road network congestion, the data does not dynamically match the topological evolution characteristics, making it difficult to accurately locate the correlation between congestion and node operation efficiency and channel connectivity, resulting in a lack of targeted governance measures; in emergency material dispatch, static data that has not been topologically corrected cannot adapt to hub load fluctuations and channel status changes, resulting in insufficient alignment of dispatch decisions. Summary of the Invention

[0005] The technical problem to be solved by this invention is to provide a method and system for dynamically constructing high-quality datasets for transportation and logistics, which solves the problem of existing fused data being disconnected from the actual operation status of the logistics network, and provides high-quality data support for accurate decision-making and rational governance in the industry.

[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0007] Firstly, a method for dynamically constructing a high-quality dataset for transportation and logistics, the method comprising:

[0008] Data from multiple authoritative data sources is accessed and securely cached to obtain data in the secure cache area;

[0009] The data in the security cache is fused and standardized based on industry indicators to form standardized data with spatiotemporal fusion.

[0010] The spatiotemporal distribution of standardized data from spatiotemporal fusion is analyzed to identify key topological feature points of the logistics network. These key topological feature points correspond to core operating areas within logistics hubs, multimodal transport transshipment connection points, and bottleneck locations of main channels. Based on these key topological feature points, an association vector model is constructed, and dynamic correction parameters reflecting the evolution of the network topology are generated according to the association vector model.

[0011] For standardized spatiotemporal fusion data, quality control and labeling are performed based on business rules and artificial intelligence. Dynamic correction parameters are used to correct the data labeling and repair process to obtain optimized data with quality labels.

[0012] The optimized data with quality labels is subjected to quality assessment and classification for thematic applications, and the data of different quality levels are stored in the corresponding thematic database to generate a graded thematic database.

[0013] The data in the hierarchical subject database is dynamically released in a versioned manner and provided with traceable data services to obtain the released dataset, while collecting application feedback on the released dataset;

[0014] Based on application feedback and dynamic correction parameters of the released dataset, targeted data requirements and governance instructions are generated in reverse to drive the optimization and evolution of the dataset.

[0015] Furthermore, the data in the security cache is fused and standardized based on industry indicators to form standardized spatiotemporal fusion data, including:

[0016] Data elements related to key indicators such as vehicle travel time, freight origin and destination, and channel traffic are extracted from the data in the security cache; alignment and association operations are performed on the data elements to obtain initial fused data;

[0017] Using a unified spatial grid and time reference, traffic flow, speed information and freight trajectory information in the initial fusion data are integrated to generate logistics traceability chain data.

[0018] The system invokes predefined standardized rules to standardize and transform the data elements, formats, and codes in the logistics traceability chain data. The predefined standardized rules are formulated based on the standards of the transportation government data resource catalog and the comprehensive transportation master data catalog, and the transformed data forms a spatiotemporal fusion standardized data.

[0019] Furthermore, the spatiotemporal distribution of standardized data from spatiotemporal fusion is analyzed to identify key topological feature points of the logistics network. These key topological feature points correspond to core operating areas within logistics hubs, multimodal transport transshipment connection points, and bottleneck locations of major channels. Based on these key topological feature points, an association vector model is constructed, and dynamic correction parameters reflecting the evolution of the network topology are generated according to the association vector model, including:

[0020] Analyze standardized data that integrates spatiotemporal data to identify the aggregation and flow patterns of logistics activities in the spatiotemporal dimensions, and obtain the spatiotemporal distribution characteristics of logistics activities;

[0021] Based on the spatiotemporal distribution characteristics of logistics activities, key topological feature points are identified in the core operating area, multimodal transport transshipment connection points, and bottleneck locations of major channels within the logistics hub.

[0022] Using the key topological feature points as nodes, construct association vectors that characterize the strength and direction of logistics associations between nodes, and form an association vector model from all association vectors;

[0023] By analyzing the changes in the direction and intensity of vectors in the correlation vector model, dynamic correction parameters reflecting the evolution of network topology are generated.

[0024] Furthermore, the standardized spatiotemporal fusion data undergoes quality control and labeling based on business rules and artificial intelligence. Dynamic correction parameters are used to correct the data labeling and repair process, resulting in optimized data with quality labels, including:

[0025] Based on predefined industry business rules, the standardized spatiotemporal fusion data is automatically verified and its logical compliance is determined to obtain preliminary labeled data.

[0026] A pre-trained artificial intelligence model is used to classify and identify complex patterns and impute missing values ​​in the initial labeled data to generate AI-enhanced labeled data.

[0027] For complex anomalies in AI-enhanced labeled data, domain expert knowledge is integrated to review and correct them, forming expert-verified labeled data.

[0028] By using dynamic correction parameters, the spatiotemporal logical consistency and topological correlation in the expert-verified annotation data are corrected to obtain optimized data with quality annotations.

[0029] Furthermore, the optimized, quality-labeled data undergoes a quality assessment and grading based on thematic applications. Data of different quality levels are then stored in corresponding thematic databases to generate a tiered thematic database, including:

[0030] Extract the spatiotemporal coordinates and quality labels from the optimized quality-labeled data, analyze their distribution characteristics in the preset key logistics areas, and obtain spatial distribution characteristic data.

[0031] Based on predefined key area weights and spatial distribution characteristic data, an adaptive coverage integrity assessment under spatiotemporal constraints is performed to generate spatial coverage quality indicators.

[0032] By combining the spatial coverage quality indicators with the quality labels of the optimized quality-annotated data itself, a comprehensive quality assessment and classification of the data is performed, generating data quality classification results.

[0033] Based on the data quality level classification results, the data are stored in the corresponding thematic databases to generate a hierarchical thematic database.

[0034] Furthermore, the data in the hierarchical topic database will be dynamically released in a versioned manner and traceable data services will be provided to obtain the released dataset. At the same time, application feedback on the released dataset will be collected, including:

[0035] Extract the data to be published from the hierarchical topic database and encapsulate it into a data packet that conforms to the preset service interface format to obtain the data packet to be published;

[0036] Add version identifiers and data traceability information to the data packages to be released, and generate versioned release packages;

[0037] Versioned release packages are published through the data service interface, and query services based on data source information are provided to obtain the released dataset;

[0038] Monitor and collect feedback information from analytical reports or decision-making applications based on published datasets, as application feedback for published datasets.

[0039] Furthermore, based on application feedback from the released dataset and dynamic correction parameters, targeted data requirements and governance instructions are generated in reverse to drive the optimization and evolution of the dataset, including:

[0040] Analyze application feedback on the published dataset to identify issues such as insufficient data support or decision-making biases, and obtain feedback analysis results;

[0041] By combining dynamic correction parameters, the correlation between the feedback analysis results and the logistics network structure and evolution characteristics is corrected, and the corrected optimization requirements are generated.

[0042] Based on the corrected optimization requirements, generate targeted data collection and governance instructions for specific data sources, processing stages, or quality indicators;

[0043] The targeted data collection and governance instructions are fed back to the corresponding data access, processing, or quality control links to drive the optimization and evolution of the dataset construction process.

[0044] Secondly, a system for dynamically constructing high-quality transportation and logistics datasets includes:

[0045] The data caching module is used to access data from multiple authoritative data sources and perform secure caching to obtain data in the secure cache area;

[0046] The fusion processing module is used to perform fusion and standardization processing on the data in the security cache based on industry indicators, forming standardized data with spatiotemporal fusion.

[0047] The model building module is used to analyze the spatiotemporal distribution of standardized data from spatiotemporal fusion, determine key topological feature points of the logistics network, where key topological feature points correspond to the core operating areas, multimodal transport transshipment connection points, and bottleneck locations of main channels within the logistics hub; construct an association vector model based on the key topological feature points, and generate dynamic correction parameters reflecting the evolution of the network topology based on the association vector model;

[0048] The optimization and correction module is used to perform quality control and labeling on standardized spatiotemporal fusion data based on business rules and artificial intelligence, and to use dynamic correction parameters to correct the data labeling and repair process to obtain optimized data with quality labels.

[0049] The assessment and grading module is used to perform quality assessment and grading of the optimized quality-labeled data for thematic applications, and to store data of different quality levels into the corresponding thematic database to generate a graded thematic database.

[0050] The data publishing and feedback collection module is used to dynamically publish data from the hierarchical subject database in a versioned manner and provide traceable data services, obtain the published dataset, and collect application feedback on the published dataset.

[0051] The data optimization and evolution module is used to generate targeted data requirements and governance instructions based on application feedback and dynamic correction parameters of the released dataset, thereby driving the optimization and evolution of the dataset.

[0052] The above-described solution of the present invention has at least the following beneficial effects:

[0053] By employing compliant access and standardized integration of multi-source authoritative data, dynamic correction parameter generation driven by topology evolution, quality control through business rules, artificial intelligence, and expert collaboration, and thematic-oriented hierarchical data entry and closed-loop optimization, this technology effectively overcomes the core technical problems of existing technologies lacking the ability to adapt to the evolution of logistics network topology and the disconnect between integrated data and actual operating status. This results in the generation of a standardized, reliable, and continuously evolving high-quality dataset for transportation and logistics, accurately supporting precise decision-making in scenarios such as tracing the source of road network congestion and dispatching emergency supplies, and effectively improving the foresight and efficiency of industry governance. Attached Figure Description

[0054] Figure 1 This is a flowchart illustrating a method for dynamically constructing a high-quality traffic and logistics dataset according to an embodiment of the present invention.

[0055] Figure 2 This is a schematic diagram of a dynamic construction system for high-quality traffic and logistics datasets provided by an embodiment of the present invention. Detailed Implementation

[0056] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0057] like Figure 1 As shown, an embodiment of the present invention proposes a method for dynamically constructing a high-quality traffic and logistics dataset, the method comprising the following steps:

[0058] Step 1: Access data from multiple authoritative data sources and perform secure caching to obtain data in the secure cache area;

[0059] Step 2: Perform fusion and standardization processing on the security cache data based on industry indicators to form standardized data with spatiotemporal fusion;

[0060] Step 3: Analyze the spatiotemporal distribution of the standardized data from spatiotemporal fusion to determine the key topological feature points of the logistics network. These key topological feature points correspond to the core operating areas, multimodal transport transshipment connection points, and bottleneck locations of major channels within the logistics hub. Based on the key topological feature points, construct an association vector model and generate dynamic correction parameters that reflect the evolution of the network topology structure according to the association vector model.

[0061] Step 4: Perform quality control and labeling on the standardized spatiotemporal fusion data based on business rules and artificial intelligence, and use dynamic correction parameters to correct the data labeling and repair process to obtain optimized data with quality labels.

[0062] Step 5: Perform quality assessment and grading on the optimized data with quality labels for thematic applications, and store data of different quality levels into the corresponding thematic database to generate a graded thematic database.

[0063] Step 6: Dynamically publish the data in the hierarchical topic database in a versioned manner and provide traceable data services to obtain the published dataset, while collecting application feedback on the published dataset.

[0064] Step 7: Based on the application feedback and dynamic correction parameters of the released dataset, reverse-engineer targeted data requirements and governance instructions to drive the optimization and evolution of the dataset.

[0065] In this embodiment of the invention, compliant access and secure caching of multi-source authoritative data ensure the reliability and security of data sources; the fusion and standardized processing based on industry indicators achieves the unification of data formats and standards; key topological feature point identification and dynamic correction parameter generation enable data to accurately match the operational evolution of the logistics network; quality control through collaboration of business rules, artificial intelligence, and expert knowledge effectively improves data quality and credibility; thematic hierarchical data entry meets the differentiated data needs of different scenarios; versioned release and traceability services enhance the standardization and transparency of data use; and closed-loop optimization based on application feedback drives the continuous adaptation of datasets to industry needs, ultimately providing stable data support for accurate industry decision-making and efficient governance.

[0066] In a preferred embodiment of the present invention, step 1 above may include:

[0067] Step 1.1 involves real-time access to multiple authoritative industry data sources from the fields of transportation operations, logistics activities, and environment and events to obtain multi-source raw data. Specifically, this includes: firstly, verifying the qualifications of authoritative industry data sources in the three major fields of transportation operations, logistics activities, and environment and events to confirm the official authority and timeliness of the data sources, and selecting target data sources that meet the core business needs of the industry; then, building a standardized data access gateway and establishing a stable data transmission channel, configuring appropriate access protocols for the transmission characteristics of different types of data sources to ensure continuous and delay-free access to multiple types of data, such as traffic operation data at the ministerial and provincial levels, key freight platform data, and real-time meteorological data; simultaneously establishing a data access status monitoring mechanism during the access process, providing real-time alarms for abnormal situations such as data transmission interruptions and data loss, and automatically triggering a retry mechanism, ultimately aggregating multi-source raw data covering the core business scenarios of the industry.

[0068] Step 1.2 involves de-identifying and cleaning the multi-source raw data according to data security and privacy protection standards to obtain de-identified data. Specifically, this includes: developing a tiered de-identification strategy based on industry-standard data security and privacy protection guidelines; targeting personal identity information and sensitive business information in the multi-source raw data using compliant de-identification methods such as mask replacement, scope generalization, and field splitting to ensure privacy is not leaked; subsequently, systematic data cleaning is conducted, performing cross-field logical verification based on mature industry business rules; identifying redundant and duplicate records, incorrectly formatted data, and invalid data unrelated to core business through field format comparison, data range verification, and correlation information matching; to avoid over-cleaning, reasonable logical verification thresholds are set based on the actual needs of core business scenarios; and filtering is performed by judging the degree of deviation between the data and the thresholds. Data that slightly deviates from the thresholds and does not affect the core business relationship is analyzed, retained, and labeled; only abnormal data that significantly exceeds the thresholds is removed. This ensures data cleanliness while fully preserving the core business relationships between data, ultimately forming de-identified data that meets compliance requirements and has high business usability.

[0069] Step 1.3 involves storing the anonymized data in a temporary storage area with a secure isolation mechanism to form a secure cache data. This includes: constructing a temporary storage area based on a distributed architecture, employing a security protection scheme combining physical and logical isolation. Physically, this involves deploying an independent server cluster and dedicated storage devices, physically isolating the data from unauthorized external networks. Logically, this involves dividing the storage into independent domains through virtual network partitioning and data isolation boundaries to avoid cross-interference between different types of data. Furthermore, access control lists are used to clearly define the data access scope for each role, strictly controlling data access permissions and granting data access permissions only to authorized subsequent processing steps for the corresponding business scenario, prohibiting irrelevant operations. The storage area is also configured with data encryption storage functionality, performing real-time encryption processing on the stored anonymized data and recording complete operation logs to trace the entire process of data storage, retrieval, and modification. A storage directory structure is designed according to classification rules such as data domain and access time, storing the anonymized data in an orderly partitioned manner. This ensures both the security and stability of data storage and provides support for subsequent rapid data retrieval and efficient access, ultimately forming a structured secure cache data.

[0070] In a preferred embodiment of the present invention, step 2 above may include:

[0071] Step 2.1: Extract data elements related to key indicators such as vehicle travel time, freight origin and destination, and channel traffic from the safety buffer data; perform alignment and association operations on the data elements to obtain initial fused data. Specifically, this includes: first, combining the business logic of the transportation and logistics industry, sorting out the business connotations of the three key indicators—vehicle travel time, freight origin and destination, and channel traffic—establishing a precise mapping relationship between the indicators and the data elements in the safety buffer, and confirming the core mandatory data elements and auxiliary data element ranges corresponding to each indicator; from the traffic operation and logistics activity classification data in the safety buffer, batch-screen target data elements according to the mapping relationship, and simultaneously perform data validity verification during the screening process, eliminating... Invalid data elements include those that are empty, have values ​​outside the reasonable business scope (e.g., negative travel time or abnormally high traffic volume). Using unique vehicle identifiers (e.g., license plate number and vehicle identification number) and waybill numbers as core association keys, the field names of data elements from different data sources are first standardized; for example, the start time and departure time are standardized to "departure time". Data precision and encoding formats are checked to align data elements across data sources. Then, traffic operation data elements and logistics activity data elements are associated through key-value matching. If multiple conflicting data sets correspond to the same association key, the data from the data source with the higher authority level is prioritized to ensure logical consistency between data elements, ultimately resulting in initial fused data with unified fields and valid associations.

[0072] Step 2.2: Using a unified spatial grid and time reference, traffic flow, speed information, and freight trajectory information from the initial fusion data are integrated to generate logistics traceability chain data. Specifically, this includes: combining transportation channel density, logistics hub distribution, and business application scenarios such as real-time monitoring and path analysis; based on the industry-standard geographic grid coding system, flexibly adjusting the grid cell size to divide spatial grid cells covering the entire logistics network; simultaneously setting UTC time as the unified time reference; and setting time slices ranging from one to five minutes according to business needs to complete the standardized conversion of timestamps for all initial fusion data; and then... The dual dimensions of inter-grid cells and time slices are used to break down and classify traffic flow and speed information in the initial fused data. Then, the data such as freight trajectory, cargo type, waybill status, and hub entry and exit records in the corresponding dimensions are accurately matched by spatiotemporal coordinates. If there are data gaps in the spatiotemporal dimension, the intermediate node data is filled by neighboring data interpolation to ensure the continuity of the cargo flow chain. The entire process data of cargo from the place of origin, through various traffic nodes and hub stations to the destination is connected in chronological order to form a complete logistics traceability chain data including spatiotemporal coordinates, traffic status, logistics information and flow nodes.

[0073] Step 2.3 involves invoking predefined standardization rules to standardize the data elements, formats, and codes in the logistics traceability chain data. These predefined standardization rules are formulated based on the standards of the Transportation Government Data Resource Catalog and the Comprehensive Transportation Master Data Catalog. The resulting standardized data is spatiotemporally integrated. Specifically, this includes: using the Transportation Government Data Resource Catalog and the Comprehensive Transportation Master Data Catalog as core bases, and combining them with the actual application needs of core transportation and logistics business scenarios such as statistical reporting and monitoring and early warning, refining the predefined standardization rules, confirming unified data element names, format requirements (e.g., date and time formats and numerical units), the correspondence of classification codes (e.g., cargo type codes and administrative division codes), and detailed rules for handling abnormal data. Industry experts are then organized to review and calibrate the rules. These rules are then used to perform field-by-field standardization processing on the logistics traceability chain data, unifying date and time data in different formats into YYYY-MM-DD. The HH:MM:SS format is used to uniformly convert speed data in different units to kilometers per hour, and replace non-standard codes such as freight origin and destination points and channel types with industry-standard codes. During processing, if data elements are found to be inconsistent with the rules, such as incorrect format or mismatched code, abnormal data is marked in time and the reason for the abnormality is recorded to ensure that core business information is not lost. After processing, a full consistency check is first performed on the data to check whether the data element names, formats, and codes fully comply with the rules. Then, 10% to 20% of the sample data is selected for sampling verification to verify the standardization effect. If problems are found, the processing method is adjusted and optimized again to finally form standardized spatiotemporal fusion data with unified format, standardized code, and consistent logic.

[0074] In a preferred embodiment of the present invention, step 3 above may include:

[0075] Step 3.1 involves analyzing the standardized spatiotemporal fusion data to identify the aggregation and flow patterns of logistics activities in the spatiotemporal dimensions, thereby obtaining the spatiotemporal distribution characteristics of logistics activities. Specifically, this includes introducing an innovative dynamic Riemannian manifold vector field co-evolution algorithm. The core of this algorithm is to characterize the dynamic characteristics of the curved spatial topology of the logistics network using Riemannian manifolds, combined with vector fields to capture cargo flow trends, achieving collaborative analysis to accurately uncover the spatiotemporal patterns of logistics activities. Firstly, an equidistant embedding mapping method is used to accurately map the spatial coordinates in the standardized spatiotemporal fusion data to a unified Riemannian manifold space. During the mapping process, the actual spatial topology of the core logistics channels across the entire region and regional core logistics hubs serves as constraints, defining the effective domain of the manifold analysis to ensure that the manifold structure aligns with the true... The physical logistics network has a consistent spatial distribution. Subsequently, an initial manifold grid structure is constructed, with grid nodes strictly corresponding to the spatial coordinates of logistics infrastructure sites such as highway entrances and exits, hub park loading and unloading points, and port terminal operation areas. The initial manifold grid is generated using the Delaunay triangulation method to ensure the spatial coverage integrity and topological rationality of the grid. Dynamic weights are assigned to each node of the manifold grid, and a hierarchical weighted calculation mechanism is established: first, cargo flow, circulation frequency, and operational busyness are divided into core business indicator layers, and the indicator weight ratio is set according to industry business priority. Then, the initial weight value is obtained through weighted summation. The weight update adopts a real-time iteration mechanism, calling the latest accessed standardized data every hour to recalculate, ensuring that the weights can dynamically match changes in logistics activities.

[0076] Next, a threshold for triggering weight changes is set. When the change in node weight exceeds the preset threshold for two consecutive iterations, the manifold structure adaptive adjustment mechanism is activated: For areas with dense logistics activities, the ability to depict the details of spatial topology is enhanced by increasing the local curvature of the manifold. At the same time, a grid node splitting strategy is adopted to achieve grid densification, that is, the original grid nodes are subdivided and new nodes are added to supplement key detail areas of logistics activities. For areas with sparse logistics activities, the topology structure is simplified by reducing the local curvature of the manifold, and a grid node merging strategy is adopted to reduce redundant calculations, merging adjacent low-weight nodes into a comprehensive node. Meanwhile, all cargo flow trajectories corresponding to each manifold grid node are extracted. A trajectory direction vector fitting method is used to cluster the flow trajectories of the same node at different time periods, and the average direction of each category of trajectory is calculated as the local direction vector of the node. The local direction vectors of all nodes are concatenated according to the spatial topological relationship of the manifold grid to form a continuous logistics vector field covering the entire manifold domain.

[0077] Finally, combining the spatial topological connectivity of Riemannian manifolds with the directional distribution patterns of vector fields, the spatiotemporal clustering patterns of logistics activities are identified through the vector field convergence density statistical method. This involves statistically analyzing the number and density of nodes where vectors converge within a unit space, and identifying areas with densities higher than a set benchmark as clustering regions. The core paths of cargo flow are then traced through the continuous vector pointing method of adjacent manifold nodes. This involves selecting node chains with continuous and consistent vector directions to form complete cargo flow paths. Finally, by integrating the clustering patterns and path characteristics, accurate spatiotemporal distribution characteristics of logistics activities are obtained.

[0078] Step 3.2: Based on the spatiotemporal distribution characteristics of logistics activities, determine the corresponding key topological feature points in the core operating area, multimodal transport transshipment connection points, and bottleneck locations of main channels within the logistics hub. Specifically, this includes: based on the generated dynamic manifold structure and spatiotemporal distribution characteristics, determining the key topological feature points according to the following refined process: First, activate the high-density manifold node cluster screening mechanism, traverse the dynamic weight data of all manifold nodes in the entire area, calculate the average weight, and use 1.2 to 1.5 times this average as the initial weight threshold range. Then, combine industry business experience, such as peak operating times of the logistics hub, etc. To determine the flow saturation threshold, three sets of typical business scenario data were selected for multiple rounds of threshold calibration. After each round of calibration, the number of selected clusters, the proportion of logistics nodes covered, and the overlap with actual business hotspots were statistically analyzed. With an overlap of no less than 85% as the core criterion, a unique optimal weight threshold was finally determined. Based on this threshold, all manifold nodes were traversed, and nodes with weight values ​​higher than the threshold were divided into the same cluster according to the principle of spatial proximity. By calculating the manifold geodesic distance between nodes, nodes with a distance less than the preset proximity threshold were grouped into one cluster, forming multiple high-density manifold node clusters with clear boundaries and close internal connections.

[0079] Next, the core region is extracted using a weighted center coordinate calculation method. The specific calculation process is as follows: First, the total dynamic weights of all manifold nodes in each high-density cluster are calculated and retained to two decimal places. Then, for each manifold node in the cluster, the weight percentage of each node is calculated by: dynamic weight of a single node ÷ total weight of the cluster × 100%, with the weight percentage retained to one decimal place. Subsequently, weighted average calculations are performed on the x-axis and y-axis components of the node's spatial coordinates: When calculating the x-axis core coordinates, the x-axis coordinate value of each node (retaining six decimal places) is multiplied by its corresponding weight percentage, and the percentage value is directly used in the calculation. All product results are summed and retained to six decimal places. The y-axis core coordinates are calculated using the same logic. Finally, the core coordinates of the cluster are composed of the x-axis and y-axis core coordinates. After the calculation is completed, it is additionally verified whether the core coordinate points are within the geometric center range of the cluster's spatial distribution. If the deviation exceeds 10% of the maximum span of the cluster, the node weight data and calculation process are re-checked to ensure that the core coordinate points can truly reflect the core clustering area of ​​the cluster. The physical spatial location corresponding to this coordinate point is the key topological candidate feature point.

[0080] Finally, combining the distribution and evolution trend of vector fields on the Riemannian manifold, a multi-dimensional candidate feature point functional attribute determination mechanism is constructed: For the region where the candidate point is located, the vector directions of the region and its three adjacent grids are clustered. If the clustering results show that there are four or more vectors converging on the candidate point in different directions, and the fluctuation range of the manifold node weights in the region is lower than the calibration threshold for three consecutive time periods (the calibration threshold is set at 30% of the standard deviation of the weights of the entire region), then the candidate point is determined to be a key topological feature point corresponding to the core operating area inside the logistics hub. If the candidate point is located in the intersection area of ​​two or more different high-density node clusters, the vector directions of the three adjacent nodes before and after the candidate point are extracted through the vector direction deviation calculation logic, and the adjacent direction deviation is calculated. The sum of the angles between the quantities is considered a significant turn if it exceeds 45 degrees. Simultaneously, multimodal transport data is retrieved to check for transshipment records between different modes of transport, such as road and rail, or road and waterway. If both conditions are met, the point is identified as a key topological feature point corresponding to a multimodal transport transshipment connection point. If the vector field of the manifold region where the candidate point is located exhibits unidirectional flow characteristics, the vector density of the candidate point and its three downstream adjacent nodes is calculated using vector density statistics. If the average vector density of the downstream nodes decreases by more than 50% compared to the candidate point (this percentage is calibrated based on historical channel bottleneck data), and the weight of the manifold node corresponding to the candidate point exceeds twice the threshold, it indicates that cargo flow is obstructed, and the point is identified as a key topological feature point corresponding to a bottleneck location of the main channel.

[0081] Step 3.3: Using the key topological feature points as nodes, construct association vectors representing the strength and direction of logistics associations between nodes, and form an association vector model from all association vectors. Specifically, the association vector model constructed in this step is based on a Riemannian manifold-enhanced graph attention vector model, which is an improvement on graph neural networks (GNNs). This model is based on the core attention mechanism of graph attention networks (GATs) and is adapted to the topological characteristics of the Riemannian manifold space of logistics networks. It adds a manifold metric adaptation layer and a multi-dimensional association coupling layer, effectively solving the problems of vector representation distortion and association strength calculation deviation in traditional GATs under curved spatial topologies, and accurately matching the modeling needs of the dynamic evolution of logistics networks. The specific construction and training process is as follows:

[0082] The specific construction process of the associated vector model needs to be advanced step by step based on key topological feature points and the Riemannian manifold topology. First, the model architecture is adapted and adjusted, using the identified key topological feature points as the core nodes of the model, and building the basic framework of the model in conjunction with the reconstructed Riemannian manifold topology. A Riemannian manifold embedding adaptation layer is added between the input layer and attention layer of the traditional GAT, transforming the Euclidean space coordinates of the core nodes into Riemannian manifold space coordinates through equidistant embedding mapping, ensuring that the model can accurately adapt to the curved spatial topology characteristics of the logistics network. Simultaneously, a multi-dimensional association coupling layer is added after the attention layer, specifically for integrating multi-dimensional business features related to cargo flow, strengthening the business relevance of the vector representation.

[0083] Subsequently, input layer feature construction was carried out, building a model input feature set based on key topological feature points. The input features of each node contain three types of core information. The basic node features cover the Riemannian manifold coordinates, dynamic weight values, and functional attribute labels of key topological feature points, such as the core operating area of ​​a port, the bottleneck section of a trunk channel, and the multimodal transport transshipment point; the flow association features include the historical cargo flow volume, average flow time, and transportation mode ratio of the node and its potential adjacent nodes; the spatiotemporal auxiliary features include data such as the node's logistics activity activity level and surrounding traffic conditions over the past three time periods. In a specific scenario, for a core operating area node in a regional logistics hub cluster, its input features include the manifold coordinates of the operating area, the dynamic weights over the past hour, the corresponding functional attribute labels, historical flow data with surrounding logistics parks and trunk channel nodes, as well as information such as cargo handling volume over the past three hours and the traffic speed of surrounding evacuation channels.

[0084] The construction of the core graph attention vector calculation layer is a crucial step in model building. It requires activating the model's graph attention mechanism, calculating the manifold geodesic distance between any two core nodes based on the metric tensor of the Riemannian manifold, and simultaneously constructing the attention weight calculation logic by combining the flow association features from the node input feature set. Generally, node pairs that are closer in distance and have higher historical flow frequency are assigned higher attention weights. For example, in multimodal transport scenarios, major coastal port nodes and adjacent railway logistics center nodes are assigned higher attention weights due to their close geodesic distance and frequent sea-rail intermodal flow. After aggregating the feature information of adjacent nodes through the attention mechanism, this information is input into the multi-dimensional association coupling layer. Combined with data such as cargo flow, average flow rate, and manifold node weight coefficients, the core parameters of the association vector between nodes are obtained through weighted coupling calculation. The vector direction is determined by the dominant cargo flow direction guided by the attention weights and aligned with the geodesic direction on the Riemannian manifold; the vector magnitude is obtained by normalizing the coupling calculation results, thus accurately representing the strength of the logistics association between nodes.

[0085] Finally, the output layer is constructed, employing a dual-branch structure to output the direction and magnitude parameters of the association vectors. A vector concatenation mechanism then integrates the association vectors of all adjacent core nodes into a vector set. Combining this with the connection logic of Riemannian manifold topology, a complete association vector model is ultimately formed. This model is essentially a manifold topological vector network with key topological feature points as nodes and association vectors as edges, capable of intuitively and accurately reflecting the association relationships among the core nodes in the logistics network.

[0086] The specific training process of the correlation vector model needs to be carried out systematically based on historical data from typical logistics scenarios. In the training data preparation phase, historical spatiotemporal fusion standardized data from three typical logistics scenarios nationwide are selected as the training set. These scenarios include regional logistics hub clusters, cross-regional multimodal transport channels, and port-road linkage channels. The data covers key topological feature point flow data and manifold topology data for six consecutive months. Simultaneously, data from one month of the same period is selected as the validation set, and data from another month is selected as the test set. Training set labels are generated using a combination of manual annotation and real-world records from the business system. The labels contain the actual cargo flow correlation direction and the actual correlation strength for each node pair, where the actual correlation strength is represented by the coupling value of actual flow volume and throughput efficiency per unit time.

[0087] In the model initialization and hyperparameter setting phase, the mapping parameters of the Riemannian manifold embedding adaptation layer and the attention weight matrix of the graph attention layer are initialized first, using the Xavier initialization method to ensure uniform parameter distribution. Then, hyperparameters are set appropriately, with a batch size of 32, an initial learning rate of 0.001, an maximum of 100 iterations, 4 attention heads, and initial values ​​for the manifold curvature adaptation coefficients based on the average curvature of the training set data. During forward propagation training, the training set data is input into the model in batches. Coordinate mapping is completed through the Riemannian manifold embedding adaptation layer, attention weights are calculated and features are aggregated through the graph attention layer, and the association vector parameters are calculated through the multi-dimensional association coupling layer. Finally, the predicted association vector direction and magnitude are output. Taking a typical port-highway linkage scenario as an example, the training data of major coastal port nodes and surrounding highway hub nodes are input into the model to predict the association vector parameters between each port node and the surrounding highway hub nodes.

[0088] In the loss calibration and backpropagation stages, a two-dimensional calibration mechanism is constructed, consisting of Riemannian manifold geodesic direction deviation calibration and association strength deviation constraint. Direction deviation calibration, based on the geodesic metric logic of the Riemannian manifold, quantifies the degree of deviation between the predicted association vector direction and the actual cargo flow direction, applying targeted adjustment forces to the direction prediction deviation and guiding model parameters towards optimization in a dimension that aligns with the actual flow direction. Association strength deviation constraint establishes a deviation threshold control rule by comparing the deviation magnitude between the predicted magnitude and the actual association strength. When the deviation exceeds the threshold, a parameter adjustment instruction is triggered, constraining the prediction accuracy of association strength. A Riemannian stochastic gradient descent (SGD) optimizer is employed, adjusting the parameters of each layer of the model round-by-round along the gradient direction of the Riemannian manifold, focusing on optimizing the manifold embedding mapping parameters and the attention weight matrix. By continuously calibrating the direction and strength deviations, the accuracy of the model's representation of node association relationships is gradually improved.

[0089] During the iterative training and validation optimization phase, the model performance is evaluated using validation set data after every 5 iterations, with orientation prediction accuracy and intensity prediction mean squared error as the core evaluation metrics. If the validation set metrics show no improvement for three consecutive iterations, a learning rate decay mechanism is activated, halving the learning rate. If the metrics improve by more than a preset threshold, the current model parameters are saved as the optimal intermediate model. For example, during training, when the iteration reaches the 35th iteration, the validation set orientation prediction accuracy improves from the initial 68% to 92%, and the intensity prediction mean squared error decreases to 0.03. At this point, the parameters for that round are saved, and iteration continues.

[0090] In the model convergence and solidification phase, training stops when the number of iterations reaches the upper limit or the overall deviation of the dual-dimensional calibration mechanism falls below the preset convergence threshold (0.02). The final model's performance is validated using test set data to ensure that in unfamiliar logistics scenarios (such as logistics channels in the central and western regions), the direction prediction accuracy is no less than 90% and the mean square error of intensity prediction does not exceed 0.05. After successful validation, the model parameters are solidified, completing the construction and training of the association vector model. The final model accurately adapts to the Riemannian manifold topological characteristics of the logistics network. By learning historical cargo flow patterns, it generates an association vector model that truly reflects the node relationships, providing reliable model support for subsequent dynamic correction parameter generation.

[0091] Step 3.4 involves analyzing the changes in the direction and intensity of vectors in the correlation vector model to generate dynamic correction parameters reflecting the evolution of the network topology. Specifically, this includes: first, constructing a real-time monitoring mechanism to continuously track changes in the Riemannian manifold topology corresponding to the correlation vector model, focusing on monitoring three core indicators: Geodesic length calculation uses the Riemannian manifold metric tensor adaptation method, based on the manifold coordinates of core vertices and the metric attributes of the manifold space, to calculate the geodesic length between adjacent core vertices in real time, comparing it with the length data from the previous monitoring period to obtain the length scaling amplitude; manifold curvature calculation uses a local surface fitting method, based on the spatial topological relationship between core vertices and their adjacent nodes, to estimate the curvature value by fitting the curvature of the local manifold surface, and extracting the curvature change value by comparing it with historical curvature data; topological connectivity detection is achieved by continuously comparing the adjacency lists of core vertices. When two core vertices have direct cargo flow records for three consecutive monitoring periods, a new adjacency relationship is determined; when there are no direct cargo flow records and no potential correlation logic for three consecutive monitoring periods, the adjacency relationship is determined to be terminated, accurately capturing the adjustment of topological connectivity relationships.

[0092] Simultaneously, the magnitude fluctuation and directional offset angle of each associated vector are continuously tracked to construct a cooperative coupling coefficient calculation mechanism: First, the changes in vector features (magnitude, direction) and manifold topological features (geodesic length, curvature) within two adjacent monitoring periods are extracted, and the dimensional differences of data in different dimensions are eliminated through normalization; then, based on the correlation analysis of the two types of features in historical data, the weight of vector feature changes is set to 60%, and the weight of manifold topological feature changes is set to 40%; then, the normalized changes are multiplied by their corresponding weights and accumulated to obtain the cooperative coupling coefficient. The coefficient difference between adjacent periods is the rate of change of the cooperative coupling coefficient, which can accurately reflect the degree of cooperative evolution between the vector field and the Riemannian manifold.

[0093] Combining the magnitude of changes in the Riemannian manifold topology with the rate of change in the cooperative coupling coefficient, an evolution trend and magnitude analysis process is initiated: A time synchronization verification mechanism is constructed using a timestamp alignment method, unifying the time dimension of vector field, manifold topology change data, node operation efficiency data, channel traffic status data, and logistics demand change data to a monitoring granularity of 5 minutes / time. By calculating the time difference of the peak values ​​of various data changes, if the time difference is less than one monitoring granularity, it is determined to be a synchronous change. Based on the synchronization results, the core driving factors of topology evolution are determined. For example, if the decrease in node operation efficiency data is synchronized with the change in local topological curvature and there are no other data anomalies, the driving factor is determined to be a change in node operation efficiency; if the channel congestion data is synchronized with the expansion and contraction of geodesic length, it is determined to be an evolution caused by the adjustment of channel traffic status; if the logistics demand transfer data is synchronized with the vector direction offset, it is determined to be an evolution caused by the transfer of logistics demand.

[0094] Based on the above analysis results, the key indicator extraction process was initiated: To quantify the degree of topological evolution, the normalized geodesic stretching and curvature changes were weighted and summed at a 50% ratio to obtain the manifold topological evolution coefficient; a larger coefficient indicates more significant topological evolution. To correct data spatiotemporal bias, the vector field offset correction value was set segmented according to the directional offset angle: 0.1 for offset angles from 0 to 15 degrees, 0.3 for 15 to 30 degrees, and 0.5 for angles above 30 degrees, ensuring the correction value matches the degree of offset. The node association strength attenuation factor was set based on the magnitude fluctuation amplitude; the attenuation factor was set to 0.8 when the fluctuation amplitude was greater than 20%, and 1.0 when it was less than or equal to 20%, accurately reflecting the stability of the association strength.

[0095] Finally, a dynamic correction parameter integration mechanism is constructed: First, the indicator weights are set according to the business characteristics of different evolution scenarios. In the node operation adjustment scenario, the manifold topology evolution coefficient is weighted at 50%, the vector field offset correction value at 30%, and the node association strength attenuation factor at 20%. In the channel state change scenario, the manifold topology evolution coefficient is weighted at 40%, the vector field offset correction value at 50%, and the attenuation factor at 10%. In the demand transfer scenario, the manifold topology evolution coefficient is weighted at 30%, the vector field offset correction value at 20%, and the attenuation factor at 50%. Then, each key indicator is multiplied by the corresponding scenario weight, and the initial dynamic correction parameters are obtained by summing them. Finally, it is verified whether the parameters are within a reasonable range of (0.1, 1.0). If they exceed the range, they are corrected according to the boundary value to ensure the validity of the parameters. The final dynamic correction parameters can accurately match the dynamic evolution characteristics of the logistics network topology.

[0096] In a preferred embodiment of the present invention, step 4 above may include:

[0097] Step 4.1: Based on predefined industry business rules, automatically verify and determine the logical compliance of the standardized spatiotemporal fusion data to obtain preliminary labeled data. Specifically, this includes: first, combining the core business scenarios of the transportation and logistics industry, constructing a complete predefined industry business rule system through business scenario decomposition, rule sorting, expert verification, and iterative optimization. This system covers three core rule categories: data validity rules, spatiotemporal logic rules, and business association rules. The sorting of data validity rules requires clarifying the reasonable value range and format requirements for each data element. After the sorting is completed, industry experts will conduct multiple rounds of verification to eliminate redundant rules and supplement rules for special scenarios. The spatiotemporal logic rules need to combine the topological characteristics of the logistics network to define the normal spatiotemporal association range between different modes of transportation and different nodes. The business association rules are based on typical business processes such as multimodal transport and hub operations to standardize the matching relationship of cross-type data.

[0098] Based on this rule system, a progressive automatic verification is carried out on the standardized spatiotemporal fusion data: The first step involves checking each field according to data validity rules, using a combination of field value comparison and format verification to mark data with values ​​outside the range or with non-standard formats. The second step verifies the spatiotemporal continuity of the data through spatiotemporal logic rules, using the normal time range and flow path constraints between logistics nodes to determine whether the location data of the same goods at different times conforms to the actual flow pattern. The third step verifies the matching of cross-type data based on business association rules, checking the adaptability of cargo types, transportation modes, and operation nodes in multimodal transport data. After verification, a logical compliance judgment is made, using a classification labeling + detailed explanation method to label the data into four categories: compliant, data abnormal, spatiotemporal logic abnormal, and business association abnormal. Simultaneously, the location of abnormal fields, the type of abnormality, and the preliminary judgment reason are accurately recorded, forming preliminary labeled data with complete fields and clear labeling.

[0099] Step 4.2 involves using a pre-trained artificial intelligence model to classify and identify complex patterns and impute missing values ​​in the initially labeled data, generating AI-enhanced labeled data. Specifically, the pre-trained artificial intelligence model used in this step is an improved spatiotemporal fusion Transformer enhancement model: based on the self-attention mechanism of the standard Transformer, it is adapted to the spatiotemporal and topological characteristics of logistics data, adding a spatiotemporal attention layer and a topological adaptation layer. This solves the problems of traditional Transformer's inaccurate capture of spatiotemporal sequence data and difficulty in adapting to the topological relationships of logistics networks, accurately matching the labeling requirements of logistics data in this invention. The specific construction and training process is as follows:

[0100] The model construction process needs to be advanced step by step based on the characteristics of logistics data. First, the model architecture is adapted and adjusted. Using the standard Transformer as the basic framework, a spatiotemporal attention layer is added between the encoder input layer and the self-attention layer. By introducing a time decay factor and spatial distance weights, the model's ability to capture spatiotemporal sequence dependencies is strengthened. A topology association adaptation layer is added after the encoder output layer, incorporating the logistics network topology features generated in step 3, enabling the model to fully utilize node relationships to assist in annotation. Next, input feature construction is carried out. Using a single record of the initially labeled data as a unit, an input feature set containing three core features is constructed: first, basic data features, i.e., the original information such as the value and format of data elements; second, spatiotemporal context features, i.e., the time slice, spatial grid coordinates, and associated data of adjacent spatiotemporal units corresponding to the current data; and third, topology association features, i.e., the attributes, weights, and adjacent node association information of the corresponding key topology feature points. In terms of core layer design, the spatiotemporal attention layer aggregates key spatiotemporal information by calculating the attention weights of data at different spatiotemporal locations; the topology association adaptation layer transforms topology features into feature vectors, which are then fused with the encoder output features to improve the model's ability to perceive the topology associations of the logistics network. The output layer adopts a dual-branch structure, which respectively realizes the complex pattern classification and recognition and the missing value imputation. The classification branch outputs the probability distribution of the anomaly type, and the imputation branch outputs the prediction result of the missing value.

[0101] The model training process must be carried out systematically based on data from typical logistics scenarios. In the training data preparation phase, massive historical spatiotemporal fusion data from three typical logistics scenarios nationwide (multimodal transport channels, regional logistics hub clusters, and inter-regional trunk lines) were selected as the training foundation. The data covers normal data and various abnormal data from January 2023 to June 2024; it was divided into training, validation, and test sets in a 7:2:1 ratio. The training set labels were generated using a combination of manual annotation and real records from the business system, including labels for abnormal types and missing value true values. In the model initialization and hyperparameter setting phase, the Xavier initialization method was used to initialize the parameters of each layer to ensure uniform parameter distribution. Hyperparameters were set as follows: batch size of 64, initial learning rate of 0.001, maximum number of iterations of 80, number of spatiotemporal attention heads of 8, and initial value of topological feature fusion weight of 0.3.

[0102] The training process adopts a phased model of pre-training and fine-tuning: In the pre-training phase, the model is trained using massive amounts of general logistics historical data to acquire basic logistics data pattern recognition capabilities; in the fine-tuning phase, the model is trained using preliminary labeled data and corresponding labels specific to the present invention's scenario, and the model parameters are optimized through gradient descent to adapt to the labeling requirements of the current data. In the loss calibration phase, a multi-task loss mechanism of classification loss and imputation loss is employed. Classification loss constrains the accuracy of anomaly type identification, while imputation loss controls the bias in missing value prediction. After every 5 iterations, the model performance is evaluated using a validation set, with anomaly identification accuracy and missing value imputation error as the core indicators. If the indicators show no improvement for three consecutive iterations, a learning rate decay mechanism is initiated; if the indicators meet the standards, the optimal model parameters are saved. When the number of iterations reaches the upper limit or the loss value falls below the preset convergence threshold, training stops, and the model performance is verified using a test set to ensure that the anomaly identification accuracy is not less than 92% and the imputation error is below the preset threshold. After successful verification, the model parameters are solidified.

[0103] After model deployment, the initially labeled data is input into the trained model according to scenario classification. The model first performs deep classification and identification of complex patterns that were not clearly identified in the initial labeling, such as identifying process anomalies in multimodal transport transshipment, path deviation anomalies in cross-regional transfers, and sudden traffic fluctuation anomalies at hub nodes, supplementing and refining anomaly type labels. Then, for missing values ​​in the data, it performs accurate imputation by combining the spatiotemporal correlation and topological correlation characteristics of the data. For example, when the traffic data of a key topological feature point is missing for a certain period of time, the model will integrate data from adjacent time periods, adjacent node data, and topological correlation features to generate imputed values. After imputation, the model performs a second verification of the initially labeled anomaly labels, correcting misjudged labels, supplementing label details, and finally generating AI-enhanced labeled data.

[0104] Step 4.3: For complex anomaly cases in AI-enhanced labeled data, integrate domain expert knowledge to review and correct them, forming expert-verified labeled data. This includes: first, building a systematic expert knowledge integration mechanism, selecting experts with over 5 years of experience in the transportation and logistics industry, covering sub-sectors such as multimodal transport, hub operation, and trunk transportation to form a review group; through expert interviews, historical case compilation, implicit rule extraction, and rule standardization, forming an expert knowledge system, including judgment criteria for complex anomaly scenarios, special data logic under different business scenarios, solutions for typical historical anomaly cases, and key points for identifying scenarios that AI models are prone to misjudging, etc. This knowledge is compiled into a standardized review manual to provide a unified basis for the review work; based on this mechanism, complex anomaly cases are screened from the AI-enhanced labeled data. The screening process is as follows: first, extract the label credibility score output by the model, set a credibility threshold, and determine it in combination with the performance of the validation set, screening out cases with credibility below the threshold; then, conduct preliminary manual screening to supplement and screen out composite anomaly cases with large data gaps, spanning multiple business scenarios, and ambiguous anomaly features.

[0105] The review panel conducts tiered reviews of the selected cases: First, experts in the corresponding sub-fields conduct a single-expert initial review. These experts, combining the review manual, their own business experience, and historical data from similar cases, verify the accuracy of the AI ​​annotation results for each case, correcting misjudged anomalies and supplementing the annotation basis for cases without clear annotations. For missing values ​​in AI imputation, experts check the reasonableness of the imputation results by reviewing similar data from the same period and business logic constraints. If deviations exist, reasonable values ​​are redefined, and the reasons for the adjustment are explained. Second, a multi-expert review is conducted. For cases with disputes in the initial review, three or more experts in relevant fields are organized to discuss and form a final review conclusion based on the opinions of all parties. After the review and correction are completed, a mapping relationship between expert opinions and data adjustments is established. Anomaly labels and missing values ​​in the annotated data are updated according to the expert conclusions. Simultaneously, expert review opinions and adjustment basis are recorded to form expert-verified annotated data, ensuring the accuracy and reasonableness of the annotations for complex anomaly cases.

[0106] Step 4.4 involves using dynamic correction parameters to correct the spatiotemporal logical consistency and topological correlation in the expert-verified labeled data, obtaining optimized data with quality labels. Specifically, this includes: calling the generated dynamic correction parameters and performing precise correction on the expert-verified labeled data according to the process of spatiotemporal logical consistency correction, topological correlation correction, and verification iteration; in the spatiotemporal logical consistency correction stage, a timestamp calibration + spatial coordinate correction method is adopted: based on the vector field offset correction value in the dynamic correction parameters, the timestamps of cargo flow data with time offsets are adjusted, the deviation between the actual data flow time and the labeled time is calculated, and the timestamps are gradually corrected according to the correction value to ensure that the time dimension is consistent with the actual cargo flow sequence; for data with slight spatial location deviations, the spatial coordinates corresponding to the data are adjusted by combining Riemannian manifold topological features and correction parameters to ensure that the spatial location accurately matches the actual node distribution of the logistics network.

[0107] In the topology correlation correction stage, based on the manifold topology evolution coefficient and node correlation strength attenuation factor in the dynamic correction parameters, a two-step verification and adjustment is carried out: The first step verifies whether the node correlation strength reflected by the data matches the topology correlation strength quantified by the correction parameters. By comparing the cargo flow volume, flow frequency and correlation strength attenuation factor between nodes, it is determined whether the data conforms to the current topology evolution trend. The second step adjusts the value or correlation relationship of the flow data according to the correction parameters for mismatched data. For example, for data where the correlation strength is underestimated, the flow volume value is appropriately corrected, and for flow data that contradicts the topology evolution trend, its associated node labels are adjusted.

[0108] After calibration, a mechanism of full verification, deviation judgment, and iterative adjustment is adopted to verify the calibration effect: a full spatiotemporal logical coherence and topological correlation matching check is carried out on the calibrated data, and the deviation between the data and the actual logistics network status is calculated; if the deviation exceeds the preset threshold, the calibration range is adjusted according to the cause of the deviation, and calibration is carried out again; if the deviation is below the threshold, the calibration is completed; finally, optimized data with quality labeling is obtained, which has accuracy, spatiotemporal consistency and topological correlation, and can be directly used for subsequent logistics network operation status analysis, business decision support and model optimization training.

[0109] In a preferred embodiment of the present invention, step 5 above may include:

[0110] Step 5.1: Extract the spatiotemporal coordinates and quality labels from the optimized quality-labeled data, analyze their distribution characteristics in the preset key logistics areas, and obtain spatial distribution characteristic data. Specifically, this includes: first, confirming the delineation criteria for the preset key logistics areas; combining the distribution of key topological feature points determined in Step 3 and the core scenario requirements of logistics business, delineating the preset key logistics areas, covering types such as core hub operation areas, multimodal transport transshipment connection areas, and key road section coverage areas of trunk corridors; simultaneously determining the spatial boundary range of each area and the corresponding Riemannian manifold coordinate range; from the optimized quality-labeled data, using the field association extraction + spatiotemporal matching method, extract the spatiotemporal coordinates and quality labels corresponding to each data point: the spatiotemporal coordinates include the Riemannian manifold corresponding to the data. The x-axis and y-axis coordinates and time slice information, along with quality labels including core information such as data compliance annotations, anomaly type annotations, and data integrity annotations, are used to match the extracted spatiotemporal coordinates with the spatial boundaries of preset key logistics areas to determine the key logistics area to which each data point belongs. Subsequently, for each key logistics area, methods such as density statistics, distribution uniformity analysis, and time period distribution characteristic analysis are used to analyze the data distribution patterns within that area. For example, the data coverage density at different time periods, the distribution ratio of anomaly labels within the area, and the spatial clustering range of compliant data are statistically analyzed. The distribution analysis results of all key logistics areas are integrated to form spatial distribution characteristic data containing information such as regional data density, time period coverage characteristics, label distribution ratio, and clustering area location.

[0111] Step 5.2: Based on predefined key area weights and spatial distribution characteristic data, perform an adaptive coverage integrity assessment under spatiotemporal constraints to generate spatial coverage quality indicators. Specifically, this includes: first, constructing a predefined key area weight system; combining the importance and priority of logistics operations; and determining the weight values ​​of each predefined key logistics area through business impact assessment and expert review: core business areas such as core hub operation areas and multimodal transport transshipment connection areas are set to high weights; supporting areas such as key trunk line coverage areas are set to medium weights; and peripheral auxiliary logistics areas are set to low weights. The weight values ​​are standardized to ensure a sum of 1. Based on this weight system and the obtained spatial distribution characteristic data, conduct an adaptive coverage integrity assessment under spatiotemporal constraints: first, set spatiotemporal constraints, with the time constraint being coverage of all business operations. The system is divided into segments, such as daytime peak operating hours and nighttime transportation periods, with spatial constraints defined by the core functional areas of each key region. Subsequently, the coverage integrity of each region is calculated using weighted averages based on regional weights. Evaluation indicators include spatial coverage compliance rate (the ratio of actual coverage area to the preset area range), time-period coverage integrity rate (the ratio of actual covered time period to the preset time period), and data density uniformity (the degree of difference in data density among sub-units within the region). The weights of the evaluation indicators are adaptively adjusted according to the business characteristics of different regions. For example, the weight of time-period coverage integrity rate is prioritized in core hub areas, while the weight of spatial coverage compliance rate is prioritized in trunk corridor areas. The weighted coverage integrity indicators of each region are integrated to generate spatial coverage quality indicators reflecting the overall data coverage quality, including overall coverage compliance rate, key area coverage excellence rate, and spatiotemporal coverage balance coefficient.

[0112] Step 5.3: Combining spatial coverage quality indicators with the quality labels of the optimized quality-annotated data, a comprehensive quality assessment and grading of the data is performed, generating data quality grading results. Specifically, this includes: constructing a two-dimensional comprehensive quality assessment system of spatial coverage quality and data quality itself, clarifying the assessment indicators and weight allocation rules for each dimension: spatial coverage quality accounts for 40% of the weight, and data quality itself accounts for 60%; the spatial coverage quality dimension is based on the spatial coverage quality indicators generated in Step 5.2, including overall coverage compliance rate, excellent coverage rate of key areas, etc.; the data quality itself dimension is based on the quality labels of the optimized quality-annotated data, and the assessment indicators include the proportion of compliant data, the accuracy rate of abnormal labels, and data completeness (the proportion of data without missing values), etc.; a weighted summation method is used to calculate the weighted summation of each data point and the overall data. The overall quality score of the data is obtained by first standardizing the indicators of each dimension and converting them into scores from 0 to 100, and then weighting them according to the set weights. Based on the overall quality score, quality levels are determined. The level classification standards are determined after expert review and calibration: an overall score of 90 or above is excellent, 75 to 89 is good, 60 to 74 is acceptable, and below 60 is unacceptable. During the level classification process, boundary scores such as 60, 75, and 90 are subject to secondary verification, and the level is adjusted according to the quality requirements of specific business scenarios. The basis for the level classification of each data point is recorded, including the specific scores of each evaluation indicator and explanations of deduction items. Finally, the data quality level classification results are generated, clearly defining the quality level of each data point, the overall distribution of each level, and the distribution of different quality levels in each key area.

[0113] Step 5.4: Based on the data quality level classification results, store the data into the corresponding thematic databases to generate a hierarchical thematic database. Specifically, this includes: constructing a hierarchical thematic database system based on business themes and data quality levels. Business themes include multimodal transport data themes, hub operation data themes, trunk line transport data themes, and logistics node-related data themes, etc. Each business theme is further divided into four sub-databases according to quality level: excellent, good, acceptable, and unacceptable. For the obtained data quality level classification results, preprocessing is performed before data entry: data at different levels is standardized in format, i.e., data field naming and encoding formats are unified; related information is supplemented, i.e., quality level labels, evaluation criteria, and information on the key regions to which the data belongs are added; for unacceptable data, the problem type and improvement suggestions are additionally marked. To facilitate subsequent data governance, the preprocessed data streams are stored in corresponding sub-databases according to the principle of matching business themes and classifying by quality level: excellent data is stored in the excellent database of the corresponding theme for high-precision business analysis and decision support; good and qualified data are stored in the good and qualified databases of the corresponding themes for routine business monitoring and model training; and unqualified data is stored in the unqualified database of the corresponding theme for data quality optimization analysis and problem tracing only. After the data is stored, an index system for the hierarchical theme database is constructed, including indexes such as quality level, business theme, spatiotemporal range, and data type, to facilitate quick subsequent queries and calls. At the same time, a database storage report is generated, recording information such as the data volume, data source, and quality overview of each theme and level, ultimately forming a hierarchical theme database with a clear structure and well-defined levels.

[0114] In a preferred embodiment of the present invention, step 6 above may include:

[0115] Step 6.1: Extract the data to be published from the hierarchical thematic database and encapsulate it into a data packet conforming to the preset service interface format to obtain the data packet to be published. Specifically, this includes: first, confirming the extraction rules for the data to be published; then, combining the actual application needs of logistics operations, extracting data through a three-dimensional filtering method based on business theme selection, quality level matching, and spatiotemporal range limitation by connecting to the index system of the hierarchical thematic database. Business theme selection can accurately match specific application scenarios such as multimodal transport analysis, hub operation monitoring, and trunk line transportation scheduling. Quality level matching selects data of excellent, good, or qualified levels under the corresponding theme according to the application's requirements for data accuracy. Spatiotemporal range limitation clarifies the time interval and geographical area covered by the data. After extraction, preprocessing is performed on the data before publication, using methods such as... The method of re-validation, format unification, and logical consistency verification first removes duplicate records through the unique data identifier field, then unifies the data encoding format and field naming conventions to ensure consistency with the data format within the hierarchical thematic database, and finally verifies the spatiotemporal logic and topological correlation of the data to eliminate invalid data found in the preprocessing. Based on the established standardized data service interface format, a data field mapping mechanism is constructed to disassemble and reassemble the preprocessed raw data according to the interface requirements, and encapsulate it into a data package containing three core modules: data body, quality level metadata, and spatiotemporal attribute description. The interface format pre-sets specifications such as field order, data type, and transmission encoding to ensure that the data package can be compatiblely parsed by downstream application systems, ultimately resulting in a structurally standardized and content-complete data package to be published.

[0116] Step 6.2: Attach version identifiers and data traceability information to the data packets to be released, generating versioned release packages. This includes: establishing a version identifier generation and management mechanism. Version identifiers are generated using a combination of timestamps, business topic codes, and version numbers. The timestamps are accurate to the second and are used to mark the time node of data release. The business topic codes correspond to the business topic classifications in the hierarchical topic database; for example, the multimodal transport topic code is MSLY, and the hub operation topic code is SNZY. The version number increases with the number of releases of the same topic and the same time dimension, with the initial version being V1.0. Simultaneously, organize the full-link data traceability information of the data packets to be released, forming a standardized traceability information list. The list includes the original... The initial data sources include logistics enterprise business systems and traffic monitoring equipment. The entire data processing record includes the quality labeling process in step 4, the grading basis in step 5, the specific sub-database information of the hierarchical subject database to which the data belongs, and the core indicator scores of the quality assessment. Using metadata association, the generated version identifier and traceability information list are attached to the data package to be released. This is stored as the metadata module of the data package and associated with the business data module, ensuring that the version identifier is unique and traceable and that the traceability information is complete and verifiable. After attachment, the data package is checked for integrity, verifying the standardization of the version identifier format, the completeness of the traceability information, and the validity of the association between the metadata and the business data. After the verification is passed, a versioned release package is generated.

[0117] Step 6.3 involves publishing a versioned release package through the data service interface, while simultaneously providing a query service based on data traceability information to obtain the published dataset. Specifically, this includes: first, completing the adaptation preparation for the data service interface; deploying the release channel using an industry-standard RESTful interface architecture; pre-configuring the interface's transmission protocol, data encryption rules, and access control policies; using HTTPS as the transmission protocol to ensure data transmission security; and assigning different access keys based on the type of application entity, such as logistics companies, regulatory departments, and research institutions; adapting and converting the versioned release package according to the interface's transmission requirements to generate a data stream conforming to the interface's data transmission specifications; completing the release through interface calls; and monitoring the data transmission status in real time during the release process, automatically handling any anomalies such as transmission interruptions or data packet loss. A retransmission mechanism is triggered to ensure the stability of the release process. Simultaneously, a query service system is built based on additional traceability information, providing multi-dimensional accurate queries and full-link traceability display. It supports single-condition or multi-condition queries based on fields such as version identifier, business theme, data quality level, spatiotemporal range, and data source. Query results include details of the corresponding versioned release package, a complete traceability information list, and a data quality assessment report. For user convenience, the traceability information is displayed in a chronological order, presenting key information from each stage of the data processing flow. After the release and query services are deployed, interface connectivity testing is conducted to verify the downstream application systems' ability to acquire the released dataset and the response efficiency of the query service. Once the tests are passed, the service is officially opened to the public, providing a release dataset that can be accessed by applications.

[0118] Step 6.4 involves monitoring and collecting feedback information from analysis reports or decision-making applications generated based on the published dataset. This feedback serves as application feedback for the published dataset and specifically includes: constructing a two-way feedback information acquisition mechanism that combines proactive monitoring and passive collection. Proactive monitoring is achieved by connecting to the interface logs of downstream application systems to capture basic information such as the frequency of calls to the published dataset, call scenarios, and data usage scope of the application systems in real time. Simultaneously, it monitors the publication status of analysis reports generated based on the published dataset and the implementation progress of decision-making applications. Passive collection involves setting up feedback entry points on the data service platform, such as online forms, suggestion boards, and interface feedback channels, to receive targeted feedback submitted by application users. The feedback content covers data quality evaluation. The process includes verifying the accuracy of analysis results, evaluating interface user experience, and providing suggestions for feature optimization. Collected feedback information is categorized and organized into archives based on the released dataset version, feedback type, and issue level. Feedback information is precisely linked to the corresponding versioned release package. Feedback types are categorized by data quality, interface service, and application effect, and issue levels are marked as urgent, general, or recommended. Simultaneously, core requests and issues are extracted from the feedback information to form structured application feedback records. These records include the feedback subject, associated dataset version, feedback type, issue description, requests / suggestions, and collection time, providing direct evidence for subsequent optimization and iteration of released datasets and upgrades to data service functions.

[0119] In a preferred embodiment of the present invention, step 7 above may include:

[0120] Step 7.1 involves analyzing application feedback from the released dataset to identify issues such as insufficient data support or decision-making biases. The analysis results include: first, establishing a deep feedback analysis mechanism; based on the structured application feedback records, filtering information directly related to data support and decision-making applications by feedback type, and eliminating feedback solely focused on interface usage experience; second, employing a problem attribution analysis method to compare the filtered feedback information with the quality level, processing flow, and spatiotemporal coverage characteristics of the corresponding version of the released dataset, deeply analyzing the root causes of the problems mentioned in the feedback; third, for issues related to insufficient data support, focusing on verifying whether there are missing data for specific business scenarios, data accuracy not meeting application requirements, or mismatched spatiotemporal coverage; fourth, for decision-making bias issues, comparing the core indicators of the released dataset with actual logistics business operation data to determine whether the bias is caused by data quality defects or insufficient data adaptability to business scenarios; fifth, simultaneously associating the application scenarios and data usage methods of the feedback subjects to clarify the scope and severity of the problem's impact; and finally, forming a structured feedback analysis result, covering core information such as problem type, associated released dataset version, root cause attribution, scope of impact, and corresponding business scenario.

[0121] Step 7.2: Combine the dynamic correction parameters to perform a correlation correction on the feedback analysis results for the logistics network structure and evolution characteristics, and generate the corrected optimization requirements, which specifically include: Extract the core information of the generated dynamic correction parameters, including the manifold topology evolution coefficient, vector field offset correction value, node association strength attenuation factor, etc., and confirm the logistics network structure characteristics and evolution trends represented by these parameters, such as changes in node association relationships, adjustments to the transfer status of core channels, and shifts in logistics demand distribution; Construct a correlation matching mechanism between the feedback analysis results and the logistics network characteristics, compare the problems in the feedback analysis results with the network evolution characteristics reflected by the dynamic correction parameters, and determine whether the problems are caused by the evolution of the logistics network structure; For problems with associations, carry out correlation correction in combination with the dynamic correction parameters, adjust the problem attribution conclusion, and exclude the feedback cognitive biases caused by the natural evolution of the network; For problems unrelated to network evolution, further clarify them as inherent defects in the data collection or processing links; After the correction is completed, transform the feedback analysis results into precise optimization requirements, clarify whether the optimization direction is to supplement specific data sources, optimize the data processing process, or improve certain quality indicators, and at the same time mark the degree of association between the optimization requirements and the logistics network evolution characteristics, and generate a list of corrected optimization requirements.

[0122] Step 7.3: Generate targeted data collection instructions and governance instructions for specific data sources, processing links, or quality indicators according to the corrected optimization requirements, which specifically include: First, disassemble the list of corrected optimization requirements, and refine the optimization requirements into specific goals that can be implemented according to the corresponding relationship between the problem root cause and the responsible link; If the optimization requirement points to insufficient data sources, generate targeted data collection instructions, and the instruction content clearly specifies the type of data source to be collected, such as the monitoring data of specific logistics nodes, the transfer records of a certain type of transportation method, etc. The spatio-temporal range of collection needs to match the business scenario corresponding to the problem, and the collection frequency and accuracy requirements are determined in combination with the application requirements; At the same time, mark the sub-library of the hierarchical thematic database to which the collected data needs to be connected; If the optimization requirement points to defects in the data processing or quality control links, generate targeted governance instructions, which clearly specify the processing link to be optimized, such as data annotation, spatio-temporal correction, quality assessment, etc. The governance goal needs to be quantified, such as improving the abnormal recognition accuracy of a certain type of data to a specific level, and the governance measures need to be tailored to the problem root cause, such as adjusting the parameters of the data annotation model, optimizing the adaptation logic of spatio-temporal correction, etc.; After the instructions are generated, verify their compatibility with the business specifications of each link to ensure that the instructions can be executed by the corresponding link, and finally form a set of targeted instructions including information such as instruction type, execution object, target requirements, spatio-temporal constraints, and associated optimization requirements.

[0123] Step 7.4 involves feeding back targeted data collection and governance instructions to the corresponding data access, processing, or quality control stages to drive the optimization and evolution of the dataset construction process. Specifically, this includes: constructing a full-link optimization-driven mechanism that enables precise instruction distribution, end-to-end execution tracking, dynamic stage updates, and closed-loop effect verification. First, by connecting with the business management systems and process scheduling platforms of each stage (data access, data processing, quality control, etc.), an instruction distribution mapping relationship is established, confirming the receiving modules and responsible entities corresponding to different types of targeted instructions. During instruction distribution, priority identifiers are added simultaneously, categorizing optimization needs into three levels: urgent, routine, and optimization suggestion. Urgent instructions trigger the receiving module's immediate response mechanism, routine instructions are executed according to daily process scheduling, and optimization suggestion instructions are included in phased iteration plans. Simultaneously, instruction interpretation explanations are provided, clearly outlining the instruction objectives, execution standards, and their logical connection to optimization needs, ensuring the recipient accurately understands the execution requirements.

[0124] During the instruction execution tracking phase, a multi-dimensional progress and effect monitoring system is established: For targeted data collection instructions, the access progress of new data sources is tracked in real time through data access logs. Monitoring indicators include data access completion rate, data format compliance rate, and preliminary quality pass rate. A progress report is generated every hour. If access delays or data quality failures occur, an early warning is automatically triggered and pushed to the responsible party to assist in troubleshooting issues such as data source connection failures or deviations in collection parameter settings. For targeted governance instructions, key monitoring points are set in the corresponding data processing or quality control stages to capture the parameter adjustments in the processing flow and the dynamic changes in quality indicators in real time. For example, when optimizing the data annotation process, the real-time fluctuations in the anomaly identification accuracy are monitored to ensure that governance measures are implemented in the preset direction.

[0125] After execution, the startup process is dynamically updated: For targeted data collection instructions, after the new data source is standardized and qualified through quality assessment, it is accurately added to the corresponding sub-database of the hierarchical theme database according to business theme and quality level. At the same time, the database index system and data traceability template are updated to include the collection information of the new data in the traceability chain. For targeted governance instructions, the optimized processing logic, such as the adjustment of data annotation rules, optimization of spatiotemporal correction and adaptation logic, or quality control parameters, are solidified into the core process of the corresponding link. The business operation manual and process specifications are updated simultaneously, and special training is organized for relevant personnel to ensure that the optimized process is executed in a standardized manner.

[0126] Finally, a closed-loop effect verification mechanism is constructed, employing a combination of phased quality comparison and multi-round application feedback verification to evaluate the optimization effectiveness. Phased quality comparison involves extracting samples of the optimized dataset and comparing them with similar datasets from the same period before optimization, focusing on verifying whether issues related to optimization requirements, such as missing data in specific scenarios or insufficient quality and accuracy, have been resolved. Multi-round application feedback verification continuously tracks application feedback from subsequently released datasets, statistically analyzing changes in the frequency of feedback related to optimization issues, and generating an optimization effect verification report. Based on the verification report, the operational efficiency of the optimization-driven mechanism is regularly reviewed. If incomplete optimization or new problems arise, the targeted instruction strategy is adjusted promptly, continuously improving the iterative chain of feedback analysis, requirement correction, instruction execution, and effect verification. This achieves dynamic adaptation of the dataset construction process to logistics business needs and network evolution characteristics, driving the continuous optimization and evolution of the entire dataset construction system.

[0127] like Figure 2 As shown, embodiments of the present invention also provide a system for dynamically constructing high-quality traffic and logistics datasets, comprising:

[0128] The data caching module is used to access data from multiple authoritative data sources and perform secure caching to obtain data in the secure cache area;

[0129] The fusion processing module is used to perform fusion and standardization processing on the data in the security cache based on industry indicators, forming standardized data with spatiotemporal fusion.

[0130] The model building module is used to analyze the spatiotemporal distribution of standardized data from spatiotemporal fusion, determine key topological feature points of the logistics network, where key topological feature points correspond to the core operating areas, multimodal transport transshipment connection points, and bottleneck locations of main channels within the logistics hub; construct an association vector model based on the key topological feature points, and generate dynamic correction parameters reflecting the evolution of the network topology based on the association vector model;

[0131] The optimization and correction module is used to perform quality control and labeling on standardized spatiotemporal fusion data based on business rules and artificial intelligence, and to use dynamic correction parameters to correct the data labeling and repair process to obtain optimized data with quality labels.

[0132] The assessment and grading module is used to perform quality assessment and grading of the optimized quality-labeled data for thematic applications, and to store data of different quality levels into the corresponding thematic database to generate a graded thematic database.

[0133] The data publishing and feedback collection module is used to dynamically publish data from the hierarchical subject database in a versioned manner and provide traceable data services, obtain the published dataset, and collect application feedback on the published dataset.

[0134] The data optimization and evolution module is used to generate targeted data requirements and governance instructions based on application feedback and dynamic correction parameters of the released dataset, thereby driving the optimization and evolution of the dataset.

[0135] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A traffic logistics high-quality data set dynamic construction method, characterized in that, The method includes: Data from multiple authoritative data sources is accessed and securely cached to obtain data in the secure cache area; Data elements related to three key indicators—vehicle travel time, freight origin and destination, and channel flow—are extracted from the secure cache data. Alignment and association operations are performed on these data elements to obtain initial fused data. Using a unified spatial grid and time reference, traffic flow, speed information, and freight trajectory information from the initial fused data are integrated to generate logistics traceability chain data. Predefined standardization rules are invoked to standardize the data elements, formats, and codes in the logistics traceability chain data. These predefined standardization rules are formulated based on the standards of the Transportation Government Data Resource Catalog and the Comprehensive Transportation Master Data Catalog, resulting in spatiotemporally fused standardized data after the conversion. Standardized spatiotemporal fusion data is analyzed to identify the aggregation and flow patterns of logistics activities in the spatiotemporal dimension, thus obtaining the spatiotemporal distribution characteristics of logistics activities. Based on the spatiotemporal distribution characteristics of logistics activities, corresponding key topological feature points are determined in the core operating area, multimodal transport transshipment connection points, and bottleneck locations of main channels within the logistics hub. Using these key topological feature points as nodes, correlation vectors characterizing the strength and direction of logistics connections between nodes are constructed, and a correlation vector model is formed from all correlation vectors. By analyzing the changes in the direction and intensity of vectors in the correlation vector model, dynamic correction parameters reflecting the evolution of the network topology are generated. Based on predefined industry business rules, the standardized spatiotemporal fusion data is automatically verified and its logical compliance is judged to obtain preliminary labeled data. A pre-trained artificial intelligence model is used to classify and identify complex patterns in the preliminary labeled data and impute missing values ​​to generate AI-enhanced labeled data. For complex anomalies in the AI-enhanced labeled data, domain expert knowledge is integrated to review and correct them to form expert-verified labeled data. Through dynamic correction parameters, the spatiotemporal logical consistency and topological correlation in the expert-verified labeled data are corrected to obtain optimized data with quality labels. The optimized data with quality labels is subjected to quality assessment and classification for thematic applications, and the data of different quality levels are stored in the corresponding thematic database to generate a graded thematic database. The data in the hierarchical subject database is dynamically released in a versioned manner and provided with traceable data services to obtain the released dataset, while collecting application feedback on the released dataset; Based on application feedback and dynamic correction parameters of the released dataset, targeted data requirements and governance instructions are generated in reverse to drive the optimization and evolution of the dataset.

2. The method according to claim 1, wherein, The optimized, quality-labeled data undergoes a subject-specific quality assessment and grading, and data of different quality levels are stored in corresponding subject databases to generate a tiered subject database, including: Extract the optimized data with quality labels, analyze its distribution characteristics in the preset key logistics areas, and obtain spatial distribution characteristic data; Based on predefined key area weights and spatial distribution characteristic data, an adaptive coverage integrity assessment under spatiotemporal constraints is performed to generate spatial coverage quality indicators. By combining the spatial coverage quality indicators with the quality labels of the optimized quality-annotated data itself, a comprehensive quality assessment and classification of the data is performed, generating data quality classification results. Based on the data quality level classification results, the data are stored in the corresponding thematic databases to generate a hierarchical thematic database.

3. The method of claim 2, wherein, The data in the hierarchical subject database is dynamically released in a versioned manner and provided with traceable data services to obtain the released dataset. At the same time, application feedback on the released dataset is collected, including: Extract the data to be published from the hierarchical topic database and encapsulate it into a data packet that conforms to the preset service interface format to obtain the data packet to be published; Add version identifiers and data traceability information to the data packages to be released, and generate versioned release packages; Versioned release packages are published through the data service interface, and query services based on data source information are provided to obtain the released dataset; Monitor and collect feedback information from analytical reports or decision-making applications based on published datasets, as application feedback for published datasets.

4. The method for dynamically constructing a high-quality transportation and logistics dataset according to claim 3, characterized in that, Based on application feedback and dynamic correction parameters of the released dataset, targeted data requirements and governance instructions are generated in reverse to drive the optimization and evolution of the dataset, including: Analyze application feedback on the published dataset to identify issues such as insufficient data support or decision-making biases, and obtain feedback analysis results; By combining dynamic correction parameters, the correlation between the feedback analysis results and the logistics network structure and evolution characteristics is corrected, and the corrected optimization requirements are generated. Based on the corrected optimization requirements, generate targeted data collection and governance instructions for specific data sources, processing stages, or quality indicators; The targeted data collection and governance instructions are fed back to the corresponding data access, processing, or quality control links to drive the optimization and evolution of the dataset construction process.

5. A system for dynamically constructing high-quality datasets for transportation and logistics, wherein the system implements the method as described in any one of claims 1 to 4, characterized in that, include: The data caching module is used to access data from multiple authoritative data sources and perform secure caching to obtain data in the secure cache area; The fusion processing module is used to perform fusion and standardization processing on the data in the security cache based on industry indicators, forming standardized data with spatiotemporal fusion. The model building module is used to analyze the spatiotemporal distribution of standardized data from spatiotemporal fusion, determine key topological feature points of the logistics network, where key topological feature points correspond to the core operating areas, multimodal transport transshipment connection points, and bottleneck locations of main channels within the logistics hub; construct an association vector model based on the key topological feature points, and generate dynamic correction parameters reflecting the evolution of the network topology based on the association vector model; The optimization and correction module is used to perform quality control and labeling on standardized spatiotemporal fusion data based on business rules and artificial intelligence, and to use dynamic correction parameters to correct the data labeling and repair process to obtain optimized data with quality labels. The assessment and grading module is used to perform quality assessment and grading of the optimized quality-labeled data for topic applications, and to store data of different quality levels into the corresponding topic databases to generate a graded topic database. The data publishing and feedback collection module is used to dynamically publish data from the hierarchical subject database in a versioned manner and provide traceable data services, obtain the published dataset, and collect application feedback on the published dataset. The data optimization and evolution module is used to generate targeted data requirements and governance instructions based on application feedback and dynamic correction parameters of the released dataset, thereby driving the optimization and evolution of the dataset.