Method and system for port group multi-source heterogeneous data fusion and transportation channel digitization

By constructing a unified multi-source heterogeneous data fusion framework and a digital method for transportation channels, the problems of data heterogeneity and model localization in the collaborative operation of port clusters have been solved. This has enabled digital modeling and intelligent collaborative management and control of transportation channels, thereby improving the overall operational efficiency and supply chain resilience of port clusters.

CN122020546APending Publication Date: 2026-05-12TRANSPORT PLANNING & RES INST MINIST OF TRANSPORT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TRANSPORT PLANNING & RES INST MINIST OF TRANSPORT
Filing Date
2026-01-31
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies in port cluster collaborative operations suffer from problems such as difficulty in integrating heterogeneous data, localized optimization models, and fragmented application scenarios, resulting in low efficiency in cross-domain collaboration and global decision-making, and failing to achieve overall optimization and resilience of transportation channels.

Method used

A unified multi-source heterogeneous data fusion framework is constructed. Using the national 2000 geographic coordinate system and UTC time as a unified spatiotemporal benchmark, a metadata model and a data source abundance deduplication algorithm are adopted to achieve data standardization. The Lambda stream-batch integrated architecture is used for multi-frequency data collaborative processing to identify and classify transportation groups whose operational characteristics meet the preset requirements, and to extract and digitize transportation channels.

Benefits of technology

It has achieved semantic unification and real-time collaboration of port cluster data, improved the digital modeling and performance evaluation capabilities of transportation channels, enhanced the overall operational efficiency and supply chain resilience of port clusters, and enabled them to cope with sudden disturbances and dynamic risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020546A_ABST
    Figure CN122020546A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of harbor planning digital intelligence and integrated fusion, and discloses a harbor group multi-source heterogeneous data fusion and transportation channel digitization method and system, and the method comprises the steps: building a systematic data fusion framework based on harbor group multi-source heterogeneous data, and building a systematic data fusion framework based on the systematic data fusion framework. Establishing a unified space-time reference and a data specification, and forming a fused data product; on the basis of the fused data product, constructing a vehicle-ship polymer, and on the basis of the vehicle-ship polymer, identifying and classifying transportation groups with operation characteristics meeting preset requirements; and based on the track data of the transportation group, extracting a transportation channel meeting a preset requirement, and realizing digital and structured expression of the transportation channel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of port planning digitalization and integration, specifically involving a method and system for the fusion of multi-source heterogeneous data of port clusters and the digitalization of transportation channels. Background Technology

[0002] Currently, smart ports have become a core direction for the transformation and upgrading of ports worldwide. Smart technologies within individual ports—such as automated quay crane control, unmanned horizontal transport, and port area digital twin systems—have gradually entered a stage of large-scale, mature application. However, with the increasing demands for "efficient collaboration" and "global resilience" from the global supply chain, port development is facing a strategic transformation from "single-point intelligence" to "port cluster collaboration," and from "internal port optimization" to "global optimization of transport channels." In this process, existing technological systems have revealed the following three significant limitations in supporting cross-domain collaboration and global decision-making: (1) The problem of heterogeneity and information silos at the data level The coordinated operation of port clusters relies on massive amounts of data across entities, dimensions, and modalities, including but not limited to: frequently updated vessel AIS dynamic data, structured data from port operation scheduling systems, spatial vector data of port clusters, and data on rail and road transport connecting the ports. These data exhibit significant differences in structure, temporal sequence, and semantics. Semantic heterogeneity makes alignment difficult: different ports and data sources have different standards. Taking AIS data as an example, key fields (such as "ship arrival time") may have different definitions in different ports or service providers. There is a lack of unified mapping with the port planning data system, which leads to a lot of manual intervention during the integration process, resulting in low efficiency and a high risk of errors.

[0003] Multi-frequency data is difficult to coordinate: High-frequency data (such as AIS, updated in seconds) and low-frequency data (such as spatial planning data, static or updated annually) have different processing paradigms, and existing systems are unable to achieve effective coordination of real-time access and on-demand updates between the two.

[0004] (2) Localization traps and lack of robustness in model and algorithm layers Although data-driven optimization research is constantly emerging, such as using mixed-integer programming and adaptive large neighborhood search algorithms based on AIS data for fleet route planning, or employing genetic algorithms to optimize container truck scheduling within ports, the optimization scope is mostly limited to a single link or a single transportation mode, failing to connect all the links in the "sea transport - port operations - rail transport - road connection". Furthermore, existing models are mostly based on ideal assumptions, lacking robustness and adaptability to common dynamic disturbances in transportation channels (such as extreme weather and equipment failures).

[0005] (3) The fragmentation of application scenarios and the lack of evaluation The current application of smart port technologies is characterized by "fragmented scenarios." For example, in container terminals, digital twins can already achieve good joint scheduling of equipment; however, in bulk cargo terminals, the model granularity is insufficient, and applications are mostly limited to the macro level. More importantly, the existing technology system lacks a unified definition and digital modeling standards for "transportation channels," making it impossible to monitor, comprehensively evaluate, and dynamically extrapolate the overall efficiency, resilience, and sustainability (such as carbon emissions) of the channels in real time.

[0006] In summary, the limitations of existing technologies at the three core levels of "data fusion, algorithm optimization, and scenario application" have created mutually reinforcing technical bottlenecks: heterogeneous data makes it impossible to effectively integrate multi-source information, thus limiting the input quality of the entire chain optimization algorithm; the localization of algorithms and the fragmentation of scenarios make it difficult to implement the global control requirements of transportation channels. These shortcomings directly lead to low collaborative efficiency of port clusters, insufficient overall efficiency of transportation channels, and a lack of resilience in actual operations. Therefore, there is an urgent need in this field for a systematic technical solution that can deeply integrate multi-source heterogeneous data from port clusters and achieve digital modeling and intelligent collaborative control of the entire transportation channel process, thereby effectively improving the overall operational efficiency of port clusters and the resilience of the supply chain in complex and ever-changing environments. Summary of the Invention

[0007] To address the challenges of data heterogeneity and integration, localized optimization models, and fragmented application scenarios in existing technologies, this invention provides a method and system for fusing multi-source heterogeneous data from port clusters and digitizing transportation channels. The aim is to construct a unified data fusion framework that integrates multi-source heterogeneous data, including port infrastructure, dynamic and static attributes of ships and vehicles, and spatial vectors of port clusters, to form a high-quality collaborative data foundation. Based on this, an innovative "vehicle-ship aggregation" analysis unit is proposed. By identifying and classifying groups of ships, vehicles, and trains with similar operational characteristics, their operational trajectory features are extracted, accurately extracting and digitally representing the macroscopic patterns of major maritime transportation channels. This provides accurate and reliable data support and decision-making basis for achieving coordinated scheduling of port clusters and overall optimization of transportation channels.

[0008] To achieve the above objectives, the present invention provides the following solution: A method for fusing multi-source heterogeneous data from port clusters and digitizing transportation channels, the method comprising: Based on multi-source heterogeneous data from port clusters, a systematic data fusion framework is constructed. Based on the systematic data fusion framework, a unified spatiotemporal benchmark and data specifications are established to form fused data products. Based on the fused data product, a "vehicle-vehicle aggregation" is constructed. Based on the "vehicle-vehicle aggregation", transportation groups whose operational characteristics meet preset requirements are identified and classified. Based on the trajectory data of transportation groups, transportation channels that meet preset requirements are extracted, enabling the digital and structured representation of transportation channels.

[0009] Preferably, the method for constructing the systematic data fusion framework includes: Establish the National Geographic Coordinate System 2000 and UTC time as a unified spatiotemporal reference; Based on a unified spatiotemporal benchmark, a metadata model for port cluster business is constructed, and the semantic definition of key business fields is unified. Based on the unified semantic definition of key business fields, a deduplication algorithm based on data source abundance is used to merge duplicate data from multiple sources. Based on the Lambda stream-batch integrated architecture, collaborative processing of multi-frequency data after merging duplicate points is realized, thus completing the construction of the systematic data fusion framework.

[0010] Preferably, the metadata model uses "entity-field-term-rule" as its core logic, defining entities such as ships, container trucks, railway trains, and port facilities. Each entity is associated with "mandatory core fields + optional extended fields," forming a standardized data dictionary. The expression is as follows: ;in, It is a collection of core entities including ships, container trucks, railway trains, and port facilities; For each entity, it is a collection of fields; A standard terminology set for entity fields; A set of mapping rules from standard terms to standard terms; U Iterative update rules for the basic information table For numeric fields, for Entity fields, for Entity mapping rules for Standard terminology for entities.

[0011] Preferably, the "vehicle-ship aggregation" is a four-dimensional structured unit containing "core members, association rules, business attributes, and hierarchical identifiers," mathematically expressed as: ; Where M is the core member set, containing all instances of transportation vehicles, categorized by type, and formally expressed as: ,in For ship crew, For the i-th ship member, For container truck crew, For the j-th container truck member, As a member of the railway freight train, For the first Members of the railway freight train; R For the set of association rules, define the constraints between members based on "spatiotemporal collaboration + business association", which can be expressed as follows: Among them, spatiotemporal coordination rules Business association rules are defined as follows: The difference between the ship's arrival time and the truck / train transfer time is ≤Δt, and the spatial distance between the transfer point and the port berth / hub is ≤Δs. To allow members to share the same container number, transport task number, or service port cluster: , in, For port cluster identification, , For container number, , To serve the port cluster, , For transportation tasks; A represents the set of core business attributes, symbolizing the overall functional characteristics of the aggregate. It is generated by aggregating member characteristics and can be expressed as a formula: In this context, RouteScope represents the route's coverage area, CargoType represents the type of cargo transported, TransMode represents the intermodal transport mode, and Capacity represents the total transport capacity. L The hierarchy is categorized by functional scale, with the following levels: Ocean Trunk Line Cluster_A: Centered on 10,000-15,000 TEU large container ships, supported by 40-foot heavy-duty highway trucks and 50-60 car trainsets, undertaking intercontinental container transport; Feeder Line Cluster_B: Centered on 500-3,000 TEU medium-sized container ships, supported by 20-30-foot feeder trucks and 30-40 car trainsets, undertaking container transshipment between core ports and surrounding feeder ports; River-Sea Direct Cluster_C: Centered on 2,000-5,000 DWT river-sea direct container ships, supported by 15-20-foot small highway trucks, without railway support, undertaking direct port transport.

[0012] Preferably, the method for identifying and classifying transport groups whose operational characteristics meet preset requirements based on the "vehicle-vehicle aggregation" includes: Extract multidimensional features of transportation entities from fused data, including spatiotemporal features, business attribute features, and behavioral pattern features; The DBSCAN density clustering algorithm is used to perform unsupervised learning on the multidimensional features to achieve automatic identification and classification of transport groups.

[0013] Preferably, the method for extracting transportation channels that meet preset requirements includes: Using the historical trajectory of the "vehicle-ship aggregation" as input, a trajectory density clustering algorithm is used to identify frequently used path corridors; The identified path corridors are abstracted into a digital network model, and each corridor is assigned basic attributes, dynamic performance attributes, and sustainability attributes.

[0014] The present invention also provides a system for fusing multi-source heterogeneous data of port clusters and digitizing transportation channels. The system is used to implement the aforementioned method and includes: a fusion module, a classification module, and an extraction module. The fusion module is used to construct a systematic data fusion framework based on multi-source heterogeneous data of the port cluster, and based on the systematic data fusion framework, to establish a unified spatiotemporal benchmark and data specifications to form fused data products. The classification module is used to construct a "vehicle-vehicle aggregation" based on the fused data product, and to identify and classify transportation groups whose operational characteristics meet preset requirements based on the "vehicle-vehicle aggregation". The extraction module is used to extract transportation channels that meet preset requirements based on the trajectory data of transportation groups, thereby realizing the digital and structured representation of transportation channels.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention provides a method for fusing multi-source heterogeneous data from port clusters and digitizing transportation channels. It constructs a systematic fusion framework based on multi-source heterogeneous data to achieve semantic unification and real-time collaboration of port cluster data. Based on the fused data, it innovatively constructs a meso-level analysis unit of "vehicle-ship aggregation," achieving a leap from micro-level individual behavior to macro-level logistics patterns. Through trajectory mining and cluster analysis, it extracts major transportation channels, completing digital modeling and performance evaluation of these channels. This invention constructs an intelligent analysis system for port cluster transportation channels covering the entire chain of "data fusion—group identification—channel representation," effectively enhancing the utilization value and collaborative analysis capabilities of multi-source data, optimizing the allocation of transportation channel resources, ensuring the overall operational efficiency and supply chain resilience of port clusters, and strengthening the ability to adapt to sudden disturbances and dynamic risks. Attached Figure Description

[0016] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1This is a schematic diagram of a method for fusing multi-source heterogeneous data of port clusters and digitizing transportation channels according to an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] Example 1 like Figure 1 As shown, this method provides a way to fuse multi-source heterogeneous data from port clusters and digitize transportation channels. It revolves around three core stages: "data fusion - aggregate analysis - channel digitization," forming a complete technical chain from underlying data governance to top-level decision support, including: Based on multi-source heterogeneous data from port clusters, a systematic data fusion framework is constructed. Based on the systematic data fusion framework, a unified spatiotemporal benchmark and data specifications are established to form fused data products. Based on the fused data product, a "vehicle-vehicle aggregation" is constructed. Based on the "vehicle-vehicle aggregation", transportation groups whose operational characteristics meet preset requirements are identified and classified. Based on the trajectory data of transportation groups, transportation channels that meet preset requirements are extracted, enabling the digital and structured representation of transportation channels.

[0021] In this embodiment, the method for constructing the systematic data fusion framework includes: Establish the National Geographic Coordinate System 2000 and UTC time as a unified spatiotemporal reference; Based on a unified spatiotemporal benchmark, a metadata model for port cluster business is constructed, and the semantic definition of key business fields is unified. Based on the unified semantic definition of key business fields, a deduplication algorithm based on data source abundance is used to merge duplicate data from multiple sources. Based on the Lambda stream-batch integrated architecture, collaborative processing of multi-frequency data after merging duplicate points is realized, thus completing the construction of the systematic data fusion framework.

[0022] In this embodiment, the metadata model uses "entity-field-term-rule" as its core logic, defining entities such as ships, container trucks, railway trains, and port facilities. Each entity is associated with "mandatory core fields + optional extended fields," forming a standardized data dictionary. The expression is as follows: ;in, It is a collection of core entities including ships, container trucks, railway trains, and port facilities; For each entity, it is a collection of fields; A standard terminology set for entity fields; A set of mapping rules from standard terms to standard terms; U Iterative update rules for the basic information table.

[0023] In this embodiment, the "vehicle-ship aggregation" is a four-dimensional structured unit containing "core members, association rules, business attributes, and hierarchical identifiers," mathematically expressed as: ; Where M is the core member set, containing all instances of transportation vehicles, categorized by type, and formally expressed as: ,in For ship crew, For container truck crew, For members of railway freight trains; R For the set of association rules, define the constraints between members based on "spatiotemporal collaboration + business association", which can be expressed as follows: Among them, spatiotemporal coordination rules Business association rules are defined as follows: The difference between the ship's arrival time and the truck / train transfer time is ≤Δt, and the spatial distance between the transfer point and the port berth / hub is ≤Δs. To allow members to share the same container number, transport task number, or service port cluster: , in, For port cluster identification; A represents the set of core business attributes, symbolizing the overall functional characteristics of the aggregate. It is generated by aggregating member characteristics and can be expressed as a formula: In this context, RouteScope represents the route's coverage area, CargoType represents the type of cargo transported, TransMode represents the intermodal transport mode, and Capacity represents the total transport capacity. LThe hierarchy is categorized by functional scale, with the following levels: Ocean Trunk Line Cluster_A: Centered on 10,000-15,000 TEU large container ships, supported by 40-foot heavy-duty highway trucks and 50-60 car trainsets, undertaking intercontinental container transport; Feeder Line Cluster_B: Centered on 500-3,000 TEU medium-sized container ships, supported by 20-30-foot feeder trucks and 30-40 car trainsets, undertaking container transshipment between core ports and surrounding feeder ports; River-Sea Direct Cluster_C: Centered on 2,000-5,000 DWT river-sea direct container ships, supported by 15-20-foot small highway trucks, without railway support, undertaking direct port transport.

[0024] In this embodiment, the method for identifying and classifying transportation groups whose operational characteristics meet preset requirements based on the "vehicle-vehicle aggregation" includes: Extract multidimensional features of transportation entities from fused data, including spatiotemporal features, business attribute features, and behavioral pattern features; The DBSCAN density clustering algorithm is used to perform unsupervised learning on the multidimensional features to achieve automatic identification and classification of transport groups.

[0025] In this embodiment, the method for extracting transportation channels that meet preset requirements includes: Using the historical trajectory of the "vehicle-ship aggregation" as input, a trajectory density clustering algorithm is used to identify frequently used path corridors; The identified path corridors are abstracted into a digital network model, and each corridor is assigned basic attributes, dynamic performance attributes, and sustainability attributes.

[0026] Example 2 This embodiment uses the identification of a container transport channel network of a port cluster as an example to demonstrate the complete implementation process of the present invention.

[0027] Step 1: Construction of a systematic multi-source heterogeneous data fusion framework This step is the data foundation of the entire methodology, aiming to transform disordered raw data into high-quality, analyzable fused data products. It converts scattered and disordered raw data (ship dynamics, vehicle positioning, operation scheduling, etc.) into fused data products that are "spatiotemporally unified, semantically consistent, and of controllable quality," providing highly reliable data input for subsequent "vehicle-ship aggregation" construction and the digitalization of transportation channels.

[0028] 1.1 Multi-port data access It also accesses AIS data of ships, GIS data of container trucks, container train data, port scheduling operation data, and basic information attribute data of port clusters from ports I, II, III, IV, V, VI, VII, and VIII.

[0029] Ship AIS data, container truck GIS data, and container train data can be accessed in real time through communication and purchasing services, or historical data can be uploaded periodically. Port scheduling operation data and basic information attribute data of port clusters are updated regularly by the operations of each port.

[0030] Ship AIS data includes mandatory fields such as MMSI (Unique Association Key), latitude and longitude in the national 2000 coordinate system, speed, heading, draft, ship type, deadweight tonnage, port of destination, estimated time of arrival (ETA), and data collection time, supplemented by optional fields such as IMO number and navigation status.

[0031] Container truck GIS data: Using the license plate number as the association key, it includes core positioning fields such as latitude and longitude in the national 2000 coordinate system, positioning accuracy, vehicle speed, heading direction, driving status, cumulative and current transport mileage, data collection and upload time, and positioning device information. It also associates fields such as transport task number, container number and type, service port, origin and destination, estimated time of arrival (ETA), approved and actual load, and overload status.

[0032] The container train data includes: using the train's unique number as the associated key, covering key information such as the origin / destination station code and name, a list of intermediate stops, departure / estimated arrival / actual arrival time, train formation type, number of carriages, approved load capacity, number of containers and container number list, and operating status.

[0033] Port scheduling operation data: Associated information includes vessel MMSI, berth number, berthing and departure time, container number and weight, yard location, loading and unloading equipment number, operation type and transshipment method, operation start and end time, etc.

[0034] Basic information and attribute data of port clusters: including remote sensing imagery, port code, name and region, berth list and purpose, waterway parameters and origin and destination, supporting railway freight station and highway logistics park information, yard and loading and unloading equipment parameters, operating entity and other basic fields.

[0035] Define a field validation function to ensure that required fields are not missing. The formula is as follows: Where D represents a single data record. This is a set of required fields for various types of data (such as MMSI, latitude and longitude, and recording time for ship AIS data); format validity checks include: latitude and longitude range checks (longitude 73°E-135°E, latitude 4°N-53°N), time format checks (compliant with ISO 8601 standard), and numerical field range checks (such as ship speed 0-60 knots, vehicle speed 0-120km / h). This is a numeric field.

[0036] By performing access verification and preprocessing, we ensure that the original data is "complete, non-duplicate, and uniform in format".

[0037] The core associated keys (MMSI, license plate number, train unique number, and operation serial number) are deduplicated using the following formula: in, As a key value; This is the serial number.

[0038] The data that passes validation is formatted (e.g., spaces are removed from string fields, numeric fields are rounded to two decimal places, and date fields are formatted as "YYYY-MM-DDTHH:MM:SS") to generate a standardized original dataset. .

[0039] in, For ship AIS data, For port operation data, For railway transportation data, For truck GPS data, This is the basic information attribute data for the port cluster.

[0040] 1.2 Standardization Processing of Multi-Source Data To address the issues of "spatiotemporal misalignment and format chaos" in data from different ports, data standardization is achieved through "spatiotemporal unification": ① Coordinate System One All spatial data (latitude and longitude, remote sensing image coordinates, berth / hub coordinates, etc.) are uniformly converted to the National Geodetic Coordinate System 2000 (CGCS2000). The conversion is divided into two steps: "three-dimensional geodetic coordinate conversion → two-dimensional latitude and longitude extraction".

[0041] Step 1: 3D Geodetic Coordinate Transformation: Transform the original coordinate system's 3D coordinates ( , Convert to CGCS2000 3D coordinates ( , The Bursa-Wolf seven-parameter model is used, and the formula is as follows: in, The translation parameters (unit: meters) are taken from the regional transformation parameters published by the State Bureau of Surveying and Mapping: for example, for a certain region , , . The rotation parameter (unit: radians) has a range of 10. -6 ~10 -5 m is the scale factor (dimensionless), with a value ranging from 10. -6 ~10 -5 If the original data only provides two-dimensional latitude and longitude (without elevation Z), then the default will be used. (Nearshore / land area data errors are negligible).

[0042] Step 2: Two-dimensional latitude and longitude extraction: Convert the CGCS2000 three-dimensional coordinates to latitude and longitude using the CGCS2000 ellipsoid parameters: semi-major axis. Flatness Calculated using the geodetic coordinate forward calculation formula: in, It represents the longitude of the Earth.

[0043] The final output latitude and longitude coordinates are retained to 6 decimal places to ensure that the spatial data alignment error across ports is ≤0.5 meters.

[0044] in, , It represents the flatness.

[0045] ②Time System One All time-series fields (data acquisition time, berthing time, departure time, etc.) are uniformly converted to Coordinated Universal Time (UTC) and standardized to ISO 8601 format (YYYY-MM-DDTHH:MM:SSZ).

[0046] Time zone conversion formula: in, Local time To coordinate time for the world This refers to the time zone difference (unit: hours). The local time at the port is UTC+8. If the original time does not have a time zone identifier, it will be filled in by default according to the time zone of the port where the data comes from (e.g., all ports are processed as UTC+8).

[0047] 1.3 Quantitative Methods for Data Quality Control To eliminate noise, missing data, and duplicates in the original data and ensure the reliability of the fused data, a hierarchical quantization processing strategy was designed.

[0048] ① Data completion Using unique identifiers (MMSI code for ships, license plate number for vehicles, and unique train number) as the association key, and combining them with a pre-set port group metadata model (including a standard terminology set of core fields such as "ship type" and "vehicle load"), data completion and semantic unification are achieved.

[0049] The port cluster metadata model uses "entity-field-term-rule" as its core logic, defining four core entities (ships, container trucks, railway trains, and port facilities). Each entity is associated with "mandatory core fields + optional extended fields," forming a standardized data dictionary. Mathematically, this is expressed as... .

[0050] in, It is a collection of core entities (ships, container trucks, railway trains, and port facilities). This is a collection of fields for each entity (including required and optional fields). for The fields of the entity. A standard terminology set for entity fields. This is a set of mapping rules from standard terms to standard terms. for Entity mapping rules. U is the iterative update rule for the basic information table.

[0051] Define "required core fields + optional extended fields" for each core entity. The required fields are the key keys for data association and completion, as shown in Table 1 below: Table 1 Define a unique set of standard terms for each core field Standard terminology must conform to industry standards (such as IMO ship classification standards and GB / T 22147-2008 freight vehicle classification) to ensure unambiguous semantics across ports.

[0052] Example of a standard terminology set for key fields: Ship type : {"Container ship", "Bulk carrier", "Crude oil tanker", "LNG carrier", "River-sea direct vessel", "Feeder barge"} Truck models :{“40-foot container trucks”, “20-foot container trucks”, “semi-trailer container trucks”, “heavy-duty container trucks”} Homework type {"Loading", "Unloading", "Transit", "Collection", "Returning"} Define a two-way mapping rule for the non-standard representation of the same field in data from different ports. This achieves semantic uniformity. The mapping rules adopt a one-to-one / many-to-one mapping from "non-standard terms to standard terms". The priority of the mapping rules is: industry-general non-standard terms (such as "container ship") > port-defined terms (such as "I port container ship") > vague expressions (such as "large cargo ship", which needs to be combined with other fields for auxiliary judgment).

[0053] Example of key field mapping rules, taking ship type mapping as an example: ("Container ship") = "Container ship" ("Container") = "Container Ship" (“I Port Feeder Container Vessel”) = “Feeder Barge” Taking a multi-port container ship with MMSI 477XXXXXXX as an example: The system detects that the ship's "vessel type" is recorded as "container ship" in the AIS data of port I and as "cargo ship" in the AIS data of port II. Through the terminology mapping rules of the metadata model (the non-standard expression "cargo ship" is "container ship"), it automatically unifies it into the standardized "container ship". At the same time, it uses the MMSI code to associate with the port group's basic information attribute table to complete its static attributes (length 299 meters, beam 40 meters, deadweight tonnage 50,000 tons, draft 12.5 meters).

[0054] To ensure the timeliness and accuracy of the port cluster metadata model, the completed attribute values ​​need to be iteratively updated in reverse to the port cluster basic information attribute table. Simultaneously, to avoid data corruption, version number management is introduced. This provides a data foundation of "semantic consistency, attribute completeness, and traceability" for subsequent multi-source data fusion.

[0055] ② Merging duplicate points Multiple sources: When data from sources A and B are merged, it is inevitable that highly overlapping duplicate data in time and space will be introduced. These duplicates will increase storage and computing overhead and may distort the accuracy of traffic statistics analysis, and must be removed.

[0056] The fused data is first grouped by MMSI / license plate number to obtain the complete trajectory sequence of a single ship / vehicle; then it is judged based on the two dimensions of "maximum likelihood estimation + data source abundance".

[0057] Step 1: Divide the merged data into several independent subsets, each subset corresponding to the full trajectory data of a single boat / vehicle. The grouping function is expressed as: in, Optional parameters for delimitation. k This is the block key. This is the total dataset after multi-source fusion. This is a subset of the trajectory dataset for a single boat / vehicle.

[0058] Step 2: For each group Data points are sorted in ascending order by "UTC timestamp" to generate continuous trajectory sequences for a single ship / vehicle. , For the first The continuous trajectory of a single boat / vehicle ensures the temporal continuity of the trajectory, providing a temporal basis for subsequent determination of repetition points.

[0059] Step 3: Same trajectory sequence In the middle, two data points from different data sources (Data source) )and (Data source) If the conditions of "time difference ≤ Δt and spatial distance ≤ Δs" are met, then it is determined to be a duplicate point pair. , The time threshold Δt can be taken as 30s, and the spatial threshold Δs can be taken as 50m.

[0060] Step 4: Process the trajectory sequence All points identified as duplicates are clustered according to their "duplicate association" to form duplicate point clusters. , For the first Duplicate points (all data points within the same cluster are duplicates of each other, coming from two or more data sources).

[0061] Step 5: For each cluster of repeating points The "historical accuracy (maximum likelihood estimation)" and "in-cluster abundance" of each participating data source are calculated to form a two-dimensional evaluation index.

[0062] Using "authoritative data sources" as the truth benchmark, the historical data matching rate of each data source is calculated through maximum likelihood estimation to quantify the reliability of data quality (the higher the accuracy, the higher the data authenticity). For authoritative data sources, the published data (S0) is selected as the authoritative benchmark for ship AIS data; the data from the freight management system (S0) is selected as the authoritative benchmark for vehicle GPS data; and the data from the railway dispatch center (S0) is selected as the authoritative benchmark for railway train data.

[0063] Select a 3-month historical data sample (sample size) (to ensure statistical significance) for data sourcesS Calculate the number of matching points between it and the authoritative data source S0.

[0064] Statistical data source S In the current repeating point cluster The percentage of data points in the data source quantifies the completeness of the data source's coverage of the current trajectory segment (the higher the percentage, the better the trajectory continuity).

[0065] Step 6: For duplicate point clusters For each data source S, calculate the overall score. Select the data source with the highest overall score. .reserve exist All data points are analyzed, and duplicates from other data sources are removed.

[0066] Step 7: After removing all duplicate data clusters, integrate the remaining data points, reorder them according to UTC timestamps, and generate a "non-repeating, high-quality, continuous" clean trajectory sequence for a single ship / vehicle. Finally, the cleaning trajectories of all groups are aggregated to form a comprehensive cleaning dataset for the port cluster. , For the first k A clean trajectory sequence.

[0067] On one hand, the historical data accuracy of each data source is statistically analyzed (e.g., source A has an accuracy rate of 92% for the past 3 months, while source B has 88%). On the other hand, the number of data points from each data source in the trajectory sequence is calculated (e.g., source A provides 85 points, while source B provides 42 points). Finally, the data points corresponding to the data source with "high accuracy + abundant data points" (source A in this example) are retained, and duplicate points are removed. This method can reduce storage overhead by about 30%, avoid statistical deviations in transportation flow caused by duplicate data (the deviation rate is reduced from 18% to below 5%), and ensure trajectory continuity and analytical accuracy.

[0068] 1.4 Multi-source data collaborative processing The enhanced Lambda architecture uses a unified port cluster database as its storage foundation to achieve real-time perception of high-frequency dynamic data and deep integration of low-frequency static data.

[0069] The unified port cluster database adopts a "hierarchical partitioning + version control" storage architecture, defined as... .

[0070] in, Real-time dynamic data output from the storage stream processing layer (partition key: unique identifier + hourly UTC time partition); Store the fused data output from the batch processing layer and the cleaned historical data (partition key: unique identifier + UTC time day-level partition). Store the port cluster metadata model (standard terminology set, mapping rules, association keys) defined above to provide a unified semantic benchmark for stream and batch processing; Clustering models and anomaly detection thresholds trained in batch processing are stored to support real-time decision-making in stream processing.

[0071] The stream processing layer and the batch processing layer are connected through and To achieve collaboration, the specific process is as follows: the stream processing layer cleans and aligns data in real time based on metadata, the batch processing layer trains models / optimizes rules based on historical data, and the results of both are written to the database, forming a closed loop of "real-time processing → batch optimization → real-time application".

[0072] ① Stream processing layer: It adopts the "Source→Transform→Sink" stream processing link of Apache Flink, combined with the metadata model to achieve real-time semantic alignment, process high-frequency AIS and vehicle GPS data in real time, realize real-time dynamic perception of ships / vehicles, generate real-time data such as the dynamic position and operating status of ships / vehicles, and store them in the real-time partition of the database to support the real-time scheduling and monitoring of port clusters.

[0073] Step 1: Parse the data into a Flink DataStream, extract the core fields (unique identifier MMSI / license plate number, spatiotemporal coordinates, status field, data source identifier), and generate the raw stream. .

[0074] Step 2: Call The standard terminology set and mapping rules in the document are for Perform real-time semantic unification while filtering obviously anomalous data (such as coordinates outside the geographical range or invalid unique identifiers). The initial anomaly screening formula can be found here. .

[0075] Step 3: Employ an adaptive sliding window (the window size dynamically adjusts according to the data update frequency) to... Calculate the real-time status indicators of ships / vehicles (average speed, cumulative mileage, current position cluster), with a sliding step size of 1 / 2 of the window size to balance real-time performance and smoothness.

[0076] Window parameter adaptive formula: Step 4: Integrate the Flink CEP (Complex Event Processing) engine, based on The system detects abnormal thresholds (such as ships exceeding 60 knots or vehicles deviating from preset routes) and monitors abnormal states in real time; it also writes "normal status data + abnormal alarms" into the system. Partition.

[0077] The anomaly detection threshold formula (dynamically adapted to ship / vehicle type) can be referenced from the following formula, which is based on the maximum value of the field in the metadata standard terminology set and is increased accordingly: For example: maximum standard speed for container ships Therefore, the abnormal threshold is 33 knots. If a speed greater than 33 knots is detected, an alarm will be triggered.

[0078] ② Batch processing layer: Using engines such as Spark, batch calculations and deep fusion are performed on low-frequency port operation data, spatial vector data and cleaned historical trajectories to generate a complete multimodal transport data chain, which is stored in the historical partition of the database to support long-term trend analysis and model training.

[0079] Step 1: Incremental data extraction is based on database partitioning and version control, extracting only newly added / updated data. The formula for determining this is as follows: in, t This is the current batch execution time. T =24h is the batch processing time window, covering all data from the previous day. This is the data version number (automatically incremented during stream processing writes). For the previous batch processing of entities e Maximum version number (to avoid duplicate extraction). For timestamps.

[0080] The extracted data includes: real-time data from the previous day's stream processing ( The data includes daily merged data, newly added port operation data, updated spatial vector data, and historical trajectory data that has not yet been merged.

[0081] Step 2: Deep multi-source data association modeling. Using the "unique identifier + business association field" defined in the metadata model as the key, a multi-source data association model is constructed, which can be expressed as follows: in, For datasets to be correlated (such as historical trajectory data and port operation data). This represents the set of associated keys (e.g., ship MMSI + container number, vehicle license plate number + transport mission number). The key matching confidence level is (perfect match = 1.0, partial match such as fuzzy container number = 0.6). A matching threshold is used to filter low-confidence association results and ensure data reliability.

[0082] Example of association: Linking ship AIS trajectory data by "container number" Port loading and unloading operation data Truck GPS data This interconnection forms a complete intermodal transport chain: "ship berthing → loading and unloading operations → truck transfer".

[0083] Step 3: Multimodal transport link construction. Based on the correlation results, the various stages of multimodal transport are connected in chronological order to generate multimodal transport link data. Chain = [ship berthing, unloading operations, truck consolidation, rail freight transfer, final destination delivery], and the time consumption of each stage is calculated. in s This refers to intermodal transport links (such as unloading operations and truck consolidation).

[0084] Related Port facility parameters (such as berth water depth and yard capacity) and static attributes of transportation vehicles (such as ship deadweight tonnage and vehicle rated load) are used to supplement the static attributes of intermodal transport links and improve data richness.

[0085] Step 4: Write the completed historical trajectory and full intermodal transport link data into the system. The system is divided into two partitions based on "intermodal transport type + time" to support long-term trend analysis and model training.

[0086] Update based on batch fusion results The stream processing rules in the code. For example, supplementing metadata mapping rules (such as adding new mapping rules when a new non-standard term "container truck" is discovered). .

[0087] Output: A standardized, high-quality fusion dataset of ship and vehicle trajectories covering a major port. It features spatiotemporal uniformity (CGCS2000 coordinate system + UTC time), semantic consistency (metadata model standardization), and controllable quality (accuracy ≥95% after deduplication and completion), and can be directly used as the core data foundation for the construction of "vehicle-ship aggregates" and the analysis of transportation channels.

[0088] Step Two: Construction and Calibration of the "Vehicle-Ship Aggregate" This step is the core bridge connecting the micro-level individual data of port clusters (ships, vehicles, train tracks) with the macro-level transportation modes (ship routes, multimodal transport links). Its core innovation lies in breaking through the limitations of traditional isolated analysis of single modes of transportation, and constructing a new type of analysis unit with "logistics chain collaboration" as the logical link, providing a precise group analysis carrier for subsequent extraction of main transportation channels and optimization of multimodal transport efficiency. In practice, the trajectory of a single ship or vehicle is easily affected by factors such as weather and temporary scheduling, exhibiting randomness. However, a group of vehicles serving the same trade route and transporting the same type of cargo will show stable common characteristics in their spatiotemporal behavior and business attributes. For example, the container intermodal transport group between Port I and Port II will form regular behaviors around fixed routes for berthing and loading and unloading of the same type of cargo. Within the port cluster, different vehicles will naturally form functionally differentiated transport groups based on the route's radiation range, the order of berthing at ports, and the correlation of collection and distribution businesses. The distribution and cooperation patterns of these groups essentially reflect the functional division of labor between ports (such as core ports and feeder ports) and the linkage of cargo flow. Therefore, accurately identifying such groups is a prerequisite for achieving a global analysis of transport channels.

[0089] 2.1 Conceptual Definition and Specific Structure of "Vehicle-Shipboard Aggregate" The innovative concept of "vehicle-ship aggregation" breaks through the traditional paradigm of analyzing only a single ship, truck, or train. It defines it as "a collection of ships, freight vehicles, and operating trains serving the same logistics chain based on spatiotemporal synergy and business relevance." Here, "business relevance" specifically refers to "cargo flow connection" and "capacity synergy." For example, the group providing collection and distribution services for the 10,000 TEU ocean-going container ship in Yangshan Port Area of ​​Port I includes feeder barges of 500-1,000 TEU, 40-foot container trucks, and 50-car railway container trains. Because they jointly support the "ocean-feeder-distribution" full-chain transportation of the same batch of containers, they have a clear business synergy logic and are therefore classified as the same "vehicle-ship aggregation."

[0090] The "vehicle and vessel aggregation" is a four-dimensional structured unit containing "core members, association rules, business attributes, and hierarchical identifiers," mathematically expressed as: M is the core member set, containing all instances of transportation vehicles, categorized by type, and expressed as a formula: ,in Ship crew (each crew member includes a unique identifier MMSI and a feature vector) ), For the i-th ship member, Container truck members (each member contains a unique identifier MMSI and a feature vector) ), For the j-th container truck member, Railway train members (each member includes a unique identifier MMSI and a feature vector) ), For the first Members of a railway freight train.

[0091] R For the set of association rules, define the constraints between members based on "spatiotemporal collaboration + business association", which can be expressed as follows: Among them, spatiotemporal coordination rules The time difference between vessel arrival and truck / train transfer must be ≤Δt (vehicle-truck Δt=4h, vessel-train Δt=6h), and the spatial distance between the transfer point and the port berth / hub must be ≤Δs (Δs=20km). Business Association Rules To allow members to share the same container number, transport task number, or service port cluster: in, , For container number, , To serve the port cluster, , For transportation tasks.

[0092] For port cluster identification, such as "Port I-Port II core cluster".

[0093] A This is a set of core business attributes, representing the overall functional characteristics of the aggregate. It is generated by aggregating member features and can be formally expressed as: .in, RouteScope The route coverage area is defined as follows: (e.g., intercontinental, core port-feed port, coastal-inland waterway); CargoType is the type of cargo transported; TransMode is the intermodal transport mode (e.g., sea-road, sea-rail, sea-road-rail); and Capacity is the total transport capacity (total deadweight tonnage of ships + total rated load capacity of trucks + total rated transport capacity of trains).

[0094] LThis is a hierarchical classification based on functional scale (Ocean Trunk Line Group / Feeder Line Group / River-Sea Direct Group). Specifically, the Ocean Trunk Line Group (Cluster_A) is centered around 10,000-15,000 TEU large container ships, supported by 40-foot heavy-duty road trucks (within a transport radius of 500km) and 50-60 car trains, primarily handling intercontinental container transport, characterized by "long distance, large capacity, and high efficiency." The Feeder Line Group (Cluster_B) is centered around 500-3000 TEU medium-sized container ships, supported by 20-30-foot road trucks (within a transport radius of 50-100km). 0km), 30-40 car trains are used for railway branch line freight trains, which are responsible for the transshipment of containers between core ports (such as Port I) and surrounding feeder ports (such as Port IV and Port IX). The core function is to "supplement the decentralized and connect the trunk line". River-sea direct group (Cluster_C): With 2000-5000DWT river-sea direct container ships as the core, and 15-20 foot small road trucks as supporting equipment, there is no railway coordination. It mainly undertakes direct transportation between a coastal port (such as Port I) and a port along the line (such as Port X). The cargo is mainly industrial raw materials and consumer goods.

[0095] 2.2 Characteristic identification of "vehicle-ship aggregate" First, determine the scope of the analysis data: From the results of multi-source heterogeneous data fusion, screen all transportation data that have called at a certain port cluster (including core ports such as Port I, Port II, Port III, and Port IV) in the past three months, including ocean-going container ships, feeder vessels, container trucks, and railway freight trains, to ensure that the data covers the entire multimodal transport scenario of "sea transport-road transport-rail transport".

[0096] Subsequently, feature vector construction was carried out, and feature dimensions that fit the actual business were designed for different types of transportation to ensure that the features can accurately reflect the group correlation: Vessel Feature Vector (F_"vessel"): includes "vessel type (e.g., 10,000 TEU ocean-going vessel, 500 TEU feeder vessel), deadweight tonnage (e.g., 50,000 DWT), call frequency (number of times it calls a certain port group in the past 3 months), fixed port of origin / port of destination (e.g., port I-II), average speed (e.g., 15.2 knots), and carrier (e.g., COSCO)". Among them, "call frequency" and "fixed port of origin / port of destination" are directly related to the stability of the shipping route and are the core features for determining the group affiliation. Truck feature vector (F_"truck"): includes "vehicle type (40-foot container truck, 20-foot light truck), rated load (e.g., 30 tons), transportation radius (e.g., short distance within 200km, long distance within 500km), core port served (e.g., serving only Yangshan Port Area, covering multiple port areas), transportation cycle (average time to complete a single port-to-factory transportation, e.g., 4 hours)". "Service port" and "transportation radius" determine its cooperative relationship with ships in the collection and distribution of goods. Train feature vector (F_"train") includes "train type (container-only train, mixed freight train), approved capacity (e.g., a 50-car train can transport 250 40-foot containers), fixed timetable, loading and unloading stations, and frequency (e.g., 5 trains per week)". The "timetable" and "loading and unloading stations" are key to matching ship arrival times and achieving intermodal transport connections.

[0097] 2.3 The specific quantitative process for identifying the "vehicle-ship aggregation" group After the feature vectors are constructed, the group identification is carried out through a quantitative process of "feature standardization → business weight allocation → weighted distance calculation → DBSCAN clustering → multi-level verification". The algorithm parameter settings are combined with the business scale and data characteristics of the port group to ensure that the clustering results have practical business significance.

[0098] Step 1: Feature Standardization. Differences in the units of measurement of different features (such as speed in knots, deadweight tonnage in DWT) will affect the clustering results. Z-score standardization is used to convert all features into standard values ​​with a mean of 0 and a variance of 1.

[0099] Step 2: Business Weight Allocation. Different features have varying degrees of influence on "group affiliation." Weights are assigned based on business logic (core features have higher weights; innovation: different from traditional equal-weight clustering), formula: Satisfy the weight normalization condition: For example, the total weight of ship characteristics = 0.3 + 0.3 + 0.2 + 0.1 + 0.1 = 1.

[0100] Step 3: Quantify the feature similarity between two entities using weighted Euclidean distance. The smaller the distance, the stronger the correlation. Formula: in, For entities, , These are the eigenvalues.

[0101] Step 4: Execute the DBSCAN algorithm based on the weighted distance results. Parameter settings are optimized in conjunction with the business scenario. For each entity, search its... If the number of samples in the neighborhood is greater than or equal to min_samples, it is marked as a "core point". Clusters are formed by the connectivity of the core points, and finally the initial population cluster set is obtained. , For the first An initial population cluster.

[0102] The neighborhood distance is defined as “the Euclidean distance after feature standardization is 0.5”. This value is based on the transportation data of a port group over the past year. After verification by 100 sets of samples, the Euclidean distance of 0.5 can ensure that the core features of transportation vehicles (such as routes, service ports, and cargo types) within the same group have an overlap of ≥85%, effectively distinguishing groups with different business attributes (such as ocean intermodal transport groups and feeder transshipment groups).

[0103] min_samples (minimum sample size): This is set based on the actual capacity requirements of multimodal transport services. It requires ≥10 ships (to cover the basic capacity of a feeder route), ≥100 trucks (to meet the daily transport needs of a core port area), and ≥10 trains (to support the stable frequency of trains on a railway trunk line). This is to avoid identifying "pseudo-groups" due to an insufficient sample size.

[0104] After the algorithm was run, 85 ocean-going vessels of the same type, 128 40-foot container trucks, and 20 50-car railway trains were grouped into the same cluster (C1). All vehicles in this cluster mainly transport containers and form a "ocean-going vessel-container truck-railway train" intermodal transport collaboration. The system assigned a unique identifier ID "CT_001" to it ("CT" represents ContainerTransport container transport, and "001" is the first trunk-level aggregate identified in a certain region).

[0105] Ultimately, based on the scale of transportation vehicles, the coverage of air routes, and the intermodal transport cooperation mode, the algorithm identifies three levels of "vehicle-ship aggregates," achieving functional differentiation and accurate classification of the groups.

[0106] Output: Multi-level, multi-scale container truck-ship aggregates. Each aggregate is associated with a unique ID, a complete list of members (ship MMSI, vehicle license plate, train number), and core business attributes (route range, cargo type, carrier), providing a clear group analysis object for subsequent accurate extraction of transportation channels and evaluation of multimodal transport efficiency.

[0107] Step 3: Digital Extraction and Representation of Major Transportation Corridors This step is the core link in transforming the collective trajectory analysis of "vehicle-ship aggregates" into the strategic knowledge needed for port cluster operation decisions. The core objective is to extract the main transportation channels in actual operation from massive individual trajectories, and through structured modeling and performance evaluation, form a "quantifiable, controllable, and optimizable" digital channel network, providing direct support for port cluster collaborative scheduling and supply chain resilience. In practice, ships, vehicles, and trains within the same "vehicle-ship aggregate" serve the same logistics chain, and their historical trajectories spatially converge on a few channel paths—these paths are the optimal routes selected by the group based on efficiency, cost, and safety in long-term operation, thus becoming the core basis for extracting main transportation channels.

[0108] 3.1 Channel Pattern Mining Based on Group Trajectory Based on the principle of "collective intelligence," although the trajectory of a single ship or vehicle may have random deviations due to factors such as weather and temporary scheduling, the collective selection of trajectories by the group can offset individual randomness, highlighting the optimal and most stable route. Using the full historical trajectory of a single "vehicle-ship aggregation" as input, appropriate algorithms are employed to extract routes for three transportation modes: sea, road, and rail. These routes are then fused through multi-mode fusion to form a complete route network. This algorithm can identify frequently used route corridors, i.e., the main transportation routes, from a large number of superimposed individual trajectories.

[0109] ① Route channel identification Taking "vehicle-ship aggregation CT_001" as an example, the system first obtains 125,000 AIS location points of its 86 internal vessels over the past three months and analyzes them using a trajectory density clustering algorithm. A key parameter in the algorithm is "neighborhood distance (...)". "Proximity threshold of trajectory points" is defined in geographic space. Combined with the actual width of a shipping route (the width of a typical coastal trunk route is approximately 0.5-1 nautical mile), it is used to... The distance is set to 0.5 nautical miles—meaning AIS points with a spatial distance of less than 0.5 nautical miles are considered to belong to the same "dense trajectory area"; at the same time, a density calculation formula is introduced: in For trajectory points The distance between, The total number of trajectory points. The bandwidth parameter is used to smooth individual trajectory fluctuations, and the average distance between adjacent AIS points is 0.12 nautical miles.

[0110] By using density thresholding and connectivity determination, discrete high-density points are integrated into continuous flight path segments. Polygon boundary fitting is performed on each connected region to form standardized flight path channels. The density threshold (Thresh) is calculated using an adaptive thresholding method, automatically dividing high-density and low-density regions. The formula is: ,in, , Threshold τ The proportions of the two types of points after segmentation , The average density of the two types of points is calculated. Thresh=45, and 38,960 AIS points with a density ≥45 are selected (high density points account for 32.8%).

[0111] Spatial distance < High-density points within a 0.25 nautical mile radius ( / 2 = 0.25 nautical miles) are grouped into the same connected region, according to the formula: ,in, For the trajectory point, For trajectory points distance, for The density.

[0112] Polygonal boundary fitting is performed on each connected region to form standardized flight path channels. Shape Algorithm ( (Sea mileage) Extract the contour boundary of the connected region, formula: .

[0113] The centerline of the calculated boundary is used as the baseline for the flight path, and k is the sampling point number on the boundary. .

[0114] ② Highway truck trajectory corridor recognition For the GPS trajectories of container trucks within the aggregate, the DBSCAN trajectory clustering algorithm adapted to highway scenarios is adopted, and the formula is defined as: Highway access = DBSCAN track({GPS track points}, ) in Set to 50 meters (matching the width of the highway lane and the vehicle's driving deviation range), clustering is used to aggregate spatially densely distributed GPS trajectory points into a continuous highway corridor, while removing discrete trajectory points caused by detours and temporary stops, ensuring that the extracted corridor is consistent with the actual main freight road.

[0115] For each valid cluster, standardized corridor parameters are extracted. The centerline extraction uses a weighted average method to calculate the lateral (perpendicular to the driving direction) midpoint of the GPS points within the cluster, forming the corridor centerline. The formula is: ,in .

[0116] The corridor width is calculated as twice the standard deviation of the lateral distance from the GPS point within the cluster to the centerline (covering 95% of the driving trajectory), using the following formula: x is the direction of travel coordinate, and y is the lateral coordinate.

[0117] ③ Railway transport corridor identification method Because railway transportation features "fixed routes and fixed stations," a graph construction method is used instead of trajectory clustering: Loading and unloading stations of railway trains within an aggregate are used as nodes, and the actual railway lines along which the trains operate are used as edges. Combining this with train frequency density (e.g., lines with an average of ≥3 trains per day are considered "main lines"), a railway transportation network graph is constructed. The line with the highest frequency density in the railway transportation network graph is the railway transportation corridor. That is: Railway corridor = graph construction (stations, lines, train frequency) A basic network is constructed using the loading and unloading stations of railway freight trains within the aggregate as nodes and the actual operating routes of the trains as edges. Node set ,in, The station's container throughput (unit: TEU). The frequency of intermodal connections with ports / trucks (average number of connections per day) and the two together reflect the intermodal hub value of the node.

[0118] If there are "stations" in the actual operation of the train →site For a direct route, define an edge. The basic properties of the edges are: in, The actual length of the line (unit: km). The route type is defined as either a dedicated freight line or a mixed passenger and freight line.

[0119] The weight of a node in the intermodal transport network is quantified; a higher weight indicates that the node is a hub station. .in, This is a weighting factor (throughput has a greater impact on hub value). Node weight, For nodes.

[0120] The edge weight calculation comprehensively considers three factors: frequency, transport volume, and operating efficiency, fully reflecting the importance of the route. The formula is: in, for → The average daily frequency of the trains (times / day). This represents the average daily container volume of this route. Average delay time (minutes / trip) for this train on this route. Weighting factor: =0.5 (shift stability) =0.35 (transportation volume) =0.15 (runtime validity), which satisfies... .

[0121] After extracting the single-mode channels, based on the intermodal transport collaboration of the "vehicle-ship aggregation", the single-mode channels of sea, road and rail are integrated through three steps: "connection node matching → channel weight superposition → integrity verification" to form a complete channel network.

[0122] Step 1: Connecting Node Matching. Identify connecting nodes for different transport modes (e.g., port berth → highway freight station, railway station → port yard), requiring compliance with spatiotemporal coordination constraints: in, For different modes of transport (such as sea transport) Highway passage ). The coordinates of the end point of the channel. This represents the average time it takes for the group to reach the destination of the passage.

[0123] Step 2: Calculate the overall weight of the integrated channel to reflect its importance in the multimodal transport network.

[0124] in, These represent the weights for single-mode channels (sea, road, and rail) respectively (single-mode channel weight = average weight of all edges / clusters within that channel). The weighting coefficients can be adjusted based on the combined transport mode, such as a large-scale ocean shipping network (Cluster_A). (Maritime transport is the dominant mode).

[0125] Output a complete channel network including basic information, structural features, and visual annotations, for example: For the ocean trunk line cluster (Cluster_A), the focus is on integrating its corresponding main shipping routes, highway corridors, and railway trunk lines to form a long-distance intermodal transport channel of "ocean-port-inland hub"; For the feeder transport cluster (Cluster_B), integrate the internal feeder waterways, highway feeder corridors, and railway feeder channels to form a regional connection channel of "core port - feeder port"; The focus of the River-Sea Direct Transport Cluster (Cluster_C) is on river-sea intermodal transport channels adapted to the Yangtze River waterway, integrating direct shipping routes between coastal ports and ports along the Yangtze River and supporting highway connections, ultimately forming a multi-level channel network covering different transportation needs.

[0126] 3.2 Structuring of transportation corridors To address the issues of vague concepts and lack of quantifiable control in traditional "transportation corridors," the corridors extracted earlier (such as corridor A and corridor B) are abstracted into a standardized digital network model—with "key nodes + edges" as the core structure. Key nodes are clearly defined as specific functional entities (such as dedicated container berths in ports, railway marshalling yards, and highway logistics parks), while edges correspond to actual transportation routes (such as shipping routes, highway trunk lines, and railway sections). At the same time, the model is endowed with multi-dimensional attributes, realizing the transformation of corridors from "spatial concepts" to "digital entities."

[0127] The model is expressed using a standardized graph structure, mathematically defined as follows: M ={ N,E,Attr},in N E represents the set of key nodes (concrete functional entities); E represents the set of edges (structured transportation routes, corresponding channel segments). Attr A collection of multi-dimensional attributes (including node attributes) Attr N Edge attributes Attr E ).

[0128] Step 1: Channel Deconstruction and Core Element Extraction. The abstract channel is first broken down into three core elements: "nodes, lines, and business characteristics," providing data support for structuring.

[0129] From the port cluster infrastructure database, the "functional entity-level nodes" involved in the channel can be selected using the following criteria: The channel is split into "continuous line segments with a single transportation mode". The splitting rules are based on the transportation mode (sea transport segment, road connecting segment, rail connecting segment) and the node spacing (adjacent nodes are independent line segments).

[0130] Collect historical data associated with the channel (such as CT_001 aggregate operation data, port operation records, and meteorological interference events) to provide a data source for attribute assignment.

[0131] Step 2: Based on the extracted node elements, construct the key nodes (N) in a structured manner.

[0132] Based on their functional roles, they are divided into "starting node (N_s), ending node (N_t), and transit node (N_m)".

[0133] The encoding rule of "node type - region - sequence number" is adopted and associated with the channel attribute.

[0134] Each node is assigned "job attribute + connection attribute".

[0135] Step 3: For the split line segments, construct the transportation line (E) in a structured manner.

[0136] Each line segment corresponds to an "edge (E)" in the model, specifying the edge type (sea edge E_sea, highway edge E_road, railway edge E_rail) and the origin and destination node codes.

[0137] The "Channel ID - Edge Type - Sequence Number" encoding is used, for example: CN-SH-NB-001-SEA-001 (Shanghai-Ningbo Channel Main Shipping Route Edge), to ensure the traceability and association between the edge and the channel.

[0138] Line attribute binding: Decompose the length and avg_transit_time in the channel attributes to the corresponding edges.

[0139] Step 4: Structured implementation of the multi-dimensional attribute system (Attr).

[0140] Taking "a certain main container shipping channel (corridor A)" as an example, the system transforms it into a digital channel and stores it in the port cluster data platform. The structured attribute design fully meets the needs of business management and control, and its structured system is as follows: channel_id: The encoding rule of "country code-origin port code-destination port code-serial number" is adopted to generate a unique identifier "CN-SH-NB-001" ("CN" is the code for China, "SH" represents the core port area of ​​Port I, and "NB" represents the core port area of ​​Port II) to ensure the uniqueness of channel identification across ports and systems; channel_level: Classified based on the dual dimensions of "freight volume + radiation range" - the annual container throughput of this channel exceeds 5 million TEUs, and it radiates a core economic zone, so it is classified as "backbone" level, which clarifies its core position in the port cluster channel network; start_port / end_port: Precisely points to the specific berth for operation - the starting point is "I Port Deepwater Port Area (Container Berths 1-3)" and the ending point is "II Port Area (Container Intermodal Berths 5-7)", ensuring direct connection with port operation scheduling; connected_ports: Marks key transit nodes along the corridor, including Dushan Port Area of ​​a certain port (a feeder barge transit port, undertaking feeder feeding tasks between Port I and Port A) and Jintang Port Area of ​​a certain port (a container feeding port, diverting the loading and unloading pressure of large ships), clearly reflecting the network radiation relationship of the corridor. Length: The unit is marked according to the mode of transport - the main sea route is about 120 nautical miles (using international standard sea transport units), and the land connecting section is about 35 kilometers, to avoid operational calculation deviations caused by a single unit; avg_transit_time: Based on the actual operation data statistics of the CT_001 aggregate over the past 3 months, the “full-link timeliness” calculation logic is adopted, which includes the sea voyage time (approximately 12 hours) and the waiting time for connection between the two ports (including berthing, loading and unloading, and truck transfer, approximately 4 hours). The final weekly average transit time is 16 hours, which truly reflects the full cycle duration of the cargo flow through the channel. estimated_weekly_volume: Based on the capacity of the CT_001 aggregate, the calculation is made by combining the average deadweight of the 86 vessels in the aggregate (50,000 DWT, approximately 2,500 TEU / vessel), the average number of voyages per week (2), and the matching degree of land transport capacity. The estimated average weekly cargo volume is approximately 12,000 TEU, clarifying the correlation between channel cargo volume and the aggregate. The resilience score is calculated using a 0-1 point scale based on historical data from the past year. It includes statistics on the number of temporary closures of the channel due to typhoons and heavy fog (an average of 2 times per year, with each impact lasting ≤6 hours), combined with the recovery efficiency of interference events such as channel dredging and equipment failures (average recovery time ≤2 hours). The final resilience score is 0.92 (≥0.9 is defined as "high reliability level"), which intuitively quantifies the channel's ability to withstand risks.

[0141] Through the above-mentioned structured processing, each transportation channel forms a digital entity with "clear attributes, logical connections, and traceable data", which can directly support the efficiency comparison, dynamic capacity allocation, and risk simulation between channels.

[0142] 3.3 Channel Performance Evaluation ① Transportation timeliness assessment The "door-to-door end-to-end timeliness" calculation is used, and the formula is as follows: in The transit time for ships within the channel (taken from AIS data statistics). This refers to the loading, unloading, and transshipment waiting time for goods at both ports (taken from port operation data). The road / rail transport time for goods from the port to the inland destination (taken from truck GPS and railway train data) directly reflects the efficiency of the corridor in handling cargo flow.

[0143] ② Channel resilience score The formula for focusing channel anti-interference capability is: The “Interruption Time” refers to the duration of the corridor’s downtime due to natural disasters, equipment failures, etc., the “Total Operating Time” refers to the annual normal operating time of the corridor, the “Affected Scope” refers to the percentage of freight volume affected during the interruption, and the “Total Corridor Volume” refers to the annual total freight volume of the corridor. The closer the score is to 1, the stronger the corridor’s resilience. ③ Capacity utilization rate assessment The formula for assessing the supply and demand matching degree of corridor capacity is: The “actual flow” refers to the channel’s actual annual freight volume (based on aggregated data statistics), and the “design capacity” refers to the channel’s theoretical maximum capacity (calculated by combining waterway navigation capacity, port loading and unloading capacity, and land transport capacity). A utilization rate between 80% and 90% is considered the “optimal supply and demand state”. A utilization rate below 80% indicates that the capacity is idle, while a utilization rate above 90% indicates that the capacity is tight and needs to be expanded.

[0144] 3.4 Implementation Priority and Parameter Optimization Strategy To ensure the efficiency and accuracy of channel extraction and evaluation, implementation rules and parameter optimization methods adapted to port cluster business scenarios are designed: ① Data processing priority rules Priority = f (Update frequency, business criticality, data quality) Prioritization is based on three dimensions: update frequency, business criticality, and data quality. High-frequency dynamic data (such as AIS and GPS, with an update frequency of 1-30 seconds) has a higher priority than low-frequency static data (such as spatial vector data, with an update frequency of 1-3 months) because it directly affects real-time channel monitoring. Data on "vehicle-ship aggregation" serving ocean trunk lines (with high business criticality and related to intercontinental freight) has a higher priority than feeder aggregation data. Data sets with a data quality score of ≥95% (such as authoritative maritime bureau AIS data) have a higher priority than data sets with a quality score of <90% (such as third-party GPS data), ensuring that core business data is processed first and improving analysis efficiency.

[0145] ② Adaptive adjustment of clustering parameters For trajectory clustering algorithms Key parameters such as minimum sample size (MIN_PTs) are optimized using a "grid search + silhouette coefficient verification" method. The optimal parameter combination is found through the following formula: the silhouette coefficient (C) is used to evaluate the clustering effect, with a value range of [-1, 1]. The closer the value is to 1, the better the clustering result (high similarity of trajectory points within the same cluster and low similarity between different clusters). This method can automatically adjust parameters according to the geographical characteristics of different port clusters (such as the difference in width between coastal and inland waterway channels) and transportation modes (such as the difference in characteristics between highway and railway lines), avoiding clustering bias caused by fixed parameters.

[0146] Output: A digitally defined optimal container transport channel network with multiple performance indicators, covering multiple levels of channels such as ocean trunk lines, regional feeder lines, and direct river-sea transport for a port cluster. Each channel comes with complete structured attributes and performance evaluation indicators, which can directly support core business operations such as port cluster collaborative scheduling (e.g., adjusting cross-port vessel schedules based on channel capacity utilization), supply chain optimization (e.g., selecting the optimal channel based on transport timeliness), and resilience enhancement (e.g., developing emergency support plans for low-resilience channels).

[0147] Example 3 The present invention also discloses a system for fusing multi-source heterogeneous data of port clusters and digitizing transportation channels. The system is used to implement the method described in Embodiment 1. The system includes: a fusion module, a classification module, and an extraction module. The fusion module is used to construct a systematic data fusion framework based on multi-source heterogeneous data of port clusters, and based on the systematic data fusion framework, to establish a unified spatiotemporal benchmark and data specifications to form fused data products. The classification module is used to construct a "vehicle-vehicle aggregate" based on the fused data product, and to identify and classify transportation groups whose operational characteristics meet preset requirements based on the "vehicle-vehicle aggregate". The extraction module is used to extract transportation channels that meet preset requirements based on the trajectory data of transportation groups, so as to realize the digital and structured representation of transportation channels.

[0148] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for fusing multi-source heterogeneous data from port clusters and digitizing transportation channels, characterized in that, The method includes: Based on multi-source heterogeneous data from port clusters, a systematic data fusion framework is constructed. Based on the systematic data fusion framework, a unified spatiotemporal benchmark and data specifications are established to form fused data products. Based on the fused data product, a "vehicle-ship aggregation" is constructed. Based on the "vehicle-ship aggregation", transportation groups whose operational characteristics meet preset requirements are identified and classified. Based on the trajectory data of transportation groups, transportation channels that meet preset requirements are extracted, enabling the digital and structured representation of transportation channels.

2. The method according to claim 1, characterized in that, The method for constructing the systematic data fusion framework includes: Establish the National Geographic Coordinate System 2000 and UTC time as a unified spatiotemporal reference; Based on a unified spatiotemporal benchmark, a metadata model for port cluster business is constructed, and the semantic definition of key business fields is unified. Based on the unified semantic definition of key business fields, a deduplication algorithm based on data source abundance is used to merge duplicate data from multiple sources. Based on the Lambda stream-batch integrated architecture, collaborative processing of multi-frequency data after merging duplicate points is realized, thus completing the construction of the systematic data fusion framework.

3. The method according to claim 2, characterized in that, The metadata model uses "entity-field-term-rule" as its core logic, defining entities such as ships, container trucks, railway trains, and port facilities. Each entity is associated with "mandatory core fields + optional extended fields," forming a standardized data dictionary. The expression is as follows: ;in, It is a collection of core entities including ships, container trucks, railway trains, and port facilities; For each entity, it is a collection of fields; A standard terminology set for entity fields; A set of mapping rules from standard terms to standard terms; U Iterative update rules for the basic information table For numeric fields, for Entity fields, for Entity mapping rules for Standard terminology for entities.

4. The method according to claim 1, characterized in that, The "vehicle and vessel aggregation" is a four-dimensional structured unit containing "core members, association rules, business attributes, and hierarchical identifiers," mathematically expressed as: ; Where M is the core member set, containing all instances of transportation vehicles, categorized by type, and formally expressed as: ,in For ship crew, For the i-th ship member, For container truck crew, For the j-th container truck member, As a member of the railway freight train, For the first Members of the railway freight train; R For the set of association rules, define the constraints between members based on "spatiotemporal collaboration + business association", which can be expressed as follows: Among them, spatiotemporal coordination rules Business association rules are defined as follows: The difference between the ship's arrival time and the truck / train transfer time is ≤Δt, and the spatial distance between the transfer point and the port berth / hub is ≤Δs. To allow members to share the same container number, transport task number, or service port cluster: , in, For port cluster identification, , For container number, , To serve the port cluster, , For transportation tasks; A represents the set of core business attributes, symbolizing the overall functional characteristics of the aggregate. It is generated by aggregating member characteristics and can be expressed as a formula: In this context, RouteScope represents the route's coverage area, CargoType represents the type of cargo transported, TransMode represents the intermodal transport mode, and Capacity represents the total transport capacity. L The hierarchy is categorized by functional scale, with the following levels: Ocean Trunk Line Cluster_A: Centered on 10,000-15,000 TEU large container ships, supported by 40-foot heavy-duty highway trucks and 50-60 car trainsets, undertaking intercontinental container transport; Feeder Line Cluster_B: Centered on 500-3,000 TEU medium-sized container ships, supported by 20-30-foot feeder trucks and 30-40 car trainsets, undertaking container transshipment between core ports and surrounding feeder ports; River-Sea Direct Cluster_C: Centered on 2,000-5,000 DWT river-sea direct container ships, supported by 15-20-foot small highway trucks, without railway support, undertaking direct port transport.

5. The method according to claim 1, characterized in that, Based on the aforementioned "vehicle-vehicle aggregation," methods for identifying and classifying transportation groups whose operational characteristics meet preset requirements include: Extract multidimensional features of transportation entities from fused data, including spatiotemporal features, business attribute features, and behavioral pattern features; The DBSCAN density clustering algorithm is used to perform unsupervised learning on the multidimensional features to achieve automatic identification and classification of transport groups.

6. The method according to claim 1, characterized in that, Methods for extracting transportation channels that meet preset requirements include: Using the historical trajectory of "vehicle-ship aggregation" as input, a trajectory density clustering algorithm is used to identify frequently used path corridors; The identified path corridors are abstracted into a digital network model, and each corridor is assigned basic attributes, dynamic performance attributes, and sustainability attributes.

7. A system for fusing multi-source heterogeneous data from port clusters and digitizing transportation channels, the system being used to implement the method described in any one of claims 1-6, characterized in that, The system includes: a fusion module, a classification module, and an extraction module; The fusion module is used to construct a systematic data fusion framework based on multi-source heterogeneous data of the port cluster, and based on the systematic data fusion framework, to establish a unified spatiotemporal benchmark and data specifications to form fused data products. The classification module is used to construct a "vehicle-vehicle aggregation" based on the fused data product, and to identify and classify transportation groups whose operational characteristics meet preset requirements based on the "vehicle-vehicle aggregation". The extraction module is used to extract transportation channels that meet preset requirements based on the trajectory data of transportation groups, thereby realizing the digital and structured representation of transportation channels.