Port data processing method
By receiving logistics order information and using mapping rules and knowledge graphs to analyze port data, the problem of low processing efficiency and data lag in port cargo source analysis has been solved. This has enabled accurate analysis and real-time data support for port data, improved the ability to identify cargo fluctuation patterns, and solved the problem of lagging route resources and warehousing configuration.
Patent Information
- Application Number
- CN202511427623.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-01-06
AI Technical Summary
Existing ports rely on manual data processing for cargo analysis, which is inefficient and unable to meet the needs of in-depth data mining. They also lack the ability to systematically analyze cargo fluctuation patterns, resulting in lagging adjustments to shipping routes and warehousing configurations, making it difficult to match dynamic cargo demand.
By receiving logistics order information, determining semantic and structural information, generating structured text using preset mapping rules, and combining enterprise knowledge graphs and logistics auxiliary data to analyze logistics data, analyzing static and dynamic anomalies in time-series data, generating logistics change trend data, and finally analyzing changes in logistics nodes and cargo volume, providing data support for ports to adjust shipping routes and warehousing.
It enables rapid and standardized processing of logistics data, accurately acquires logistics node and cargo volume information, captures dynamic changes in advance, provides real-time data support, solves the problem of lagging adjustments to route resources and warehousing configuration, and improves the efficiency of port cargo source expansion and the accuracy of business decisions.
Smart Images

Figure CN121278480A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic digital processing technology, and in particular to a method for processing port data. Background Technology
[0002] In port freight operations, cargo analysis and statistics are a crucial foundation for supporting route adjustments, warehousing allocation, and business decisions. Currently, port cargo volume is frequently fluctuating due to factors such as market demand, policy changes, and supply chain volatility. This is manifested in short-term surges and drops in cargo volume, dynamic changes in the distribution of logistics nodes, and continuous adjustments in cargo type structure.
[0003] However, existing ports rely heavily on manual processing of cargo manifest data and statistics on cargo sources. This not only results in low processing efficiency and difficulty in meeting the needs of in-depth data mining of massive amounts of data, but also lacks the ability to systematically analyze the patterns of cargo fluctuations. Consequently, they are unable to accurately identify the causes of abnormal cargo volumes, leading to delays in the adjustment of shipping route resources and warehousing configurations. This makes it difficult to match dynamic cargo source demands and ultimately affects the efficiency of terminal cargo source expansion and the accuracy of business decisions. Summary of the Invention
[0004] The main purpose of this application is to provide a method for processing port data, which aims to solve the technical problem that the adjustment of shipping route resources and warehousing configuration is lagging behind and difficult to match dynamic cargo demand.
[0005] To achieve the above objectives, this application provides a method for processing port data, the method comprising:
[0006] In response to the input logistics order information, determine the semantic information and structural information corresponding to the logistics order information;
[0007] Based on preset mapping rules, the semantic information and structural information are mapped into structured text containing cargo classification information;
[0008] Logistics data is parsed based on a knowledge graph pre-constructed from enterprise information, logistics auxiliary data associated with the structured text, and the structured text itself.
[0009] The static and dynamic anomaly results of the time-series data corresponding to the logistics data are analyzed to obtain logistics change trend data, which includes logistics node change data and cargo volume change data.
[0010] The logistics node analysis results are obtained by analyzing the logistics data and the logistics change trend data.
[0011] In one embodiment, the step of determining the semantic information and structural information corresponding to the input logistics order information in response to the input logistics order information includes:
[0012] Receive input logistics order information, which includes unstructured text information and structured field information;
[0013] Unstructured text information is split using a text segmentation algorithm, and the splitting results are transformed into semantic information that can characterize the semantic association attributes of the text based on a word embedding model.
[0014] The structured field information is format-validated according to the preset data specifications to remove invalid fields and correct abnormal data, thereby obtaining the standardized structured information.
[0015] The semantic features are associated with standardized structured data through the unique identifier of the logistics order, forming an associated data set.
[0016] In one embodiment, the step of mapping the semantic information and structural information into structured text containing cargo classification information based on preset mapping rules includes:
[0017] Invoke a preset set of mapping rules, which includes association rules between semantic features and goods classification, structured data, and supplementary classification rules;
[0018] The semantic information is input into the mapping rule set, and the corresponding cargo classification candidate results are initially determined by matching the semantic features with the association rules.
[0019] Extract the cargo volume and transportation mode data from the structured information, and verify the cargo classification candidate results in conjunction with the classification supplement rules, eliminating candidate classifications that do not match the structured data;
[0020] Based on the verified cargo classification results, the semantic information, structural information and cargo classification information are integrated according to a preset field format to generate structured text containing cargo classification information.
[0021] In one embodiment, the step of parsing logistics data based on a knowledge graph pre-constructed based on enterprise information, logistics auxiliary data associated with the structured text, and the structured text includes:
[0022] The system invokes a knowledge graph pre-constructed based on enterprise business registration information and historical cooperation information, which contains the relationship between enterprise entities and regional information.
[0023] Extract the shipper's company identifier and cargo classification information from the structured text, match the target company entity in the knowledge graph based on the company identifier, and obtain the regional association data corresponding to the target company.
[0024] Retrieve logistics auxiliary data associated with the structured text, the logistics auxiliary data including historical transportation routes and warehouse node records;
[0025] By integrating and analyzing regional correlation data, logistics auxiliary data, and cargo volume and customs declaration time information in structured text, and removing data conflicts, logistics data containing logistics node information, cargo volume statistics, and transportation route correlation information is extracted.
[0026] In one embodiment, the step of parsing the static and dynamic anomaly results of the time-series data corresponding to the logistics data to obtain logistics change trend data, wherein the logistics change trend data includes logistics node change data and cargo volume change data, includes:
[0027] Extract time-series data corresponding to each dimension of the logistics data. The time-series data is divided according to a preset time period, and each time-series data is associated with corresponding logistics node information and cargo volume information.
[0028] Static anomaly analysis is performed on the time series data. The static anomaly result is determined by calculating the deviation value between the data in each time period and the historical data of the same period and comparing it with a preset deviation threshold.
[0029] Dynamic anomaly analysis is performed on the time series data. The data change trend is fitted by a time series prediction model to identify the deviation between the actual data and the fitted trend and determine the dynamic anomaly results.
[0030] By integrating the static and dynamic anomaly results, abnormal data involving changes in logistics node information are filtered out to generate logistics node change data, and abnormal data involving changes in cargo volume information are filtered out to generate cargo volume change data. The logistics node change data and cargo volume change data are combined into logistics change trend data.
[0031] In one embodiment, the step of analyzing the logistics data and the logistics change trend data to obtain logistics node analysis results includes:
[0032] Extract basic information of logistics nodes, cargo volume data corresponding to each logistics node, and transportation route data associated with logistics nodes from logistics data;
[0033] Extract logistics node change data from the logistics change trend data to determine the addition, reduction and fluctuation of each logistics node;
[0034] By linking the basic information of logistics nodes with the corresponding cargo volume data, the cargo volume ratio and cargo volume growth rate of each logistics node can be calculated.
[0035] By combining the changes in logistics nodes with the proportion of cargo volume and the growth rate of cargo volume, the activity and stability of each logistics node are analyzed.
[0036] By integrating cargo volume, activity, stability, and related transportation route data of each logistics node, a logistics node analysis result is generated, which includes the classification of logistics nodes, the identification of key logistics nodes, and the assessment of the development potential of logistics nodes.
[0037] In one embodiment, after the step of parsing the static and dynamic anomaly results of the time-series data corresponding to the logistics data to obtain logistics change trend data, wherein the logistics change trend data includes logistics node change data and cargo volume change data, the process includes:
[0038] By combining the first historical logistics data, the logistics data, and the logistics change trend data, short-term results are predicted; and by combining the second historical logistics data, the logistics data, and the logistics change trend data, long-term results are predicted.
[0039] Based on the combined short-term and long-term results, the storage planning data for each cargo category and the corresponding flight route data for each cargo category are adjusted.
[0040] In one embodiment, the step of combining the first historical logistics data, the logistics data, and the logistics change trend data to predict short-term results includes:
[0041] Call the first historical logistics data within a preset time range. The first historical logistics data includes logistics node distribution data, cargo volume fluctuation data and short-term anomaly correction records within the corresponding period.
[0042] The real-time cargo volume information and logistics node association information in the logistics data are fused with the short-term abnormal change data in the logistics change trend data to form a predictive input dataset.
[0043] The predicted input dataset and the first historical logistics data are input into the short-term prediction model, and the model learns the short-term fluctuation patterns and abnormal impact weights in the historical data.
[0044] Based on the prediction results output by the model, short-term results are generated, including short-term cargo volume forecast data, logistics node traffic change forecast data, and short-term abnormal risk warnings.
[0045] In one embodiment, the step of combining the second historical logistics data, the logistics data, and the logistics change trend data to predict long-term results includes:
[0046] Call up the second historical logistics data with a time span greater than or equal to a preset number of years. The second historical logistics data includes logistics node change data, cargo structure evolution data and long-term abnormal impact records within the corresponding year.
[0047] Extract the cargo category proportion information and long-term cooperative logistics node information from the logistics data, and integrate them with the long-term trend change data from the logistics change trend data to construct a long-term prediction input dataset;
[0048] The long-term forecast input dataset and the second historical logistics data are input into the long-term forecast model, and the model captures the cumulative impact of multi-year evolution patterns and trend anomalies in the historical data.
[0049] Based on the prediction results output by the model, long-term results are generated, including monthly cargo volume growth trend data, logistics node expansion potential data, and cargo type structure adjustment direction.
[0050] In one embodiment, the step of integrating the short-term and long-term results to adjust the storage planning data for each cargo category and the route data for each cargo category includes:
[0051] Extract short-term cargo volume forecast data and short-term abnormal risk warnings from the short-term results, and integrate them with cargo volume growth trend data and cargo type structure adjustment direction data from the long-term results for the second future quantity time period to generate a comprehensive analysis dataset.
[0052] Based on the comprehensive analysis dataset, the peak storage demand for the first quantity period and the average storage demand for the second quantity period are calculated according to the category of goods. The storage planning data for each category of goods is then adjusted. The storage planning data includes the allocation ratio of storage areas and the number of reserved storage spaces.
[0053] Extract the estimated data on logistics node traffic changes and logistics node expansion potential data for each cargo category from the comprehensive analysis dataset. Combine this with the current route capacity configuration information to adjust the route data corresponding to each cargo category. The route data includes the route frequency and the range of logistics nodes covered by the route.
[0054] The adjusted cargo category storage planning data and route data are output, wherein the duration of the first quantity time period is shorter than the duration of the second quantity time period, and the first quantity time period corresponds to the prediction period of short-term results, while the second quantity time period corresponds to the prediction period of long-term results.
[0055] This application provides a method for processing port data. First, in response to input logistics order information, the method determines the semantic and structural information corresponding to the logistics order information. Based on preset mapping rules, the semantic and structural information are mapped into structured text containing cargo classification information. Logistics data is parsed using a knowledge graph pre-constructed based on enterprise information, logistics auxiliary data associated with the structured text, and the structured text itself. Static and dynamic anomaly results of the time-series data corresponding to the logistics data are analyzed to obtain logistics change trend data, which includes logistics node change data and cargo volume change data. Finally, the logistics data and the logistics change trend data are analyzed to obtain logistics node analysis results.
[0056] In this application, after receiving logistics order information, semantic and structural information is determined, and unstructured text and structured fields are quickly decomposed to lay a standardized data foundation for subsequent analysis, avoiding the inefficiency of manual data processing. Secondly, structured text containing cargo classification is generated through preset mapping rules to clarify cargo category attribution, providing a classification basis for warehousing configuration and route planning by category. Then, logistics data is analyzed by combining enterprise knowledge graph and logistics auxiliary data to accurately obtain core information such as logistics nodes and cargo volume, solving the problem of ambiguous logistics node identification in traditional methods. Subsequently, static and dynamic anomalies in time-series data are analyzed to obtain trends in logistics node and cargo volume changes, capturing dynamic changes such as sudden increases or decreases in cargo volume and transfers of logistics nodes in advance, avoiding information lag. Finally, logistics data and change trends are analyzed to obtain logistics node analysis results, outputting key information such as logistics node activity and cargo volume fluctuation patterns, providing data support for ports to adjust routes and optimize warehousing in real time, and ultimately matching dynamic cargo demand to solve the problem of adjustment lag. Attached Figure Description
[0057] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0058] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0059] Figure 1 This is a flowchart illustrating an embodiment of the port data processing method of this application.
[0060] Figure 2 This is a flowchart illustrating steps S61-S62 in Embodiment 7 of the port data processing method of this application;
[0061] Figure 3 This is a schematic diagram of the hardware structure involved in the port data processing equipment of this application.
[0062] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0063] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0064] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0065] Currently, most ports rely on manual processing of cargo manifest data and statistics on cargo sources. This not only results in low processing efficiency and difficulty in meeting the needs of in-depth data mining of massive amounts of data, but also lacks the ability to systematically analyze the patterns of cargo fluctuations. Consequently, it is impossible to accurately identify the causes of abnormal cargo volumes, leading to delays in the adjustment of shipping route resources and warehousing configurations. This makes it difficult to match dynamic cargo source demands and ultimately affects the efficiency of terminal cargo source expansion and the accuracy of business decisions.
[0066] In this application, after receiving logistics order information, semantic and structural information is determined, and unstructured text and structured fields are quickly decomposed to lay a standardized data foundation for subsequent analysis, avoiding the inefficiency of manual data processing. Secondly, structured text containing cargo classification is generated through preset mapping rules to clarify cargo category attribution, providing a classification basis for warehousing configuration and route planning by category. Then, logistics data is analyzed by combining enterprise knowledge graph and logistics auxiliary data to accurately obtain core information such as logistics nodes and cargo volume, solving the problem of ambiguous logistics node identification in traditional methods. Subsequently, static and dynamic anomalies in time-series data are analyzed to obtain the trend of changes in logistics nodes and cargo volume, and to capture dynamic changes such as sudden increases or decreases in cargo volume and transfers of logistics nodes in advance, avoiding information lag. Finally, logistics data and change trends are analyzed to obtain logistics node analysis results, outputting key information such as logistics node activity and cargo volume fluctuation patterns, providing data support for ports to adjust routes and optimize warehousing in real time, and ultimately matching dynamic cargo demand to solve the problem of adjustment lag.
[0067] It should be noted that the executing entity in this embodiment can be a port data processing system, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a port data processing device capable of performing the above functions. This embodiment does not specifically limit it in this way. The following uses a port data processing system as the executing entity as an example to describe this embodiment and the following embodiments.
[0068] Based on this, Embodiment 1 of this application proposes a method for processing port data. Please refer to... Figure 1 The method for processing the port data includes steps S10 to S50:
[0069] Step S10: In response to the input logistics order information, determine the semantic information and structural information corresponding to the logistics order information.
[0070] In this embodiment, logistics order information refers to the document data received by the port that contains cargo transportation requirements, covering information such as cargo name, shipper company name, cargo quantity, and customs declaration time. Semantic information is data extracted from the unstructured text of the logistics order information that can represent the meaning of the text, such as the category association implied by the cargo name and the enterprise attributes implied by the shipper name. Structural information is data obtained after standardizing the structured fields in the logistics order information, including information such as a uniform format for cargo quantity, unit time, etc.
[0071] As an optional implementation, after receiving logistics order information, the system first uses the Jieba word segmentation algorithm to split the unstructured text portion, removing stop words such as "special" and "new type". Then, it uses the Word2Vec word embedding model to convert the splitting result into vector form to obtain semantic information. At the same time, the structured fields are validated according to preset standards, and the abnormal cargo volume values are corrected to complete the incorrect time format to obtain structured information.
[0072] As another optional implementation, the logistics order information is processed by calling the NLP preprocessing interface. First, key text such as the name of the goods and the name of the shipper are extracted by named entity recognition technology and semantic feature vectors are generated as semantic information. Then, the data types of fields such as the quantity of goods and the customs declaration time are converted and the units are unified to form standardized structural information. Finally, the semantic information and structural information are associated and stored through the order number.
[0073] Step S20: Based on preset mapping rules, the semantic information and structural information are mapped into structured text containing cargo classification information.
[0074] In this embodiment, the preset mapping rules are a pre-constructed set of matching rules for associating semantic information structure information with cargo classification. This set includes structured data on the correspondence between semantic features and cargo categories, as well as supplementary verification conditions for classification. Cargo classification information refers to category identifiers for goods classified according to port business standards, such as electromechanical products, fresh produce, and building materials. Structured text is a unified format text formed by integrating semantic information structure information and cargo classification information, facilitating subsequent data parsing and storage.
[0075] As an optional implementation method, preset mapping rules are retrieved from the rule base. First, the semantic vector of goods in the semantic information is compared with the semantic features of goods category in the rules to initially match the goods classification. Then, the volume and transportation mode data in the structural information are extracted to verify whether the initial matching result meets the supplementary rules. For example, goods with large volume and marked "refrigerated" need to be adjusted to fresh food. Finally, the verified goods classification information and semantic structural information are integrated according to the field format to generate structured text.
[0076] As another optional implementation, a rule engine is used to load preset mapping rules. First, semantic information is input to trigger the semantic matching module, which obtains multiple candidate results for cargo classification. Then, structural information is input to trigger the data verification module, which filters out candidate results that match the cargo volume and transportation mode, and determines the final cargo classification. Finally, a structured text containing cargo classification information is generated and stored in the database according to the fixed field order of "cargo name-cargo category-cargo quantity-shipper".
[0077] Step S30: Based on the knowledge graph pre-constructed based on enterprise information, the logistics auxiliary data associated with the structured text, and the structured text, the logistics data is parsed out.
[0078] In this embodiment, the knowledge graph pre-constructed based on enterprise information is an interconnected network built by integrating historical cooperative data of enterprise business registration information. It includes the correspondence between enterprise entities and their registered locations and business scopes, as well as historical transportation records between enterprises and ports. Logistics auxiliary data is supplementary data associated with structured text, covering historical transportation routes, warehousing node records, route and schedule information, etc. Logistics data is core business data parsed from the structured text of the logistics auxiliary data in the knowledge graph, containing key information such as logistics nodes, cargo volume, and transportation routes.
[0079] As an optional implementation, the shipper's company name and cargo classification are first extracted from the structured text. Based on the company name, the corresponding company entity is matched in the knowledge graph to obtain the company's registered address as preliminary logistics node information. Then, logistics auxiliary data such as historical transportation routes and warehousing records associated with the structured text are retrieved to verify whether the preliminary logistics node information is consistent with the transportation origin. At the same time, cargo volume data in the structured text is extracted and integrated to form logistics data containing cargo volume and transportation routes of logistics nodes.
[0080] As another optional implementation, the knowledge graph query interface is called, the shipper identifier in the structured text is input, and information such as the company's regional contact information is obtained; then, the logistics data interface is used to obtain the logistics auxiliary data associated with the order, including the expected transportation route and warehouse location; then, the cargo volume and customs declaration time in the structured text are extracted, the company's regional cargo volume and route information are integrated, and after removing data conflicts, complete logistics data is generated.
[0081] Step S40: Analyze the static and dynamic anomaly results of the time-series data corresponding to the logistics data to obtain logistics change trend data, which includes logistics node change data and cargo volume change data.
[0082] In this embodiment, the logistics data corresponding to time-series data is sequential data formed by organizing logistics data according to time periods, such as daily or monthly aggregated logistics node cargo volume data. Static anomaly results are anomalies identified by comparing the deviation between time-series data and historical data for the same period, such as a monthly cargo volume exceeding a preset threshold compared to the same period last year. Dynamic anomaly results are deviations between actual data and the fitted trend identified by fitting the data trend using a time-series prediction model, such as a sudden deviation of cargo volume from the long-term growth trend. Logistics change trend data is data reflecting logistics changes obtained by integrating static and dynamic anomaly results. Logistics node change data records the addition and reduction of logistics nodes, and cargo volume change data records sudden increases and decreases in cargo volume.
[0083] As an optional implementation method, logistics data is organized into time-series data by natural month, and the deviation of monthly cargo volume from the average monthly cargo volume of the same period in the past three years is calculated. If the deviation exceeds 20%, it is marked as a static anomaly. At the same time, the Prophet model is used to fit the cargo volume time-series trend of the past two years. If the actual cargo volume of the current month deviates from the predicted value by more than 15%, it is marked as a dynamic anomaly. Finally, the changes in logistics nodes involved in the static and dynamic anomalies are statistically analyzed to form logistics node change data, and the cargo volume fluctuations are statistically analyzed to form cargo volume change data, which together form logistics change trend data.
[0084] As another optional implementation, time-series data is generated by weekly polymer flow data. The isolated forest algorithm is used to detect anomalies in the time-series data and identify static anomalies that differ significantly from the distribution of historical weekly data. Then, the time-series data is dynamically predicted using an LSTM model, and the deviation points between the actual data and the predicted curve are marked as dynamic anomalies. Subsequently, records of the addition or disappearance of logistics nodes are extracted from the anomalies as logistics node change data, and records of cargo volume changes exceeding 30% month-on-month are extracted as cargo volume change data. The data are then integrated to obtain logistics change trend data.
[0085] Step S50: Analyze the logistics data and the logistics change trend data to obtain logistics node analysis results.
[0086] In this embodiment, the logistics node analysis results are business decision support data formed by comprehensively analyzing logistics data and logistics change trend data, including logistics node level classification, key logistics node identification, and logistics node development potential assessment.
[0087] As an optional implementation method, firstly, the cargo volume proportion and transportation frequency of each logistics node are extracted from the logistics data, and the cargo volume growth rate and change frequency of each logistics node are extracted from the logistics change trend data. Then, the comprehensive score of each logistics node is calculated according to the weights of cargo volume proportion (30%), transportation frequency (20%), growth rate (30%), and change frequency (20%). Based on the score, the logistics nodes are divided into three levels: core, key, and potential. At the same time, the top 10 logistics nodes with the highest scores are identified as key logistics nodes, and the growth trend of potential-level logistics nodes is analyzed to form a logistics node analysis result that includes a key list and potential assessment of the level classification.
[0088] As another optional implementation method, a clustering algorithm is used to group logistics nodes in the logistics data. Combined with cargo volume change data in the logistics change trend data, the cargo volume stability of each group of logistics nodes is determined. Then, the matching relationship between logistics nodes and cargo types is mined through association analysis, and the proportion of popular cargo types of each logistics node is counted. Finally, the cargo type matching results of stability assessment are integrated to generate a list of logistics nodes marked with core focus potential levels, as well as development suggestions for each logistics node, which together constitute the logistics node analysis results.
[0089] For example, a port conducts cargo source analysis at the end of a quarter. The system first receives 5,000 logistics order information for the month. After processing, it obtains semantic and structural information containing cargo semantic vectors and standardized cargo volume time. Then, through mapping, goods such as "televisions" and "refrigerators" are classified as electromechanical products, generating structured text containing cargo classification. Subsequently, based on the enterprise knowledge graph, the shipper "Company A" is matched to be registered in Zhejiang Province. Combined with logistics auxiliary data such as historical transportation routes, the logistics node is determined to be in Zhejiang Province, and the cargo volume is extracted as 1,200 tons, parsing out the logistics data. The time-series data is sorted by week, and it is found that the cargo volume in the second week increased by 25% compared with the same period last year (static anomaly) and deviated from the monthly growth trend (dynamic anomaly). It is determined that the logistics node has not changed but the cargo volume has suddenly increased, forming logistics change trend data. Finally, the cargo volume of a certain province logistics node accounts for 18% with a growth rate of 22%, and the comprehensive score is ranked as a key logistics node. At the same time, its popular cargo category is identified as electromechanical products, generating logistics node analysis results.
[0090] This embodiment processes logistics order data through a systematic process, achieving accurate extraction and analysis of cargo information and solving the problems of low efficiency and poor accuracy in traditional manual analysis. Simultaneously, by generating logistics node analysis results, it provides data support for ports to subsequently adjust shipping routes for electromechanical goods and reserve storage space, effectively addressing the issue of lagging adjustments to shipping route resources and storage configurations, and helping ports better match dynamic cargo demand.
[0091] Based on any of the above embodiments, in Embodiment 2 of this application, step S10 includes:
[0092] Step S11: Receive the input logistics order information, which includes unstructured text information and structured field information.
[0093] In this embodiment, logistics order information refers to complete data received by the port business system to record cargo transportation needs, covering various types of information related to cargo transportation. Unstructured text information refers to text content without a fixed format in the logistics order information, such as cargo name, full name of the shipper company, and delivery address remarks. Structured field information refers to field content with a fixed format and data type in the logistics order information, such as cargo quantity, customs declaration time, and transportation mode code.
[0094] As an optional implementation, the port business system connects to the customs declaration platform and the enterprise's shipping system through an interface to receive logistics order information in real time. After receiving the information, the system first performs an integrity check on the data. After confirming that the data contains unstructured text information and structured field information, the data is stored in a temporary data buffer for further processing.
[0095] As an alternative implementation method, staff can import logistics order information files in batches through the system upload port. The system parses the file content, automatically separates unstructured text information from structured field information, and marks the file number and order association identifier respectively to avoid data confusion.
[0096] Step S12: The unstructured text information is split using a text segmentation algorithm, and the splitting results are transformed into semantic information that can characterize the semantic association attributes of the text based on a word embedding model.
[0097] In this embodiment, the text segmentation algorithm is an algorithmic tool used to split continuous unstructured text into independent semantic units, capable of identifying core words in the text. The word embedding model is a model that transforms text words into vector form, and the semantic relationships between words can be represented through vector distance. Semantic information is data that reflects the meaning of unstructured text and the relationships between words, existing in vector form.
[0098] As an optional implementation, the system calls the Jieba word segmentation algorithm to process unstructured text information, splitting "XX brand smart TV" into semantic units such as "XX brand", "smart", and "TV", removing meaningless function words such as "of" and "and". The split semantic units are then input into the Word2Vec word embedding model to generate vectors for each semantic unit. The vector dimension is uniformly 256 dimensions, and these vectors together constitute semantic information.
[0099] As another optional implementation, the HanLP word segmentation tool is used to segment unstructured text information, while identifying the part of speech of words and selecting keywords such as nouns and adjectives as the segmentation results. Then, the segmentation results are input into the BERT word embedding model to generate vectors containing contextual semantic associations, forming semantic information that can represent the semantic association attributes of the text, and stored in the semantic database.
[0100] Step S13: Perform format validation on the structured field information according to the preset data specifications to remove invalid fields and correct abnormal data, thereby obtaining the standardized structured information.
[0101] In this embodiment, the preset data specification is a structured field information processing standard pre-defined by the port based on business needs, including the format requirements, data range, and unit standards for each field. Invalid fields are fields in the structured field information that have no actual data or whose data is unidentifiable, such as null-valued cargo quantity fields or garbled time fields. Abnormal data is data in the structured field information that does not conform to the preset data specification, such as cargo quantity values exceeding a reasonable range or incorrectly formatted customs declaration times. Structured information is standardized data obtained after format validation, invalid field removal, and abnormal data correction of the structured field information.
[0102] As an optional implementation, the system loads a preset data specification, in which the quantity field requires the unit to be "tons" and the value range to be 0-10000, and the customs declaration time requires the format to be "YYYY-MM-DD". The structured field information is validated field by field, the "transportation mode code" field with empty values is marked as an invalid field and removed, "quantity 100000 tons" is corrected to "quantity 9999 tons" (which meets the maximum range), and the customs declaration time in the format of "2024 / 05 / 20" is corrected to "2024-05-20", finally obtaining standardized structured information.
[0103] As another optional implementation method, the default data specification sets the validation rules for each structured field, such as the quantity field must be positive and the customs declaration time must be before the current time. The system uses a data validation tool to perform batch validation of the structured field information, automatically identify and delete invalid fields, correct abnormal data according to the reasonable average data of historical orders of the same type, generate structured information with a unified format after correction, and record the correction log at the same time.
[0104] Step S14: Associate the semantic information with standardized structured data using the unique identifier of the logistics order to form an associated data set.
[0105] In this embodiment, the unique identifier for each logistics order is a code used to uniquely distinguish each order, such as an order number or UUID code, ensuring that each order's data can be located individually. Semantic information is generated data that can represent the semantic association attributes of unstructured text. Standardized structured data is structured field information that conforms to preset data specifications. The associated data set is a complete data set formed by binding semantic information and standardized structured data through a unique identifier, facilitating subsequent unified retrieval and analysis.
[0106] As an optional implementation, the system extracts the order number of each logistics order as a unique identifier, first adds the order number field to the semantic information, and then adds the order number field to the standardized structured data. Subsequently, through database association query operation, the semantic information and standardized structured data are matched and integrated according to the order number to form an association data set containing order number, semantic vector, standardized cargo quantity, standardized time, etc., which is stored in the association database table.
[0107] As another optional implementation, a data association tool is used, which uses the UUID code of the logistics order as the unique key to map and associate semantic information with standardized structured data. During the association process, data consistency is detected to ensure that the semantic information and structured data under the same UUID code correspond to the same order. After the association is completed, a JSON format association data set is generated. Each set contains the UUID code and the corresponding semantic and structured data content, which is convenient for subsequent system interface calls.
[0108] For example, a port processes a batch of logistics orders from Company B in its daily operations. The business system first connects to Company B's shipping system to receive this batch of logistics order information. Among them, "intelligent sweeping robot" and "Company B's logistics department" are unstructured text information, while "volume 500 units," "customs declaration time 20240610," and "transportation mode code 01" are structured field information. The system first uses the HanLP word segmentation tool to split "intelligent sweeping robot" into "intelligent" and "sweeping robot," and "Company B's logistics department" into "Company B" and "logistics department." Then, the segmentation results are input into the BERT word embedding model to generate semantic information in the form of semantic vectors. Next, the structured field information is validated according to the preset data specifications. "volume 500 units" meets the unit and range requirements and does not need to be corrected. "customs declaration time 20240610" is corrected to "2024-06-10," and "transportation mode code 01" is effectively retained, finally obtaining standardized structured information. The system then extracts the order number of each order, "DY20240610001", as a unique identifier. Semantic information is then bound to standardized structured data through the order number, forming a set of associated data containing order number, semantic vector, cargo quantity, customs declaration time, etc. This data can then be directly used for cargo classification and logistics node analysis.
[0109] This embodiment achieves effective integration of unstructured and structured data through a coherent logistics order information processing flow. The generated associated data set provides a standardized and highly available data foundation for subsequent analysis, solving the problems of messy format and semantic ambiguity in traditional data processing, and improving the efficiency and accuracy of port logistics data processing.
[0110] Based on any of the above embodiments, in Embodiment 3 of this application, step S20 includes:
[0111] Step S21: Invoke a preset mapping rule set, which includes association rules between semantic features and cargo classification, structured data, and classification supplementary rules.
[0112] In this embodiment, the preset mapping rule set is a set of rules pre-built and stored in the rule base by the port to achieve data and cargo classification matching, and can be dynamically updated according to business needs. The association rule between semantic features and cargo classification is a rule that establishes the correspondence between lexical vectors and semantic attributes in semantic information and standard cargo categories of the port. For example, semantic features such as "television set" and "refrigerator" correspond to "mechanical and electrical products". The structured data and classification supplementary rules are rules that further verify cargo classification based on structured field information, used to eliminate misjudgments in semantic matching. For example, "cargo quantity marked 'refrigerated' and transportation mode is 'cold chain'" should correspond to "fresh produce".
[0113] As an optional implementation, the system calls a preset set of mapping rules through the rule engine interface. When calling, the version validity of the rule set is first verified to ensure that the latest updated rules are used. After loading, the semantic feature association rules and the structured data supplement rules are stored in a temporary cache area for quick calling in subsequent steps. At the same time, the rule calling time and version information are recorded.
[0114] As another optional implementation, staff can trigger a mapping rule set invocation command through the system management terminal. The system extracts the corresponding rule set from the distributed rule base, first performs a completeness check on the rules, and generates a rule loading report after confirming that the rules contain semantic association rules and structured supplementary rules. Then, the rule set is parsed into machine-executable code logic to prepare for subsequent classification and matching.
[0115] Step S22: Input the semantic information into the mapping rule set, and preliminarily determine the corresponding cargo classification candidate results by matching the association rules with semantic features.
[0116] In this embodiment, semantic information refers to the data generated in previous steps that represents the semantic association attributes of unstructured text in vector form, including semantic features of core text such as product name and shipper information. Semantic feature matching is the process of comparing the vector data and keyword features in the semantic information with the semantic association rules in the mapping rule set, and determining the corresponding product category by calculating similarity or keyword matching degree. The product category candidate result is a number of product category options that may match the semantic information obtained after semantic feature matching, and is not the final determined result.
[0117] As an optional implementation, the system inputs the semantic vector of the goods in the semantic information into the mapping rule set, and calculates the similarity between the vector and the standard vector of each category in the semantic association rule using the cosine similarity algorithm; sets the similarity threshold to 0.8, and filters out the categories with similarity ≥ 0.8. For example, the semantic vector of "intelligent sweeping robot" has a similarity of 0.85 with the standard vector of "electromechanical" and 0.78 with "home appliance", and finally determines "electromechanical" as the candidate result for goods classification.
[0118] As another optional implementation method, the core keywords in the semantic information are extracted and matched with the category keyword library of semantic association rules in the mapping rule set; the number of matched keywords is counted, and the two categories with the most matched keywords are taken as candidate results. For example, "photovoltaic inverter" matches 3 keywords of "new energy" and 1 keyword of "electromechanical", and "new energy" and "electromechanical" are determined as candidate results for the category of goods.
[0119] Step S23: Extract the cargo volume value and transportation mode related data from the structure information, and verify the cargo classification candidate results in combination with the classification supplement rules, and eliminate candidate classifications that do not match the structured data.
[0120] In this embodiment, the structural information is standardized structured field information, including data such as cargo quantity values, transportation mode codes, and corresponding text descriptions. The cargo quantity value is the specific data representing the quantity of goods in the structural information, and may include units. The transportation mode related data records the transportation method of the goods in the structural information, such as "sea transport," "land transport," and "cold chain transport." The classification supplementary rule verification is the process of using structured data to verify whether candidate cargo categories conform to actual transportation and quantity characteristics, used to eliminate unreasonable candidate results.
[0121] As an optional implementation, the system extracts "volume 100 tons" and "transportation mode: cold chain" from the structural information, and combines them with the classification supplementary rule in the mapping rule set - "goods transported by cold chain and with a volume ≥ 50 tons, the candidate category must include 'fresh food' or 'frozen food'". If the previous candidate result was "mechanical and electrical", which does not match the supplementary rule, the candidate category is removed. If the candidate result includes "fresh food", the candidate category is retained.
[0122] As another optional implementation, extract "volume 500 units" and "transportation mode: ordinary land transport" from the structural information, and call the classification supplement rule - "goods with a single shipment volume ≥ 300 units and transportation mode of ordinary transport, the candidate categories are prioritized as 'mechanical and electrical' and 'home appliances'"; if the candidate results include "fresh produce", then because "fresh produce is usually measured in 'tons' and often requires cold chain", the candidate category of "fresh produce" is removed, and only "mechanical and electrical" and "home appliances" are retained.
[0123] Step S24: Based on the verified goods classification results, the semantic information, structural information and goods classification information are integrated according to a preset field format to generate structured text containing goods classification information.
[0124] In this embodiment, the verified cargo classification result is a unique cargo category determined through semantic matching and structured data verification, such as "Mechanical and Electrical Products" or "New Energy Products". The preset field format is a fixed field order and format predefined by the port to integrate multiple types of data, ensuring the uniformity of the structured text, such as "Order Number - Cargo Semantic Keywords - Cargo Quantity - Transportation Mode - Cargo Classification - Customs Declaration Time". The structured text is standardized text data formed by integrating semantic, structural, and classification information, and can be directly used for subsequent logistics data parsing and business system calls.
[0125] As an optional implementation, the system determines that the verified goods classification result is "Mechanical and Electrical Products". According to the preset field format "Order Number-Semantic Keyword-Quantity-Transportation Mode-Goods Classification-Customs Declaration Time", the system fills in the corresponding fields in sequence with "DY20240615001", "Intelligent Sweeping Robot", "500 Units", "Ordinary Land Transport", "Mechanical and Electrical Products", and "2024-06-15" to generate structured text. After generation, the field format is verified to ensure that there are no missing or incorrect formats, and then stored in the structured text database.
[0126] As another optional implementation, the preset field format includes "cargo semantic vector summary - cargo quantity unit - transportation mode code - cargo classification code - structured data remarks". The system integrates the semantic information vector summary (truncated to the first 100 dimensions), "unit", "02" (ordinary land transportation code), "003" (mechanical and electrical category code), and "no abnormal data" into the corresponding fields; the system generates structured text in XML format and adds data validation tags to facilitate the verification of data integrity during subsequent system parsing.
[0127] For example, a port processes a batch of "photovoltaic modules" logistics orders from Company C. The system first calls a preset mapping rule set. This set includes association rules between semantic features such as "photovoltaic modules" and "solar panels" and the "new energy category," as well as supplementary rules for "shipments ≥ 100 pieces and transported by sea prioritize matching 'new energy category." The semantic information of the order (including the semantic vector of "photovoltaic modules" and the keyword "Company C's New Energy Business Unit") is input into the rule set. Through cosine similarity calculation, the semantic vector of "photovoltaic modules" has a similarity of 0.92 with the standard vector of "new energy category," initially determining "new energy category" as a candidate category for goods. Subsequently, the structural information "shipments 200 pieces" and "transportation method: sea" are extracted. Combined with the supplementary classification rule—"photovoltaic-related goods shipped by sea and with a shipment quantity ≥ 100 pieces must be classified as 'new energy category'"—verification confirms that "new energy category" meets the rule, and no candidate category needs to be removed. Finally, following the preset field format "order number-semantic keyword-volume-transportation mode-cargo category-customs declaration time", the system integrates "DY20240620008", "photovoltaic modules", "200 pieces", "sea freight", "new energy category", and "2024-06-20" to generate a structured text containing cargo classification information. This text will then be directly used in the logistics node analysis and route configuration stages.
[0128] This embodiment achieves accurate determination of cargo classification and data standardization through a coherent process of "rule invocation - semantic matching - structured verification - text integration". It solves the problems of low efficiency and easy misjudgment due to semantic ambiguity in traditional manual classification, and provides unified and accurate classification data support for subsequent port business processes.
[0129] Based on any of the above embodiments, in Embodiment 4 of this application, step S30 includes:
[0130] Step S31: Invoke the knowledge graph that was pre-constructed based on enterprise business information and historical cooperation information. The knowledge graph contains the relationship between enterprise entities and regional information.
[0131] In this embodiment, enterprise registration information is basic enterprise data obtained from the business registration system, including the enterprise's registered address, business scope, legal representative, etc. Historical cooperation information is logistics data generated from past cooperation between the port and the enterprise, such as historical shipping records and preferred shipping routes. The knowledge graph is a structured network of relationships built by integrating enterprise registration information and historical cooperation information, with the enterprise as the core entity, associating it with its corresponding region, business scope, and other attributes. The relationship between enterprise entities and regional information is the mapping relationship between enterprise entities and regional data such as registered address and frequently shipped areas in the knowledge graph; for example, "Company A" is associated with "Hangzhou City, Zhejiang Province" and "Shanghai City".
[0132] As an optional implementation, the system calls the knowledge graph through the graph query interface. Before the call, the update time of the graph data is verified to ensure that the version updated within the last 30 days is used. After loading, the enterprise entities and regional association data in the knowledge graph are loaded into the in-memory database, and the data indexing service is started at the same time to facilitate the rapid matching of enterprise identifiers in the future. The graph version and loading time are recorded during the call process.
[0133] Step S32: Extract the shipper's enterprise identifier and cargo classification information from the structured text, and obtain the regional association data corresponding to the target enterprise by matching the enterprise identifier with the target enterprise entity in the knowledge graph.
[0134] In this embodiment, the shipper's enterprise identifier is information in the structured text used to uniquely identify the shipper, such as the company's full name and unified social credit code. The goods classification information is the identified goods category identifier in the structured text, such as "new energy" or "electromechanical products." The target enterprise entity is the enterprise node in the knowledge graph that successfully matches the shipper's enterprise identifier, containing the enterprise's complete attribute data. The regional association data is the regional information associated with the target enterprise entity in the knowledge graph, including its registered location, historical shipping areas, and the location of its partner warehouses.
[0135] As an optional implementation, the system uses regular expressions to extract the "full name of the shipper company" from structured text as the company identifier, and extracts the "goods category" field as the goods category information; the full name of the company is input into the entity matching module of the knowledge graph, and the corresponding target company entity is found through the string fuzzy matching algorithm (similarity threshold set to 0.9); the "registered location" and "frequently shipped areas in the past year" are filtered from the attributes of the target company entity as the regional association data corresponding to the target company.
[0136] Step S33: Retrieve the logistics auxiliary data associated with the structured text, the logistics auxiliary data including historical transportation routes and warehousing node records.
[0137] In this embodiment, the logistics auxiliary data is supplementary logistics data corresponding to the current structured text, used to verify and improve logistics information, and is not part of the structured text itself. Historical transportation routes are the route data of the shipping company or similar goods during past transportation, including node information such as port of origin, transshipment port, and port of destination. Warehouse node records are data on warehousing facilities previously used by the shipping company or similar goods, including information such as warehouse location, storage capacity, and duration of cooperation.
[0138] As an optional implementation, the system extracts the order number from the structured text, uses it as the association key, and calls the query interface of the logistics auxiliary database. The query conditions are set as "enterprise identifier associated with order number + goods category", and the system filters out the transportation route data of the same type of goods of the enterprise in the past 3 months (i.e., historical transportation routes) and the records of the warehousing facilities cooperated by the enterprise (i.e., warehousing node records). The two types of data are integrated into logistics auxiliary data.
[0139] As another optional implementation, the system uses data association services to find a set of historical orders that match the structured text "shipper's company ID + goods category", extracts "transportation route details" from the historical orders as historical transportation paths, and extracts "warehouse name and address" as warehouse node records; the extracted data is deduplicated (data within the last 6 months is retained) to form non-duplicate logistics auxiliary data, and the historical order number of the data source is recorded at the same time.
[0140] Step S34 involves integrating and analyzing regional correlation data, logistics auxiliary data, and cargo volume and customs declaration time information from structured text. After eliminating data conflicts, logistics data containing logistics node information, cargo volume statistics, and transportation route correlation information is parsed out.
[0141] In this embodiment, cargo quantity information is the quantity data of goods recorded in structured text, including numerical values and units (such as "200 pieces" or "100 tons"). Customs declaration time information is the customs declaration time of goods recorded in structured text, in the format "YYYY-MM-DD". Fusion analysis is the process of integrating and logically verifying regional correlation data, logistics auxiliary data, cargo quantity information, and customs declaration time information to ensure consistency among various data types. Data conflicts are contradictory information from different data sources, such as regional correlation data showing "registered location is Shanghai" but historical transportation routes showing "port of origin is Guangzhou". Logistics data is the core business data obtained after fusion analysis, containing key information such as logistics nodes, cargo quantity, and transportation routes.
[0142] As an optional implementation, the system first compares the "frequently shipped areas" in the regional association data with the "historical transportation route origin port areas" in the logistics auxiliary data. If they match (e.g., both are "Ningbo City, Zhejiang Province"), the area is retained as a candidate logistics node. Then, the system integrates the "cargo volume 200 pieces" and "customs declaration time 2024-06-20" in the structured text to verify whether the cargo volume is within a reasonable range of the cargo volume of the historical transportation route (e.g., the cargo volume of the same historical route is 100-300 pieces, and 200 pieces meet the requirements). If the regional association data conflicts with the historical transportation route (e.g., the region shows "Shanghai" and the route origin port is "Guangzhou"), the "frequently shipped areas" in the regional association data are removed, and the origin port area of the historical transportation route is used as the standard. Finally, the system parses out the logistics data containing "logistics node: Ningbo City, Zhejiang Province", "cargo volume statistics: 200 pieces", and "transportation route association: Ningbo Port - Shanghai Port".
[0143] As another optional implementation method, a data fusion algorithm is used to associate the four types of data according to the dimensions of "region-time-volume-route": First, based on the customs declaration time, historical transportation routes and warehousing node records of the same period are filtered out; then, the "registered location" in the regional association data is compared with the "originating area" of the historical transportation routes. If there is a conflict, the warehousing node records are referenced (e.g., if the warehouse is in "Suzhou City, Jiangsu Province", then "Suzhou and surrounding areas" are selected as candidate logistics nodes); finally, the cargo volume information of the structured text is integrated to ensure that the cargo volume matches the route carrying capacity (e.g., the maximum single carrying capacity of the route "Suzhou Port-Qingdao Port" is 500 pieces, and the cargo volume of 300 pieces meets the requirements); after removing the conflicting "registered location: Hefei City, Anhui Province" data, logistics data containing logistics nodes, cargo volume statistics, and transportation route association information is parsed out.
[0144] For example, when a port processes a logistics order for "lithium batteries" from Company D, it first calls a knowledge graph pre-constructed based on Company D's business registration information (registered location: Shenzhen, Guangdong Province) and historical cooperation information (frequent shipments from Shenzhen Port in the past year). From the structured text of the order, it extracts "Shipper's Enterprise Identifier: Ding New Energy Technology Co., Ltd." and "Cargo Category: New Energy." By matching the company's full name with the entity of Company D in the knowledge graph, it obtains the regional association data "Registered Location: Shenzhen, Frequent Shipment Area: Surrounding Shenzhen Port." Next, it retrieves the logistics auxiliary data associated with the structured text, obtaining the historical transportation route "Shenzhen Port - Xiamen Port - Tianjin Port" and the warehousing node record "Shenzhen Yantian Port Warehousing Center." Subsequently, it integrates and analyzes the regional association data, logistics auxiliary data, and the structured text's "Cargo Quantity: 300 boxes" and "Customs Declaration Time: 2024-07-05": the regional association data "Shenzhen Port Surrounding Area" matches the historical transportation route "Port of Origin: Shenzhen Port," with no conflict; the cargo quantity of 300 boxes is within the reasonable range of 100-500 boxes per shipment in the historical route; and the customs declaration time matches the historical shipment time for the same period. The final parsed logistics data was "Logistics Node: Shenzhen Port, Shenzhen City, Guangdong Province; Cargo Volume: 300 TEUs; Transportation Route: Shenzhen Port - Xiamen Port - Tianjin Port." This data was used for subsequent anomaly detection and logistics node analysis. This embodiment, through the fusion of knowledge graphs and multi-source data, solves the problem of errors caused by traditional logistics node identification relying on only single data points, achieving accurate parsing of logistics data and providing a reliable data foundation for subsequent port business decisions.
[0145] Based on any of the above embodiments, in Embodiment 5 of this application, step S40 includes:
[0146] Step S41: Extract the time-series data corresponding to each dimension of the logistics data. The time-series data is divided according to a preset time period, and each time-series data is associated with the corresponding logistics node information and cargo volume information.
[0147] In this embodiment, the various dimensions of logistics data are the core dimensions that can be used for time-series analysis, including logistics node dimensions, cargo volume dimensions, and transportation route dimensions. Time-series data is a sequence of data formed by arranging the data of each dimension in chronological order, reflecting the changes in data over time. The preset time period is a pre-defined time unit for dividing the time-series data, such as day, week, or month. Associating the corresponding logistics node information and cargo volume information means that the time-series data of each time period is bound to the specific logistics node name, cargo volume value, and unit within that period.
[0148] As an optional implementation, the system extracts two core dimensions of data from logistics data: logistics nodes and cargo volume. It divides the data into preset "weekly" time periods and summarizes the cargo volume data of the same logistics node within each week to form three-dimensional time-series data of "time (week) - logistics node - cargo volume". For example, in the 25th week of 2024, the cargo volume corresponding to "Shenzhen Port, Shenzhen City, Guangdong Province" is 800 boxes, and in the 26th week of 2024, the cargo volume of the same logistics node is 950 boxes. The system generates continuous periodic time-series data in sequence, and adds a logistics node code and cargo volume unit label to each period of data.
[0149] Step S42: Perform static anomaly analysis on the time series data. Calculate the deviation value between the data in each time period and the historical data of the same period, and compare it with a preset deviation threshold to determine the static anomaly result.
[0150] In this embodiment, static anomaly analysis is an analytical method based on historical data from the same period to determine whether the current period's data is abnormal, without relying on the data's temporal trend. Historical data from the same period refers to data within the same time frame as the current period; for example, the historical data for week 25 of 2024 includes data from week 25 of 2023 and week 25 of 2022. The deviation value is the percentage difference between the current period's data and the average of historical data from the same period, relative to the average of historical data from the same period. The preset deviation threshold is a pre-set critical value for determining whether the data is abnormal, such as ±20%. The static anomaly result is a description of the time-series data exhibiting static anomalies, as determined through analysis, along with the degree of anomaly.
[0151] As an optional implementation, the system calculates the deviation of the time-series data for each period from the average of the same period data over the past 3 years. For example, in the 25th week of 2024, the cargo volume of "Shenzhen Port" was 950 TEUs, while the average for the same period over the past 3 years was 750 TEUs. The deviation is approximately (950-750) / 750≈26.7%. The deviation is compared with the preset ±20% deviation threshold. Since 26.7% exceeds the upper limit of the threshold, the data for this period is marked as "static anomaly (surge in cargo volume)," forming a static anomaly result. At the same time, the specific values of the historical average and deviation are recorded.
[0152] Step S43: Perform dynamic anomaly analysis on the time series data, fit the data change trend through the time series prediction model, identify the deviation between the actual data and the fitted trend, and determine the dynamic anomaly results.
[0153] In this embodiment, dynamic anomaly analysis is an analytical method based on the time-varying trend of data to determine whether the current period's data is abnormal, focusing on the continuous changes in the data. Time series prediction models are models used to fit the changing patterns of time series data and predict future data, such as LSTM models and Prophet models. Fitting the data changing trend involves the model learning the changing patterns of multiple historical consecutive periods of data to generate a smooth trend curve. The deviation between the actual data and the fitted trend is the difference and percentage deviation between the actual data of the current period and the model's predicted data. Dynamic anomaly results are time series data with identified dynamic anomalies and descriptions of the anomaly types.
[0154] As an optional implementation, the system uses the Prophet time series forecasting model. It takes continuous time series data from the past 52 weeks (divided by week) as input, and the model learns the weekly fluctuations and trend changes of the data to fit a predicted trend curve for the next 12 weeks. For example, the model predicts that the cargo volume of "Shenzhen Port" in the 26th week of 2024 will be 920 TEUs, while the actual data is 680 TEUs. The deviation between the actual and the prediction is approximately (680-920) / 920≈-26.1%. The deviation threshold is set at ±15%. Since -26.1% exceeds the lower limit of the threshold, the data for this period is marked as "dynamic anomaly (sudden drop in cargo volume)", forming a dynamic anomaly result. At the same time, the predicted data and the deviation value are stored.
[0155] Step S44: Integrate the static and dynamic anomaly results, filter out the anomaly data involving changes in logistics node information to generate logistics node change data, filter out the anomaly data involving changes in cargo volume information to generate cargo volume change data, and combine the logistics node change data and cargo volume change data into logistics change trend data.
[0156] In this embodiment, integrating anomaly results is the process of merging static and dynamic anomaly results by time period and logistics node, removing duplicate markers. Anomaly data involving changes in logistics node information includes data showing the addition, disappearance, or name change of logistics nodes, such as the sudden appearance of cargo volume data for "Gaolan Port, Zhuhai City, Guangdong Province" (a newly added logistics node) in a certain period. Anomaly data involving changes in cargo volume information includes data showing surges, drops, or continuous fluctuations in cargo volume. Logistics node change data is a dataset formed by summarizing all logistics node change anomaly data, and cargo volume change data is a dataset formed by summarizing all cargo volume change anomaly data. Logistics change trend data is a dataset reflecting the overall trend of logistics changes, formed by integrating the two types of change data in chronological order.
[0157] As an optional implementation, the system integrates static and dynamic anomaly results according to the "time period - logistics node" dimension, merging anomaly markers for the same logistics node within the same period into a single record. From the merged anomaly results, data categorized as "new logistics node (e.g., 'Gaolan Port' added in week 26 of 2024)" and "logistics node deactivated (e.g., no data for 'Guangzhou Port' after week 25 of 2024)" are selected and arranged chronologically to generate logistics node change data. Data categorized as "surge in cargo volume (deviation +26.7%)" and "sudden drop in cargo volume (deviation -26.1%)" are also selected, supplementing the cargo volume change magnitude and duration information to generate cargo volume change data. Finally, the two types of data are linked through the time period field and combined into logistics change trend data containing "time - logistics node change - cargo volume change," which is stored in the trend analysis database.
[0158] For example, a port conducted anomaly analysis on the logistics data of new energy goods from weeks 24 to 26 of 2024. First, logistics node and cargo volume data were extracted from the logistics data, and the time-series data was divided by week, resulting in the following time-series data: "Week 24: Shenzhen Port 780 TEUs, Week 25: Shenzhen Port 950 TEUs, Week 26: Shenzhen Port 680 TEUs + Gaolan Port 320 TEUs". During static anomaly analysis, the deviation of the cargo volume at Shenzhen Port in week 25 from the average of 750 TEUs over the same period over the past three years was calculated to be 26.7%, exceeding the ±20% threshold, and was marked as a static anomaly. For week 26, Gaolan Port had no historical data for the same period (new data), and was not initially considered a static anomaly. During dynamic anomaly analysis, the Prophet model was used to fit the data from the past 52 weeks, predicting a cargo volume of 920 TEUs at Shenzhen Port in week 26. The actual volume was 680 TEUs, a deviation of -26.1%, and was marked as a dynamic anomaly. For Gaolan Port, the newly added data had no comparison with the model's predicted value, and was not initially considered a dynamic anomaly. After integrating the abnormal results, data on logistics node changes was generated by filtering out "the addition of Gaolan Port in week 26" and cargo volume changes by filtering out "a surge in cargo volume at Shenzhen Port in week 25 and a sudden drop in cargo volume at Shenzhen Port in week 26". These were combined to form logistics change trend data. This data will be used to analyze the reasons for the shift in logistics nodes and cargo volume fluctuations of new energy goods, and to assist the port in adjusting shipping routes.
[0159] This embodiment comprehensively captures anomalies in logistics data changes by combining static and dynamic anomaly analysis. The generated logistics change trend data provides accurate information for ports to grasp cargo dynamics, solving the problem that traditional single anomaly detection is prone to missing key changes.
[0160] Based on any of the above embodiments, in Embodiment Six of this application, step S50 includes:
[0161] Step S51: Extract basic information of logistics nodes, cargo volume data corresponding to each logistics node, and transportation route data associated with logistics nodes from the logistics data.
[0162] In this embodiment, the basic information of a logistics node refers to the information in the logistics data that characterizes the basic attributes of the logistics node, including the logistics node name, administrative division affiliation, and port terminal, such as "Shenzhen Port, Shenzhen City, Guangdong Province" associated with "South China Region" and "Shenzhen Yantian International Container Terminal". The cargo volume data corresponding to each logistics node is the cargo quantity data summarized by the logistics node dimension in the logistics data, including cargo volume values and units for different periods (day / week / month), such as "Shenzhen Port's cargo volume in July 2024 was 5,000 TEUs". The transportation route data associated with the logistics node is the transportation route information bound to the logistics node in the logistics data, including the port of origin, transshipment port, port of destination, and corresponding route code, such as "Shenzhen Port - Xiamen Port - Tianjin Port (Route Code CY001)".
[0163] As an optional implementation, the system extracts data from the logistics database using SQL queries. First, it filters the "Logistics Node Name," "Administrative Division," and "Wharf" fields in the "Logistics Node Information" table as basic information for the logistics nodes. Then, it links the "Cargo Volume Statistics" table, grouping and summarizing by "Logistics Node Name + Statistical Period (Month)" to obtain monthly cargo volume data for each logistics node. Finally, it links the "Transportation Route" table, extracting the "Port of Origin - Port of Transit - Port of Destination" combination and route code corresponding to each logistics node to form logistics node-related transportation route data. The three types of data are stored together through the "Logistics Node Name" field.
[0164] Step S52: Extract logistics node change data from the logistics change trend data to determine the addition, reduction, and fluctuation of each logistics node.
[0165] In this embodiment, the logistics change trend data is a comprehensive dataset generated in previous steps that includes changes in logistics nodes and cargo volume. Logistics node change data is a subset of the logistics change trend data that specifically records changes in logistics nodes, including change time, change type, and information about the logistics nodes before and after the change. New logistics nodes refer to logistics nodes that appear for the first time in the current statistical period and have no historical records, such as "Zhuhai Gaolan Port, which first appeared in July 2024." Decreased logistics nodes refer to logistics nodes for which there is no data in the current statistical period and which have not appeared for two consecutive periods, such as "Guangzhou Nansha Port, for which there is no data in June-July 2024." Fluctuating logistics nodes refer to logistics nodes that, although neither newly added nor decreased, have cargo volume changes exceeding a preset threshold (e.g., ±30%) within a continuous period.
[0166] As an optional implementation, the system filters records with "Logistics Node Change" as the "Change Type" from the logistics change trend data to obtain logistics node change data; it groups the data by "Statistical Period (Month)" and compares the current period with the logistics node list of the previous three periods, marking newly added logistics nodes in the current period (such as Gaolan Port); marking logistics nodes with no data for two consecutive periods (such as Nansha Port) as a decrease; at the same time, it calculates the continuous periodic cargo volume change range of the remaining logistics nodes, and marks logistics nodes with a change range exceeding ±30% (such as Shenzhen Port with 4,000 TEUs in June and 2,800 TEUs in July, a decrease of 30%) as fluctuations, and finally forms a list of new, decreased and fluctuating logistics nodes.
[0167] Step S53: Associate the basic information of logistics nodes with the corresponding cargo volume data, and calculate the cargo volume ratio and cargo volume growth rate of each logistics node.
[0168] In this embodiment, the basic information of logistics nodes and the corresponding cargo volume data are linked by binding the two types of data through common fields (such as the logistics node name), forming a complete dataset of "basic attributes + cargo volume data". The cargo volume percentage is the percentage of a logistics node's cargo volume data relative to the total cargo volume of all logistics nodes, reflecting the cargo volume contribution of that logistics node. The calculation formula is "(cargo volume of a logistics node / total cargo volume of all logistics nodes) × 100%". The cargo volume growth rate is the growth rate of a logistics node's current period cargo volume relative to the previous week's futures volume, reflecting the trend of cargo volume changes. The calculation formula is "(current period cargo volume - previous week's futures volume) / previous week's futures volume × 100%".
[0169] As an optional implementation, the system uses a data association tool to bind the basic information of logistics nodes with the corresponding cargo volume data (monthly) through "logistics node name" to generate a dataset of "logistics node name - administrative division - affiliated terminal - monthly cargo volume". First, the total monthly cargo volume of all logistics nodes is calculated (e.g., the total cargo volume in July 2024 is 15,000 boxes). Then, the cargo volume percentage of each logistics node is calculated according to the formula (e.g., the cargo volume of Shenzhen Port in July is 5,000 boxes, accounting for 33.3%). At the same time, the cargo volume growth rate of each logistics node is calculated (e.g., the cargo volume of Shenzhen Port in June is 4,500 boxes and in July is 5,000 boxes, with a growth rate of 11.1%). The calculation results are then added to the corresponding fields of the dataset.
[0170] Step S54: Combine the changes in logistics nodes with the proportion of cargo volume and the growth rate of cargo volume to analyze the activity and stability of each logistics node.
[0171] In this embodiment, logistics node activity is an indicator that comprehensively reflects the frequency of logistics node operations. It is mainly determined by the proportion of cargo volume (weight 40%), the cargo volume growth rate (weight 30%), and the number of new shipments (weight 30%). Higher activity indicates more frequent operations at the logistics node. Logistics node stability is an indicator that comprehensively reflects the degree of fluctuation in logistics node operations. It is determined by the fluctuation range of the cargo volume growth rate (weight 60%) and the decrease / fluctuation rate (weight 40%). Higher stability indicates more stable operations at the logistics node.
[0172] As an optional implementation, the system sets the following activity scoring rules: the top 5 logistics nodes by cargo volume share receive 40 points (out of 40), the top 6-10 receive 30 points, and the rest receive 10-20 points; a cargo volume growth rate of ≥10% receives 30 points, 5%-10% receives 20 points, 0%-5% receives 10 points, and negative growth receives 0 points; newly added logistics nodes receive an additional 30 points, and non-new nodes receive 0 points. The sum of the three scores is the total activity score (out of 100), with 80 points or above indicating high activity, 60-80 points indicating medium activity, and below 60 points indicating low activity. Stability scoring rules: 60 points for a fluctuation range of ≤10% in cargo volume growth rate, 40 points for 10%-20%, and 20 points for more than 20%; 40 points for no decrease / fluctuation, 20 points for fluctuation, and 0 points for decrease. The two scores are added together to form the total stability score (maximum score 100). 80 points or above is considered high stability, 60-80 points is medium stability, and below 60 points is low stability. The final rating of activity and stability for each logistics node is formed.
[0173] Step S55: Integrate the cargo volume data, activity level, stability and related transportation route data of each logistics node to generate logistics node analysis results that include logistics node classification, key logistics node identification results and logistics node development potential assessment.
[0174] In this embodiment, the logistics node classification is based on a comprehensive rating of activity (weight 50%) and stability (weight 50%), divided into core logistics nodes (activity ≥ 80 and stability ≥ 80), key logistics nodes (activity 60-80 and stability 60-80, or a single indicator ≥ 80), potential logistics nodes (activity 60-80 and stability < 60, or newly added nodes with activity ≥ 60), and ordinary logistics nodes (other cases). The key logistics node identification result is a list of core and key logistics nodes selected from the classification, along with key information such as cargo volume percentage and transportation routes. The logistics node development potential assessment combines cargo volume growth rate, new additions, and transportation route completeness to judge the development expectation of potential logistics nodes, such as "newly added potential logistics nodes with high growth rate and complete transportation routes are expected to be upgraded to key logistics nodes in the next 3 months."
[0175] As an optional implementation, the system integrates monthly cargo volume data, activity rating, stability rating, and related transportation route data of each logistics node, calculates a comprehensive score based on "activity × 50% + stability × 50%", and classifies logistics nodes into levels (e.g., Shenzhen Port has an activity rating of 90 and a stability rating of 85, with a comprehensive score of 87.5, making it a core logistics node; Gaolan Port is newly added, has an activity rating of 75 and a stability rating of 70, with a comprehensive score of 72.5, making it a potential logistics node). Core and key logistics nodes are selected to form key logistics node identification results, and their cargo volume share (e.g., Shenzhen Port's share is 33.3%) and main transportation routes (e.g., Shenzhen Port-CY001 route) are marked. For potential logistics nodes, their development potential over the next 3-6 months is assessed based on cargo volume growth rate (e.g., Gaolan Port's July growth rate is 25%) and transportation route completeness (e.g., two routes have been opened). Finally, a logistics node analysis result containing a level classification table, a list of key logistics nodes, and a potential assessment report is generated and exported as a PDF for use by port business departments.
[0176] For example, a port conducted a logistics node analysis on new energy goods in July 2024. First, it extracted basic information about the logistics nodes (e.g., Shenzhen Port belongs to the South China region, Yantian Terminal), monthly cargo volume for each node (Shenzhen Port 5000 TEUs, Zhuhai Gaolan Port 1500 TEUs, Guangzhou Nansha Port 0 TEUs), and transportation routes (Shenzhen Port-CY001 route, Gaolan Port-CY003 route). Then, it extracted logistics node change data from logistics change trend data, determining that Gaolan Port was newly added in July, Nansha Port had no data for two consecutive months (a decrease), and Shenzhen Port's cargo volume decreased by 30% (fluctuation). Finally, it correlated the basic information with the cargo volume data to calculate... The total cargo volume was 15,000 TEUs, with Shenzhen Port accounting for 33.3% and a growth rate of 11.1%, and Gaolan Port accounting for 10% with no historical data (growth rate is calculated as 100% for new shipments). Combining changes and cargo volume indicators, Shenzhen Port was rated as highly active (85 points) and moderately stable (65 points), while Gaolan Port was rated highly active (90 points) and moderately stable (70 points). Finally, the data was integrated and classified into levels: Shenzhen Port was identified as a key logistics node, Gaolan Port as a potential logistics node, and Nansha Port as an eliminated logistics node. A list of key logistics nodes was identified, and Gaolan Port was assessed as likely to be upgraded to a key logistics node in the next 3 months, forming the logistics node analysis results.
[0177] This embodiment accurately outputs the level and potential assessment of logistics nodes through multi-dimensional data integration and quantitative analysis, solving the problems of traditional logistics node analysis relying on experience and being highly subjective, and providing data support for ports to adjust shipping route resources and expand cargo sources.
[0178] Based on any of the above embodiments, in Embodiment Seven of this application, after step S40, refer to Figure 2 ,include;
[0179] Step S61: Combine the first historical logistics data, the logistics data, and the logistics change trend data to predict short-term results, and combine the second historical logistics data, the logistics data, and the logistics change trend data to predict long-term results.
[0180] In this embodiment, the first historical logistics data is logistics data within a preset short-term historical period, with a period shorter than the long-term result prediction period. It includes data such as cargo volume and transportation frequency of logistics nodes over the past 3-6 months, used to capture short-term data patterns. The second historical logistics data is logistics data within a preset long-term historical period, with a period of not less than 3 calendar years. It includes data such as multi-year cargo structure and changes in logistics nodes, used to identify long-term trends. Short-term results are future short-term (e.g., 1-3 months) logistics information predicted based on short-term data patterns, including short-term cargo volume estimates and short-term anomaly risk warnings. Long-term results are future long-term (e.g., 6-12 months) logistics information predicted based on long-term trends, including long-term cargo volume growth trends and the potential for logistics node expansion.
[0181] As an optional implementation, the system first calls up the first historical logistics data of the past 6 months, and inputs it into a short-term prediction model (such as the XGBoost model) along with current logistics data and logistics change trend data (abnormal records of the past month). The model learns the short-term fluctuations in cargo volume and the short-term changes in logistics nodes, and outputs short-term cargo volume forecast data for the next 2 months (such as an average cargo volume of 4,800 boxes for new energy goods at Shenzhen Port in the next 2 months) and short-term abnormal risk warnings (such as the risk of a sudden drop in cargo volume at the end of the first month or spot volume), forming short-term results. Then, it calls up the second historical logistics data of the past 5 years, and inputs it into a long-term prediction model (such as the LSTM time series model) along with current logistics data and logistics change trend data (abnormal records of the past 6 months). The model captures the evolution of cargo type structure and the long-term changes in logistics nodes over multiple years, and outputs the long-term cargo volume growth trend for the next 10 months (such as an annual growth rate of 15% for the total cargo volume of new energy goods in the next 10 months) and the expansion potential of logistics nodes (such as the expected 30% growth in cargo volume at Zhuhai Gaolan Port in the next 10 months), forming long-term results.
[0182] Step S62: Based on the short-term and long-term results, adjust the storage planning data for each cargo category and the route data for each cargo category.
[0183] In this embodiment, the storage planning data is used to guide the warehousing configuration of various types of goods at the port, including the allocation ratio of warehousing areas for each type of goods, the number of reserved storage spaces, and the matching of warehousing facility types. The route data is used to standardize the port's cargo transportation routes, including the frequency of route services for each type of goods, the scope of logistics nodes covered by the routes, and the route capacity configuration. The comprehensive short- and long-term results combine short-term demands and risk warnings from the short-term results with long-term trends and potential assessments from the long-term results, forming a comprehensive decision-making basis that considers both the present and the future.
[0184] As an optional implementation, the system first extracts short-term cargo volume estimates from short-term results (e.g., peak cargo volume of fresh produce in the next 2 months of 3200 tons) and long-term cargo volume growth trends from long-term results (e.g., annual growth rate of fresh produce in the next 10 months of 12%). Storage demand is then calculated by cargo category: peak storage demand for fresh produce in the next 2 months is 3200 tons, and average storage demand in the next 10 months is 2800 tons. Based on this, the storage planning data is adjusted (the allocation ratio of fresh produce storage areas is increased from 15% to 18%, and 200 tons of cold chain storage space is reserved). Next, short-term logistics node risks (e.g., cargo volume fluctuation risk at Guangzhou Port in the next month) and long-term logistics node expansion potential (e.g., long-term potential of Zhuhai Gaolan Port) are extracted from short-term results. Combined with current shipping route data for each cargo category, the shipping route data is adjusted: the number of sailings on the "Shenzhen Port-Shanghai Port" route for fresh produce is increased from 3 to 4 per week, and a new "Gaolan Port-Ningbo Port" route (2 sailings per month) is added, forming the adjusted storage planning data and shipping route data.
[0185] For example, a port conducts outcome prediction and data adjustment work for new energy and fresh produce goods. First, it uses the first six months of historical logistics data (average monthly volume of new energy goods at Shenzhen Port: 4500 TEUs; average monthly volume of fresh produce goods at Guangzhou Port: 2800 tons). Combined with current logistics data (current monthly volume of new energy goods at Shenzhen Port: 5000 TEUs; current monthly volume of fresh produce goods at Guangzhou Port: 3000 tons) and logistics trend data (new additions at Gaolan Port for new energy goods; fluctuations in fresh produce volume at Guangzhou Port over the past week), the XGBoost model predicts short-term results: The average volume of new energy goods at Shenzhen Port over the next two months is expected to be 4800 TEUs, and at Gaolan Port, 1600 TEUs; there is a risk of a sudden drop in fresh produce volume at Guangzhou Port at the end of the first month or in the current spot volume. Then, by utilizing the second set of historical logistics data from the past 5 years (annual growth rate of 15% for new energy and 10% for fresh produce), and combining current data with trend data, the LSTM model is used to predict long-term results: In the next 10 months, the total cargo volume of new energy is expected to grow at an annual rate of 18%, with significant potential at Gaolan Port; the total cargo volume of fresh produce is expected to grow at an annual rate of 12%, with increased demand around Shanghai Port. Based on these two sets of results, the following data adjustments are made: For new energy, storage planning data (the allocation ratio of warehousing areas is increased from 20% to 23%, with 1500 containers reserved for related warehousing at Gaolan Port), and shipping route data (the number of sailings between Shenzhen Port and Tianjin Port increases from 4 to 5 per week, and a new shipping route between Gaolan Port and Xiamen Port is added with 2 sailings per week); for fresh produce, storage planning data (cold chain storage space is increased from 3000 tons to 3200 tons), and shipping route data (the number of sailings between Guangzhou Port and Shanghai Port remains at 3 per week for now, increasing to 4 in the third month).
[0186] This embodiment achieves the adaptation of storage and route data to short- and long-term logistics needs through periodic forecasting and comprehensive adjustment, solving the problem of traditional data adjustment lagging behind demand changes and providing a scientific basis for the optimal allocation of port resources.
[0187] Optionally, the step of predicting short-term results by combining the first historical logistics data, the logistics data, and the logistics change trend data includes:
[0188] Step S71: Call the first historical logistics data within a preset time range. The first historical logistics data includes logistics node distribution data, cargo volume fluctuation data and short-term anomaly correction records within the corresponding period.
[0189] In this embodiment, the preset time range is a pre-defined time period used to extract the first historical logistics data, typically 3-6 months, which can be adjusted according to the port's business rhythm. The first historical logistics data is historical logistics business data generated within this preset time range, used to provide historical pattern references for short-term forecasting. Logistics node distribution data is information recorded in the first historical logistics data regarding the proportion of cargo volume and shipping frequency of different logistics nodes within each time period, such as "Shenzhen Port's cargo volume accounts for 35% in the past 3 months, with 4 shipments per week". Cargo volume fluctuation data records the daily fluctuations in cargo volume of each logistics node within the historical period, including the daily / weekly cargo volume change range and the annotation of the cause of the fluctuation. Short-term anomaly correction records are records of annotating and correcting short-term abnormal data (such as temporary order surges or sudden shipping stoppages) that occur within the historical period, including the abnormal time, abnormal type, and corrected data value.
[0190] As an optional implementation, the system calls the first historical logistics data within a preset time range (nearly 4 months) through the historical data interface. During the call, all logistics records within the period are first filtered by timestamp; then, the "logistics node name-volume-time" data is extracted from the records, and the volume percentage and shipping frequency of each logistics node are statistically analyzed to form logistics node distribution data; the deviation between the daily volume and the weekly average is calculated, records with deviations exceeding ±20% are marked and the reasons for the fluctuations are added (e.g., "Shenzhen Port volume deviation +25% on April 15th, due to e-commerce promotions"), forming volume fluctuation data; simultaneously, the system extracts the anomaly handling logs within the historical period, filters out records of short-term anomalies (duration ≤ 7 days), marks the anomaly type (e.g., "temporary order addition," "equipment failure causing shipping stoppage") and the corrected data, forming short-term anomaly correction records; finally, the three types of data are integrated according to the "time-logistics node" dimension and stored in a dedicated short-term prediction database.
[0191] Step S72: The real-time cargo volume information and logistics node association information in the logistics data are fused with the short-term abnormal change data in the logistics change trend data to form a predictive input dataset.
[0192] In this embodiment, logistics data refers to the core real-time logistics data generated within the current business cycle. Real-time cargo volume information is the actual cargo volume data recorded at each logistics node in the logistics data, including the real-time cargo volume value, unit of volume, and statistical time point. Logistics node association information is auxiliary information bound to logistics nodes in the logistics data, such as the transportation method, partner company, and warehouse location corresponding to the logistics node. Short-term abnormal change data in the logistics change trend data records information on sudden increases or decreases in cargo volume at logistics nodes in the recent period (within 1-2 months), and information on temporary additions / discontinuations of logistics nodes, such as "Zhuhai Gaolan Port saw a 40% month-on-month increase in cargo volume in the past month." The prediction input dataset is a data set used to input the short-term prediction model, formed by integrating real-time data and short-term abnormal change data, and is formatted as a structured data table.
[0193] As an optional implementation, the system extracts "logistics node name - real-time cargo volume (average of the last 3 days) - transportation mode - partner company" from the current logistics data as the real-time cargo volume information and logistics node association information; it filters short-term abnormal change data within the last 1.5 months from the logistics change trend data, including "abnormal logistics node - abnormality type - cargo volume change range - abnormal duration"; it associates the two types of data through the "logistics node name" field, performs cross-validation on the real-time cargo volume and abnormal change data of the same logistics node (such as confirming whether the real-time cargo volume is affected by abnormal changes), removes duplicate and redundant data, and supplements data integrity (such as marking "no short-term abnormality" for logistics nodes without abnormal changes), and finally forms a prediction input dataset containing the fields "logistics node - real-time cargo volume - association information - short-term abnormality", and converts the format to CSV format that the model can recognize.
[0194] Step S73: Input the predicted input dataset and the first historical logistics data into the short-term prediction model, and learn the short-term fluctuation patterns and abnormal influence weights in the historical data through the model.
[0195] In this embodiment, the short-term forecasting model is an algorithmic model used to achieve short-term freight flow forecasting. Models such as XGBoost and LightGBM, which are suitable for processing structured data and offer high prediction accuracy, are typically employed. Short-term fluctuation patterns are the short-term variation patterns of freight volume and logistics node flow presented in the first set of historical logistics data, such as "freight volume on Mondays is 15% higher than on weekends, and freight volume at the end of each month is 20% higher than in the middle of the month." Anomaly impact weights are numerical values calculated by the model through learning from historical anomaly data, determining the degree of impact of different types of short-term anomalies on freight volume and logistics node flow, such as "the impact weight of temporary order surge anomalies on freight volume is +30%, and the impact weight of equipment failure anomalies is -25%."
[0196] As an optional implementation, the system divides the prediction input dataset (feature data) and the first historical logistics data (label data, with "shipment volume in the next month" as the prediction label) into a training set and a validation set in a 7:3 ratio. The training set is input into a short-term prediction model (using the XGBoost model). The model first learns the short-term fluctuation patterns in the first historical logistics data through a gradient boosting algorithm, identifying the "weekly / monthly shipment volume fluctuation cycle" and the "logistics node traffic change rhythm". Then, for short-term anomaly correction records in the historical data, the impact of different anomaly types on shipment volume is calculated, and an anomaly impact weight table is generated (e.g., "e-commerce promotion anomaly weight +28%, weather impact anomaly weight -18%)". The model is then optimized using the validation set, adjusting parameters such as the learning rate and tree depth to ensure that the model prediction error rate is below 8%, thus completing model training and pattern learning.
[0197] Step S74: Based on the prediction results output by the model, generate short-term results that include short-term cargo volume forecast data, logistics node traffic change forecast data, and short-term abnormal risk warnings.
[0198] In this embodiment, the model output is a set of predicted values for logistics data in the short term (e.g., 1-3 months), including predicted cargo volume, predicted flow changes, and anomaly probabilities for each logistics node. Short-term cargo volume estimates are generated based on the prediction results, providing estimated cargo volume information for each logistics node in the short term, including daily / weekly estimated cargo volume and estimated cargo volume ranges (e.g., "Shenzhen Port's estimated weekly cargo volume for the next month is 4800-5200 containers"). Logistics node flow change estimates are the predicted changes in shipping frequency and cargo volume percentage for each logistics node in the short term, such as "Zhuhai Gaolan Port's shipping frequency will increase from 2 times per week to 3 times per week in the next 2 months, and its cargo volume percentage will increase from 8% to 12%." Short-term anomaly risk warnings are based on the model's predicted anomaly probabilities and provide early warnings of potential short-term anomalies, including potential anomaly types, estimated anomaly times, and risk levels (high / medium / low).
[0199] As an optional implementation, the system extracts the prediction results output by the short-term prediction model and organizes them according to the "logistics node-time" dimension: it summarizes the average daily estimated cargo volume of each logistics node in the next two months, marks the range of fluctuation within 10%, and forms short-term cargo volume prediction data; it calculates the difference between the delivery frequency and cargo volume ratio of each logistics node in the next two months and the current value, marks the growth / decline trend, and forms logistics node traffic change prediction data; it filters out cases where the model predicts an anomaly probability ≥30%, and combines historical anomaly types to mark potential anomalies (such as "the cargo volume of Guangzhou Port may decrease due to warehouse maintenance in the third week of the future"), the estimated duration of the anomaly (3 days), and the risk level (medium risk), forming a short-term anomaly risk warning; finally, it integrates the three types of information according to "logistics node classification" to generate a structured short-term result report, including data tables and trend charts, and stores it in the prediction result module of the port business system.
[0200] For example, a port conducts short-term forecasting for electromechanical goods. First, it retrieves historical logistics data for a preset timeframe (the past 3 months) and extracts data on logistics node distribution (Shenzhen Port's cargo volume share 38%, Guangzhou Port 25%), cargo volume fluctuations (Monday cargo volume is 12% higher than weekend volume), and short-term anomaly correction records (Shenzhen Port's cargo volume increased by 22% two months ago due to temporary order additions; the corrected data is now labeled). Next, it extracts real-time cargo volume information from current logistics data (Shenzhen Port's average over the past 3 days is 5100 TEUs, Guangzhou Port 3200 TEUs) and logistics node association information (Shenzhen Port's transportation mode is sea freight, partner company A). This information is then combined with short-term anomaly data from logistics change trend data (Guangzhou Port's cargo volume decreased by 15% month-on-month due to partner company order adjustments) to form the prediction input dataset. This dataset, along with the historical logistics data, is then input into the XGBoost short-term forecasting model. After learning patterns such as "weekly cargo volume peak" and "temporary order addition anomaly weight +20%", the model outputs the prediction results. Finally, based on the forecast results, short-term results are generated: short-term cargo volume forecast (Shenzhen Port's weekly cargo volume is 4900-5300 TEUs and Guangzhou Port's is 3000-3400 TEUs in the next month), logistics node traffic change forecast (Shenzhen Port's shipping frequency remains at 4 times per week, while Guangzhou Port's is reduced to 3 times per week), and short-term abnormal risk warning (Guangzhou Port may experience a one-day decrease in cargo volume due to heavy rain in the second week of the future, with a low risk level).
[0201] This embodiment uses systematic historical data retrieval, data fusion, and model learning to accurately generate short-term results, solving the problems of traditional short-term forecasting relying on experience and having large errors, and providing reliable data support for port short-term warehousing and route adjustments.
[0202] Optionally, the step of predicting long-term results by combining the second historical logistics data, the logistics data, and the logistics change trend data includes:
[0203] Step S81: Call up the second historical logistics data with a time span greater than or equal to a preset number of years. The second historical logistics data includes logistics node change data, cargo structure evolution data, and long-term abnormal impact records within the corresponding year.
[0204] In this embodiment, the preset number of years is a pre-defined standard used to define the time span of the second historical logistics data, typically three calendar years or more, to ensure that the data reflects long-term trends. The second historical logistics data is historical logistics business data accumulated within this time span, used to provide multi-period pattern support for long-term forecasting. Logistics node change data records information on the addition, elimination, and long-term changes in cargo volume proportion of logistics nodes over multiple years, such as "from 2021 to 2024, the cargo volume proportion of Shenzhen Port increased from 30% to 38%, and Zhuhai Gaolan Port saw new additions in 2023 that gradually stabilized." Cargo category structure evolution data records information on the cargo volume proportion and demand changes of various types of goods over multiple years, such as "from 2021 to 2024, the cargo volume proportion of new energy goods increased from 15% to 25%, and traditional building materials decreased from 20% to 12%." Long-term abnormal impact records are records that mark abnormal events (such as policy adjustments and industry cycle changes) that have a long duration (e.g., more than 3 months) and a wide impact over multiple years. They include the type of abnormal event, the period of impact, and the specific impact on cargo volume / logistics nodes.
[0205] As an optional implementation, the system accesses second-historical logistics data spanning five calendar years (2019-2023) through a long-term historical data interface. During access, complete logistics records for each year are filtered by annual timestamps. From these records, the system statistically analyzes the proportion of cargo volume and the status of new / outdated shipments at each logistics node by year, forming logistics node change data. It also statistically analyzes the proportion of cargo volume and demand growth rate for each cargo type by year, forming cargo structure evolution data. Furthermore, it extracts logs of abnormal events lasting more than three months within each year, labeling the event type, impact period, and magnitude of impact on cargo volume (e.g., "cargo volume decreased by 18% year-on-year"), forming long-term abnormal impact records. Finally, the three types of data are integrated according to the dimensions of "year-cargo-logistic node" and stored in a dedicated long-term forecasting database, while simultaneously generating a data integrity report (ensuring no missing data for each year).
[0206] Step S82: Extract the cargo category proportion information and long-term cooperative logistics node information from the logistics data, and integrate them with the long-term trend change data in the logistics change trend data to construct a long-term prediction input dataset.
[0207] In this embodiment, the cargo category percentage information in the logistics data refers to the cargo volume percentage of various goods within the current business cycle (e.g., the past month), including specific values and month-on-month changes, such as "the current monthly cargo volume percentage of new energy goods is 26%, a month-on-month increase of 2%." The long-term cooperative logistics node information refers to information related to logistics nodes that maintain stable cooperation with the port (e.g., cooperation for more than one year), including the logistics node name, cooperation duration, and average annual cargo volume contribution, such as "Shenzhen Port has cooperated for 5 years, with an average annual cargo volume contribution of 40,000 containers." The long-term trend change data in the logistics change trend data records logistics data that shows a continuous changing trend over the past 1-2 years (not short-term fluctuations), such as "the cargo volume of Zhuhai Gaolan Port has continued to grow in the past year, with an average monthly growth rate of 5%; the demand for fresh produce has continued to rise in the past 2 years, with an annual growth rate of 12%." The long-term prediction input dataset is a structured dataset formed by integrating current data and long-term trend data, used as input to the long-term prediction model.
[0208] As an optional implementation, the system extracts "goods category name - current monthly percentage - month-on-month change" as goods category percentage information from current logistics data, and extracts "logistics node name - cooperation duration - average annual cargo volume contribution" as long-term cooperative logistics node information; it filters long-term trend change data within the past two years from logistics change trend data, including "logistics node name - average monthly growth rate - duration" and "goods category name - annual growth rate - duration"; it associates the three types of data through the "goods category name" and "logistics node name" fields, cross-validates the current data and trend data of the same goods category / logistics node (e.g., confirming whether the current goods category percentage conforms to the long-term growth trend), removes short-term fluctuation data (e.g., a logistics node temporarily experiences a surge in orders in one month), and supplements data annotations (e.g., labeling goods categories without long-term trend changes as "stable"); finally, it forms a long-term prediction input dataset containing the fields "goods category - current percentage - long-term growth rate - logistics node - cooperation duration - long-term growth rate", converts the format to JSON format that the model can recognize, and performs data standardization processing (e.g., converting growth rate data into normalized values in the 0-1 range).
[0209] Step S83: Input the long-term prediction input dataset and the second historical logistics data into the long-term prediction model, and use the model to capture the cumulative impact of multi-year evolution patterns and trend anomalies in the historical data.
[0210] In this embodiment, the long-term forecasting model is an algorithmic model used to achieve long-term logistics forecasting. It typically employs models suitable for processing time-series data and capturing long-term trends, such as LSTM (Long Short-Term Memory) networks and Prophet. The multi-year evolution pattern refers to the stable change pattern across multiple years presented in the second historical logistics data, such as "the volume of new energy goods in Q4 is always the peak of the year, increasing by 30% compared to Q1; the volume of goods from logistics nodes with cooperation for more than 3 years increases by an average of 8% annually." The cumulative impact of trend anomalies refers to the gradual accumulation of the effects of long-term abnormal events (such as policy adjustments) over multiple years.
[0211] As an optional implementation, the system divides the long-term prediction input dataset (feature data) and the second historical logistics data (labeled data, with "monthly cargo volume for the next 12 months" as the prediction label) into a training set and a validation set in an 8:2 ratio. The training set is input into the long-term prediction model (using an LSTM time series model). The model learns the multi-year evolution patterns in the second historical logistics data through a multi-layer neural network, such as identifying "seasonal cycles of cargo types" and "long-term growth curves of logistics nodes". For long-term abnormal impact records, the model calculates the cumulative impact coefficient of abnormal events over multiple years (e.g., "the annual cumulative impact coefficient of environmental protection policies on traditional cargo types is -7%) and incorporates it into the trend prediction logic. The model is optimized through the validation set, adjusting parameters such as the number of hidden layers and the learning rate to ensure that the average error rate of the model's cargo volume prediction for the next 12 months is less than 10%, thus completing model training and pattern capture.
[0212] Step S84: Based on the prediction results output by the model, generate long-term results including monthly cargo volume growth trend data, logistics node expansion potential data, and cargo structure adjustment direction.
[0213] In this embodiment, the model output is a set of predicted values for long-term (e.g., 12 months) logistics data, including predicted cargo volume for each type of cargo / logistics node in each month, growth potential score, and probability of structural changes. Monthly cargo volume growth trend data, generated based on the prediction results, shows the cargo volume growth for each type of cargo / logistics node in each month over the next 12 months, including monthly predicted cargo volume, year-on-year / month-on-month growth rate, and growth trend curve, such as "average monthly cargo volume of new energy goods in the next 12 months is 5000 boxes, with a year-on-year growth rate of 15%, peaking in Q4." Logistics node expansion potential data, generated based on the prediction results, assesses the development potential of each logistics node over the next 12 months, including a potential score (0-100 points), estimated cargo volume growth space, and expansion suggestions, such as "Zhuhai Gaolan Port has a potential score of 85 points, with an estimated cargo volume growth of 40% in the next 12 months; it is recommended to add one dedicated shipping route." The direction of cargo structure adjustment is an optimization suggestion for the cargo structure within the next 12 months, generated based on the forecast results. It includes the priority adjustment of each cargo category and the direction of resource allocation. For example, "the priority of new energy cargo is raised to level one, and it is recommended to increase cold chain warehousing resources; the priority of traditional building materials is reduced to level three, and warehousing space is appropriately reduced."
[0214] As an optional implementation, the system extracts the prediction results output by the long-term prediction model and organizes them according to the dimensions of "monthly-cargo-logistics node": It summarizes the predicted cargo volume for each cargo category / logistics node for each month of the next 12 months, calculates the year-on-year / month-on-month growth rate, generates a monthly cargo volume growth trend curve, and forms monthly cargo volume growth trend data; based on the estimated cargo volume growth potential (e.g., growth of over 30% is high potential) and stability (e.g., fluctuation range of less than 10% is stable) of each logistics node over the next 12 months, it calculates a potential score, marks high-potential logistics nodes, and provides expansion suggestions (e.g., "new..."). The process involves: increasing cooperation with enterprises and optimizing transportation routes to generate data on the potential for expanding logistics nodes; analyzing the growth rate of each cargo category over the next 12 months (e.g., growth rate above 15% is considered high growth) and market demand (e.g., policy-supported categories are considered high demand) to determine the priority of each cargo category (Level 1 / Level 2 / Level 3), providing suggestions for adjusting warehousing and shipping resources, and forming a direction for cargo category structure adjustment; finally, integrating the three types of information by "cargo category" to generate a long-term results report containing data tables, trend charts, and a list of adjustment suggestions, exporting it as a PDF and synchronizing it to the port's strategic planning department's business system.
[0215] For example, a port conducted long-term outcome forecasting for new energy and traditional building materials. First, it retrieved second-historical logistics data spanning five calendar years (2019-2023) and extracted data on changes in logistics nodes (Shenzhen Port's cargo volume share increased from 30% to 38% over five years, and Zhuhai Gaolan Port saw new additions in 2023), cargo structure evolution data (new energy cargo volume share increased from 15% to 25% over five years, while traditional building materials decreased from 20% to 12%), and records of long-term anomalies (environmental policies in 2022 led to a 10% year-on-year decrease in traditional building materials cargo volume, with the impact continuing into 2023). Next, cargo category percentage information is extracted from current logistics data (new energy currently accounts for 26% of the total, up 2% month-on-month; traditional building materials currently account for 11% of the total, down 1% month-on-month), and information on long-term cooperative logistics nodes (Shenzhen Port has cooperated for 5 years with an average annual cargo volume of 40,000 TEUs; Guangzhou Port has cooperated for 4 years with an average annual cargo volume of 25,000 TEUs). This data is then integrated with long-term trend data from logistics change trend data (Zhuhai Gaolan Port's average monthly growth rate is 5% in the past year; the annual growth rate of new energy has been 12% in the past 2 years) to construct a long-term prediction input dataset. Subsequently, this dataset, along with second-generation historical logistics data, is input into an LSTM long-term prediction model. After learning patterns such as "peak cargo volume of new energy in Q4 each year" and "the continued impact of environmental protection policies on traditional building materials," the model outputs prediction results. Finally, based on the forecast results, long-term results are generated: monthly cargo volume growth trend data (average monthly cargo volume of new energy in the next 12 months is 5,000 boxes, an increase of 15% year-on-year; average monthly cargo volume of traditional building materials is 2,000 boxes, a decrease of 8% year-on-year), logistics node expansion potential data (Zhuhai Gaolan Port has a potential score of 85 points, with an estimated growth of 40%, and it is recommended to add 1 dedicated shipping route), and cargo structure adjustment direction (new energy is prioritized to level 1 and cold chain warehousing is increased; traditional building materials are prioritized to level 3 and warehousing space is reduced by 10%).
[0216] This embodiment uses multi-year historical data and long-term trend model learning to accurately generate long-term results, solving the problem that traditional long-term forecasts are unable to capture cross-year patterns and ignore the impact of trend anomalies, thus providing a scientific basis for port long-term strategic planning and resource allocation.
[0217] Optionally, the steps of adjusting the storage planning data for each cargo category and the route data for each cargo category by combining the short-term and long-term results include:
[0218] Step S91: Extract short-term cargo volume forecast data and short-term abnormal risk warnings from the short-term results, and integrate them with the cargo volume growth trend data and cargo structure adjustment direction from the long-term results for the second future quantity period to generate a comprehensive analysis dataset.
[0219] In this embodiment, the short-term cargo volume forecast data refers to the cargo volume data of each cargo type / logistics node predicted in the first future quantity time period in the short-term results, including specific values and fluctuation ranges. The short-term anomaly risk warning is a warning of possible anomalies in the first future quantity time period, including anomaly type and risk level. The cargo volume growth trend data for the second future quantity time period is the predicted cargo volume growth pattern of each cargo type over a longer future period in the long-term results, including monthly growth rate and peak period. The cargo type structure adjustment direction is the priority and resource allocation suggestions for each cargo type given in the long-term results. The comprehensive analysis dataset is a dataset formed by integrating highly correlated data from the short-term and long-term results according to a unified dimension, used for subsequent storage and route data adjustment. The first quantity time period is the prediction period of the short-term results (e.g., 2 months), and the second quantity time period is the prediction period of the long-term results (e.g., 10 months), with the first quantity time period being shorter than the second quantity time period.
[0220] As an optional implementation, the system extracts "Goods Category Name - Estimated Monthly Goods Quantity for the First Quantity Period (2 Months) - Short-Term Anomaly Type - Risk Level" from short-term results and "Goods Category Name - Monthly Growth Rate for the Second Quantity Period (10 Months) - Goods Category Priority - Resource Allocation Direction" from long-term results. The two types of data are linked through the "Goods Category Name" field. Logical verification is performed on the short-term and long-term growth trends of the same goods category (e.g., confirming whether the estimated goods quantity for the first quantity period conforms to the growth pattern of the second quantity period). Supplementary data association annotations are added (e.g., for goods categories with short-term anomalies, the annotation "Adjustment of response strategy based on long-term trend" is used). Finally, a comprehensive analysis dataset containing the fields "Goods Category - Estimated Goods Quantity for the First Time Period - Short-Term Anomaly - Growth Rate for the Second Time Period - Goods Category Priority" is formed, stored in a dedicated data fusion database, and a data fusion report is generated (annotating data sources and association logic).
[0221] Step S92: Based on the comprehensive analysis dataset, calculate the peak storage demand for the first quantity time period and the average storage demand for the second quantity time period in the future, according to the category of goods, and adjust the storage planning data for each category of goods. The storage planning data includes the allocation ratio of storage areas and the number of reserved storage spaces.
[0222] In this embodiment, the peak storage demand for the first future time period is the maximum storage demand calculated by cargo type over a relatively short period, requiring the reservation of emergency storage space, such as "peak storage demand for new energy products in the next 2 months: 5500 containers". The average storage demand for the second future time period is the average storage demand calculated by cargo type over a relatively long period, used for planning regular storage configuration, such as "average storage demand for new energy products in the next 10 months: 5000 containers". Storage planning data is the core data guiding the allocation of port storage resources, where the storage area allocation ratio is the percentage of total storage space occupied by each cargo type, and the number of reserved storage spaces is the number of storage locations (including regular and emergency storage spaces) reserved in advance for each cargo type.
[0223] As an optional implementation, the system extracts the estimated monthly cargo volume for the first quantity period (2 months) from the comprehensive analysis dataset by cargo type, and takes the maximum value as the peak storage demand for that cargo type (e.g., the estimated cargo volume for new energy in February is 5200 boxes and 5500 boxes, with a peak of 5500 boxes); it extracts the estimated monthly cargo volume for the second quantity period (10 months) (calculated based on long-term growth rate), and takes the average value as the average storage demand (e.g., the average estimated cargo volume for new energy over 10 months is 5000 boxes); combined with the total port storage capacity (e.g., 50000 boxes), it calculates the storage demand for each cargo type. The allocation ratio of storage areas (new energy category = 5000 / 50000 × 100% = 10%) is determined, and the number of reserved storage spaces is determined based on 1.2 times the peak storage demand (new energy category reserved 5500 × 1.2 = 6600 boxes). Based on this, the storage planning data is adjusted. For example, the allocation ratio of new energy storage areas is increased from the original 8% to 10%, and the number of reserved storage spaces is increased from 5000 boxes to 6600 boxes. For traditional building materials, due to the decrease in the average demand during the second time period, the allocation ratio is reduced from 12% to 10%, and the number of reserved storage spaces is reduced from 4500 boxes to 4000 boxes.
[0224] Step S93: Extract the estimated data on logistics node traffic changes and logistics node expansion potential data for each cargo category from the comprehensive analysis dataset. Combine this with the current route capacity configuration information to adjust the route data for each cargo category. The route data includes route frequency and the range of logistics nodes covered by the route.
[0225] In this embodiment, the logistics node traffic change prediction data is a comprehensive analysis of the data, predicting the cargo flow changes of each cargo type at corresponding logistics nodes within a first time period in the future, such as "new energy category: Shenzhen Port traffic increases by 10% in the next 2 months, Guangzhou Port decreases by 5%". The logistics node expansion potential data is a comprehensive analysis of the data, identifying logistics nodes with high growth potential within a second time period in the future, such as "new energy category: Zhuhai Gaolan Port's potential score is 85 points in the next 10 months". The current route capacity configuration information is the port's existing route transportation capacity data, including the single-trip capacity, current frequency, and covered logistics nodes for each route. Route data is the core data guiding port transportation route operations, where route frequency is the number of times each route operates weekly / monthly, and the route coverage of logistics nodes is the list of logistics nodes served by each route.
[0226] As an optional implementation, the system extracts estimated data on logistics node traffic changes (e.g., a 10% increase in Shenzhen Port's traffic for new energy goods) and data on the potential for logistics node expansion (e.g., a potential score of 85 for Zhuhai Gaolan Port) from the comprehensive analysis dataset by cargo type; retrieves current route capacity configuration information (e.g., the "Shenzhen Port-Tianjin Port" route has a single-trip capacity of 2000 TEUs and currently operates 3 times per week, but does not yet cover Zhuhai Gaolan Port); and calculates the required additional trips for logistics nodes with increasing traffic (a 10% increase in Shenzhen Port's traffic requires 0.3 additional trips per week, upwards...). Rounded down to 1, the frequency of the "Shenzhen Port - Tianjin Port" route will be increased from 3 to 4 per week; for high-potential logistics nodes, corresponding routes will be added (such as the "Zhuhai Gaolan Port - Xiamen Port" route, with 2 flights per week and a single trip capacity of 1,500 TEUs), expanding the range of logistics nodes covered by the routes; due to a 5% decrease in traffic at Guangzhou Port, the frequency of the "Guangzhou Port - Shanghai Port" route for traditional building materials will be reduced from 2 to 1 per week, forming the adjusted route data, while also indicating the basis for the adjustment (such as "adjusting the frequency based on a 10% increase in traffic at Shenzhen Port").
[0227] Step S94: Output the adjusted cargo category storage planning data and route data, wherein the duration of the first quantity time period is shorter than the duration of the second quantity time period, and the first quantity time period corresponds to the prediction period of short-term results, while the second quantity time period corresponds to the prediction period of long-term results.
[0228] In this embodiment, the output of adjusted storage planning data and route data involves presenting the data adjusted in the aforementioned steps in a standardized format and synchronizing it to the relevant business systems for easy execution by various departments of the port. The first quantity time period corresponds to the prediction period for short-term results (e.g., 2 months), and the second quantity time period corresponds to the prediction period for long-term results (e.g., 10 months). The difference in duration ensures that data adjustments take into account both short-term emergency needs and long-term planning needs.
[0229] As an optional implementation, the system organizes the adjusted storage planning data into the format of "cargo type name - storage area allocation ratio - storage space reservation quantity (regular + emergency)" and the route data into the format of "cargo type name - route name - frequency of service - covered logistics nodes - single-trip capacity". The system checks the logical consistency of the two types of data using data verification tools (e.g., an increase in the storage allocation ratio of a certain cargo type requires corresponding route capacity matching). After confirming that there are no errors, the system generates a "storage planning adjustment report" and a "route adjustment report". The reports are synchronized to the port storage management system and the route operation system, and are also pushed to the storage department and the route scheduling department via email. The reports clearly indicate the definitions of the first quantity time period (2 months) and the second quantity time period (10 months), explaining that short-term adjustments (such as emergency storage space reservation) are based on the needs of the first quantity time period, and long-term adjustments (such as storage ratio allocation) are based on the needs of the second quantity time period, ensuring that all departments understand the adjustment logic.
[0230] For example, a port conducts storage planning and route data adjustments for new energy and traditional building materials goods. First, from short-term results, the estimated cargo volume (5200 TEUs, 5500 TEUs) for the first time period (2 months) of new energy goods, along with short-term anomaly risk warnings (equipment maintenance risk at the end of the second month or currently), are extracted. From long-term results, the growth rate (monthly average 1.2%) for the second time period (10 months) of new energy goods and their cargo priority (Level 1), and the growth rate (monthly average -0.7%) and cargo priority (Level 3) for traditional building materials goods in the second time period are extracted and merged to generate a comprehensive analysis dataset. Based on this dataset, the peak storage demand for new energy goods in the first time period is 5500 TEUs, and the average in the second time period is 5000 TEUs. The storage planning data is adjusted to a storage area allocation ratio of 10% and a storage space reserve of 6600 TEUs. For traditional building materials goods, the peak demand in the first time period is 3800 TEUs, and the average in the second time period is 3500 TEUs. The allocation ratio is adjusted to 10% and a reserve of 4000 TEUs. Based on the estimated 10% increase in traffic at Shenzhen Port for new energy vehicles and the potential score of 85 points at Zhuhai Gaolan Port, and considering current shipping capacity, the number of sailings on the "Shenzhen Port-Tianjin Port" route will be increased from 3 to 4, and a new "Gaolan Port-Xiamen Port" route will be added (2 sailings per week). For traditional building materials vehicles, traffic at Guangzhou Port will decrease by 5%, and the number of sailings on the "Guangzhou Port-Shanghai Port" route will be reduced from 2 to 1. Finally, the adjusted data will be output, labeled with the first time period of 2 months and the second time period of 10 months.
[0231] This embodiment achieves the adaptation of stored and route data to different periodic requirements through the fusion of short-term and long-term data and targeted adjustments. It solves the problem that traditional adjustments that focus solely on the short-term or long-term can easily lead to resource waste, and provides data support for the efficient operation of ports.
[0232] This application provides a port data processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the port data processing method in the first embodiment described above.
[0233] The following is for reference. Figure 3 The diagram illustrates a structural schematic of a port data processing device suitable for implementing embodiments of this application. The port data processing device in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablets, and in-vehicle terminals, as well as fixed terminals such as digital TVs and desktop computers. Figure 3 The port data processing equipment shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0234] like Figure 3 As shown, the port data processing device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the port data processing device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the port data processing equipment to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows a port data processing equipment with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.
[0235] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0236] The port data processing equipment provided in this application employs the port data processing method described in the above embodiments to solve the technical problem of lagging adjustments in shipping route resources and warehousing configurations, making it difficult to match dynamic cargo demand. Compared with the prior art, the beneficial effects of the port data processing equipment provided in this application are the same as those of the port data processing equipment provided in the above embodiments, and other technical features of this port data processing equipment are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0237] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0238] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0239] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the port data processing method in the above embodiments.
[0240] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), or any suitable combination thereof.
[0241] The aforementioned computer-readable storage medium may be included in the port data processing equipment; or it may exist independently and not be assembled into the port data processing equipment.
[0242] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by a port data processing device, the port data processing device: in response to input logistics order information, determines the semantic and structural information corresponding to the logistics order information; based on preset mapping rules, maps the semantic and structural information into structured text containing cargo classification information; parses logistics data according to a knowledge graph pre-constructed based on enterprise information, logistics auxiliary data associated with the structured text, and the structured text; analyzes the static and dynamic anomaly results of the time-series data corresponding to the logistics data to obtain logistics change trend data, which includes logistics node change data and cargo volume change data; and analyzes the logistics data and the logistics change trend data to obtain logistics node analysis results.
[0243] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0244] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0245] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0246] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the port data processing method described above. This solves the technical problem of lagging adjustments to shipping route resources and warehousing configurations, making it difficult to match dynamic cargo demand. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the port data processing method provided in the above embodiments, and will not be elaborated upon here.
[0247] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the port data processing method described above.
[0248] The computer program product provided in this application can solve the technical problem of lagging adjustments in shipping route resources and warehouse configurations, making it difficult to match dynamic cargo demand. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the port data processing method provided in the above embodiments, and will not be repeated here.
[0249] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A method of processing port data, characterized by, The port data processing method comprises: In response to the input logistics order information, the semantic information and the structural information corresponding to the logistics order information are determined; Based on the preset mapping rule, the semantic information and the structural information are mapped into structured text containing cargo classification information; According to the knowledge graph constructed in advance based on enterprise information, the logistics auxiliary data associated with the structured text, and the structured text, the logistics data is parsed; The static abnormal result and the dynamic abnormal result of the time sequence data corresponding to the logistics data are parsed to obtain logistics change trend data, and the logistics change trend data includes logistics node change data and cargo volume change data; The logistics node analysis result is obtained by analyzing the logistics data and the logistics change trend data.
2. The method of processing port data according to claim 1, wherein, The step of determining the semantic information and the structural information corresponding to the input logistics order information comprises: Receiving input logistics order information, the logistics order information contains unstructured text information and structured field information; The unstructured text information is split by a text segmentation algorithm, and the split result is converted into semantic information that can represent text semantic association attributes based on a word embedding model; According to the preset data specification, the structured field information is format checked to eliminate invalid fields and correct abnormal data, and the standardized structural information is obtained; The semantic features and the standardized structured data are associated through the unique identification of the logistics order to form an associated data set.
3. The method of claim 1, wherein, The step of mapping the semantic information and the structural information into structured text containing cargo classification information based on the preset mapping rule comprises: Calling a preset mapping rule set, the mapping rule set contains semantic feature and cargo classification association rules, and structured data and classification supplement rules; The semantic information is input into the mapping rule set, and the corresponding cargo classification candidate result is preliminarily determined by matching the semantic feature with the association rule; Extract the cargo volume value and transportation mode related data in the structural information, and check the cargo classification candidate result in combination with the classification supplement rule to eliminate the candidate classification that does not match the structured data; Based on the checked cargo classification result, the semantic information, the structural information, and the cargo classification information are integrated according to the preset field format to generate structured text containing cargo classification information.
4. The method of claim 1, wherein, The step of parsing the logistics data according to the knowledge graph constructed in advance based on enterprise information, the logistics auxiliary data associated with the structured text, and the structured text comprises: Calling the knowledge graph constructed in advance based on enterprise business information and historical cooperation information, the knowledge graph contains the association relationship between enterprise entities and regional information; Extract the shipper enterprise identifier and cargo classification information from the structured text, match the target enterprise entity in the knowledge graph based on the enterprise identifier, and obtain the regional associated data corresponding to the target enterprise; Call the logistics auxiliary data associated with the structured text, and the logistics auxiliary data contains historical transportation path and warehouse node record; The regional association data, the logistics auxiliary data, and the cargo volume information and the customs declaration time information in the structured text are fused and analyzed, and after eliminating data conflict items, logistics data containing logistics node information, cargo volume statistical information, and transportation path association information are parsed.
5. The method of claim 1, wherein, The step of parsing static abnormal results and dynamic abnormal results of the time sequence data corresponding to the logistics data to obtain logistics change trend data, the logistics change trend data including logistics node change data and cargo volume change data, comprises: Extracting time sequence data corresponding to each dimension data in the logistics data, the time sequence data being divided according to a preset time period, and each time sequence data being associated with corresponding logistics node information and cargo volume information; Performing static abnormal analysis on the time sequence data, determining static abnormal results by calculating the deviation value of data in each time period and historical data in the same period, and comparing a preset deviation threshold value; Performing dynamic abnormal analysis on the time sequence data, fitting a data change trend by a time sequence prediction model, identifying the deviation of actual data from the fitted trend, and determining dynamic abnormal results; Integrating the static abnormal results and the dynamic abnormal results, screening out abnormal data related to logistics node information changes to generate logistics node change data, screening out abnormal data related to cargo volume information changes to generate cargo volume change data, and combining the logistics node change data and the cargo volume change data into logistics change trend data.
6. The method of claim 1, wherein, The step of analyzing the logistics data and the logistics change trend data to obtain logistics node analysis results comprises: Extracting logistics node basic information, cargo volume data corresponding to each logistics node, and logistics node associated transportation path data from the logistics data; Extracting logistics node change data from the logistics change trend data to determine the addition, reduction, and fluctuation of each logistics node; Associating the logistics node basic information with the corresponding cargo volume data to calculate the cargo volume proportion and the cargo volume growth rate of each logistics node; Combining the logistics node change situation with the cargo volume proportion and the cargo volume growth rate to analyze the activity and stability of each logistics node; Integrating the cargo volume data, activity, stability, and associated transportation path data of each logistics node to generate logistics node analysis results including logistics node level division, key logistics node identification results, and logistics node development potential evaluation.
7. The method of claim 1, wherein, After the step of parsing static abnormal results and dynamic abnormal results of the time sequence data corresponding to the logistics data to obtain logistics change trend data, the logistics change trend data including logistics node change data and cargo volume change data, comprises: Combining the first historical logistics data, the logistics data, and the logistics change trend data to predict short-term results, and combining the second historical logistics data, the logistics data, and the logistics change trend data to predict long-term results; Adjusting the storage planning data of each cargo category and the route data corresponding to each cargo category based on the short-term results and the long-term results.
8. The method of processing port data according to claim 7, wherein, The step of combining the first historical logistics data, the logistics data, and the logistics change trend data to predict short-term results comprises: Call the first historical logistics data in the preset time range, which contains the logistics node distribution data, the cargo volume fluctuation data and the short-term abnormal correction record in the corresponding period; Fuse the real-time cargo volume information, the logistics node association information in the logistics data with the short-term abnormal change data in the logistics change trend data to form a prediction input data set; Input the prediction input data set and the first historical logistics data into a short-term prediction model to learn the short-term fluctuation rule and the abnormal influence weight in the historical data through model learning; Based on the prediction result output by the model, generate a short-term result containing short-term cargo volume estimation data, logistics node flow change estimation data and short-term abnormal risk prompt.
9. The method of processing port data of claim 7, wherein, The step of predicting a long-term result by combining the second historical logistics data, the logistics data and the logistics change trend data includes: Call the second historical logistics data with a time span greater than or equal to a preset number of years, which contains the logistics node transition data, the cargo class structure evolution data and the long-term abnormal influence record in the corresponding year; Extract the cargo class proportion information and the long-term cooperative logistics node information in the logistics data, and integrate them with the long-term trend change data in the logistics change trend data to build a long-term prediction input data set; Input the long-term prediction input data set and the second historical logistics data into a long-term prediction model to capture the multi-year evolution rule and the cumulative influence of trend anomalies in the historical data through model learning; Based on the prediction result output by the model, generate a long-term result containing monthly cargo volume growth trend data, logistics node expansion potential data and cargo class structure adjustment direction.
10. The method of claim 7, wherein the port data is processed by: The step of adjusting the storage planning data of each cargo class and the route data of each cargo class by combining the short-term result and the long-term result includes: Fuse the short-term cargo volume estimation data and the short-term abnormal risk prompt in the short-term result with the future second number of time period cargo volume growth trend data and the cargo class structure adjustment direction in the long-term result to generate a comprehensive analysis data set; Based on the comprehensive analysis data set, calculate the peak warehouse demand of the future first number of time periods and the average warehouse demand of the future second number of time periods by cargo class, adjust the storage planning data of each cargo class, and the storage planning data includes warehouse area allocation ratio and warehouse reservation quantity; Extract the logistics node flow change estimation data and the logistics node expansion potential data of each cargo class in the comprehensive analysis data set, and adjust the route data of each cargo class by combining the current route capacity configuration information, and the route data includes route frequency and route coverage logistics node range; Output the adjusted storage planning data and route data of each cargo class, wherein the first number of time periods is shorter than the second number of time periods, and the first number of time periods corresponds to the prediction period of the short-term result, and the second number of time periods corresponds to the prediction period of the long-term result.
Citation Information
Patent Citations
Method and system for automatically generating logistics based on order abnormity and storage medium
CN118246999A
Automatic financial information processing method based on AI
CN120494993A
Operation and control analysis method and device based on logistics transportation, equipment and medium
CN120598453A
Supply chain sales anomaly detection and root cause analysis system and method fused with knowledge graph
CN120631970A
System for predicting freight rates of optimized import and export cargo transfort routes and providng customized customer managing services based on artificial intelligence
KR102560210B1