A method and system for predicting throughput of a terminal container and a storage medium
By collecting and analyzing multi-dimensional operational data, a dynamic prediction model was constructed, which solved the problem of low accuracy in traditional throughput prediction and enabled the scientific allocation and efficient management of port resources.
Patent Information
- Application Number
- CN202511368621.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-09-24
AI Technical Summary
Traditional terminal throughput forecasting methods do not fully integrate multi-dimensional operational data and ignore dynamic external factors, resulting in low forecast accuracy, unreasonable resource allocation, and impact on operational efficiency and operating costs.
Collect multi-dimensional basic operation data, construct a standardized historical database, extract multi-dimensional correlation features by running correlation feature mining strategy, construct a dynamic prediction model by combining time series model and multiple regression model, and introduce influencing factor evaluation and dynamic correction mechanism to optimize the prediction model.
It improves the accuracy and stability of throughput forecasting, enables timely adaptation to changes in the external environment, achieves scientific resource optimization and operation scheduling, and enhances the reliability and practicality of port management.
Smart Images

Figure CN120851952B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of port resource scheduling technology, and more specifically, to a method, system, and storage medium for predicting port cargo throughput. Background Technology
[0002] In port operation and management, throughput forecasting is crucial for the rational allocation of resources and the improvement of operational efficiency. Accurate forecasting of future throughput provides a reliable basis for equipment scheduling, manpower scheduling, and yard resource planning, thereby reducing operating costs and improving the overall operational efficiency of the port. Traditional forecasting methods often rely on simple historical data statistics, failing to fully explore the correlations between data points, such as the combined impact of factors like ship type, cargo type, and operation time on throughput. Furthermore, existing methods often only focus on the time-series changes in historical throughput, ignoring the potential correlations between multi-dimensional data, leading to insufficient accuracy in forecast results. External factors such as fluctuations in market demand for different cargo types, dynamic changes in vessel scheduling, extreme weather, and policy adjustments also frequently cause significant deviations between forecasts and actual results. This results in unreasonable allocation of port equipment, manpower, and yard resources, affecting operational efficiency and costs, and making it difficult to adapt to complex and ever-changing port operation scenarios, leading to significant discrepancies between forecasts and actual conditions.
[0003] Insufficient forecasting accuracy often leads to unreasonable allocation of resources such as terminal equipment, manpower, and storage yards, resulting in idle or overloaded equipment, redundant or insufficient manpower scheduling, and congestion or vacancy in storage yards, thereby affecting operational efficiency and increasing operating costs.
[0004] Therefore, it is necessary to provide a method, system, and storage medium for predicting terminal cargo throughput to solve the above-mentioned technical problems. In order to solve the above problems, a technical solution is provided. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art, the present invention provides a terminal cargo throughput prediction method, system and storage medium to solve the problems of low prediction accuracy and unreasonable resource allocation caused by the failure of traditional terminal throughput prediction to fully correlate multi-dimensional operational data and ignore dynamic external factors.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A method for predicting terminal cargo throughput includes the following steps:
[0008] Collect multi-dimensional basic operation data from the terminal tallying operation system, preprocess the multi-dimensional operation data to obtain the first operation data, construct a standardized historical database, and classify and store the first operation data according to the dimensions.
[0009] Based on the first operational data, a strategy for mining operational correlation features was constructed to analyze the operational volume patterns of different ship types and corresponding cargo types. Multi-dimensional correlation features were extracted based on the operational volume patterns, a multi-dimensional feature set was constructed, and the influence weight of each dimension correlation feature on throughput data was quantified.
[0010] A dynamic prediction model is constructed by fusing time series models and multivariate regression models with multidimensional feature sets to predict the throughput data of terminal containers; the output results are corrected and judged based on the influencing factors of the dynamic prediction model.
[0011] The system outputs a forecast report on the throughput of container cargo at the terminal at a fixed period, collects actual throughput data in real time, compares and analyzes the actual throughput data with the forecasted throughput data of container cargo at the terminal to calculate the forecast error rate, and optimizes the dynamic forecast model based on the forecast error rate.
[0012] As a further aspect of the present invention, a strategy for mining operational correlation features is constructed based on the first operational data to analyze the operational volume patterns of different ship types and corresponding cargo types. The specific operational volume patterns are as follows:
[0013] Based on the analysis of the first operational data, the operational volume patterns of different ship types and corresponding cargo types are analyzed. The operational volume patterns include the first correlation pattern between cargo information data and operational information data, the second correlation pattern between cargo information data and ship information data, and the third correlation pattern between ship information data and operational information data.
[0014] As a further aspect of the present invention, the specific steps for obtaining the regularity characteristics of workload are as follows:
[0015] The time series of ship information data is used as the first sequence, the time series of cargo information data is used as the second sequence, and the time series of operation information data is used as the third sequence.
[0016] Based on the first sequence of statistics within a fixed time window, ship-related indicators are statistically analyzed; cargo-related indicators are statistically analyzed in the second sequence; and operation-related indicators are statistically analyzed in the third sequence.
[0017] Obtain the time series of ship-related indicators within a preset time period as the first correlation sequence I1 = {s1, ..., s2} t , ..., s T}, where s t Let t be the ship-related indicators at time t, and T be the preset time period length. The time series of cargo-related indicators is used as the second correlation sequence I2 = {c1, ..., c2}. t c T}, where c t Let I3 be the time series of cargo-related indicators and operation-related indicators at time t, which is used as the third correlation sequence I3 = {o1,…,o2}. t ,…,oT}, where o t For the relevant indicators of the operation at time t;
[0018] The first association rule feature F1(c,o) is obtained by performing association analysis on the second and third related sequences. The second association rule feature F2(s,c) is obtained by performing association analysis on the first and second related sequences. The third association rule feature F3(s,o) is obtained by performing association analysis on the first and third related sequences.
[0019] As a further aspect of the present invention, a dynamic prediction model is constructed by fusing a multi-dimensional feature set with a time series model and a multivariate regression model to predict the throughput data of terminal containers. The specific steps are as follows:
[0020] Historical monthly and annual cargo throughput data are obtained, and the ARIMA time series algorithm is used to capture the fluctuation trend of cargo throughput over time. Generate basic time series forecast features; these features include long-term trend characteristics, short-term fluctuation characteristics, and seasonality characteristics.
[0021] Using multidimensional correlation features as independent variables and cargo throughput as dependent variable, a multivariate regression sub-model is constructed to predict the first cargo throughput. The first cargo throughput is then fused and calibrated with the basic time series prediction features to output the second cargo throughput.
[0022] As a further aspect of the present invention, a multivariate regression sub-model is constructed to predict the first cargo throughput using multidimensional correlation features as independent variables and cargo throughput as the dependent variable. The first cargo throughput is then fused and calibrated with the basic time series prediction features to output the second cargo throughput. The specific steps are as follows:
[0023] Extract multi-dimensional correlation features, including the first correlation feature F1(c,o), the second correlation feature F2(s,c), and the third correlation feature F3(s,o), to obtain the cargo throughput time series {G1,…,G t ,…,G T}, G t Let G be the first cargo throughput at time t. T The cargo throughput within a preset time period T;
[0024] Based on multi-dimensional correlation features and cargo throughput time series, a multiple linear regression equation is constructed to output the first cargo throughput.
[0025] The basic time series forecast features are obtained, including long-term trend features, short-term fluctuation features, and seasonal features. The first cargo throughput is fused and calibrated with the basic time series forecast features to obtain the second cargo throughput.
[0026] As a further aspect of the present invention, the output results are corrected and judged according to the dynamic prediction model of influencing factors. The specific steps are as follows: obtain the historical second cargo throughput, construct a first-level correction model by combining the real cargo throughput of similar ships in history, and perform preliminary correction on the second cargo throughput to be corrected to output the first-level predicted throughput.
[0027] An influencing factor assessment model is established based on the influencing factors to output the degree of influence of the factors. Based on the degree of influence of the factors, it is determined whether a secondary correction is needed to the first-level predicted throughput.
[0028] As a further aspect of the present invention, the multi-dimensional basic operational data includes ship information data, cargo information data, and operational information data; ship information data includes ship name, voyage number, ship type, route, and designed cargo capacity; cargo information data includes cargo type, number of containers, weight, dangerous goods attributes, and storage requirements; operational information data includes operation start and end times, number of team members, and equipment usage records.
[0029] As a further aspect of the present invention, the actual throughput data is compared and analyzed with the predicted throughput data of the terminal containers to calculate the prediction error rate. The dynamic prediction model is then optimized based on the prediction error rate. Specifically, the formula for calculating the prediction error rate is as follows: Where γ is the prediction error rate, M is the actual throughput data, and M′ is the predicted throughput data of the terminal container. The prediction error rate is compared with the preset error threshold. If the prediction error rate is greater than or equal to the preset error threshold, the dynamic prediction model needs to be optimized; if the prediction error rate is less than the preset error threshold, the dynamic prediction model does not need to be optimized.
[0030] A terminal cargo throughput prediction system includes a multi-dimensional operation data acquisition module, a multi-dimensional correlation feature mining module, a cargo throughput prediction module, a correction and discrimination module, and an error optimization module;
[0031] The multi-dimensional operation data acquisition module is used to collect multi-dimensional basic operation data of the terminal tallying operation system, preprocess the multi-dimensional operation data to obtain the first operation data, build a standardized historical database, and classify and store the first operation data according to the dimensions.
[0032] The multi-dimensional correlation feature mining module is used to build and run correlation feature mining strategies based on the first operation data, analyze the operation volume patterns of different ship types and corresponding cargo types, extract multi-dimensional correlation features based on the operation volume patterns, construct a multi-dimensional feature set, and quantify the influence weight of each dimension correlation feature on throughput data.
[0033] The cargo throughput prediction module is used to construct a dynamic prediction model by fusing time series models and multivariate regression models with multidimensional feature sets to predict the throughput data of cargo containers at the terminal.
[0034] The correction and discrimination module is used to correct and discriminate the output results based on the dynamic prediction model of influencing factors;
[0035] The error optimization module is used to output a throughput forecast report for terminal containers at fixed intervals, collect actual throughput data in real time, compare and analyze the actual throughput data with the predicted throughput data of terminal containers to calculate the prediction error rate, and optimize the dynamic prediction model based on the prediction error rate.
[0036] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a method for predicting the throughput of a terminal cargo container.
[0037] The technical effects and advantages of this invention—a method, system, and storage medium for predicting terminal cargo throughput—are as follows: This invention collects multi-dimensional basic operational data on ships, cargo, operations, and the environment, preprocesses and standardizes the data, and constructs a historical database, ensuring data integrity and consistency and providing a reliable data foundation for subsequent analysis. Based on the initial operational data, an operational correlation feature mining strategy can deeply analyze the operational volume patterns of different ship types and corresponding cargo types, extract multi-dimensional correlation features, and quantify their impact on throughput. This allows for the scientific identification of key patterns hidden in historical operational data, providing a basis for optimal port resource allocation and operational scheduling. By integrating time series models with multiple regression models to construct a dynamic prediction model, it can not only capture long-term trends, seasonal fluctuations, and short-term fluctuations in throughput but also quantify the impact of various features on throughput, thereby improving prediction accuracy and stability. Furthermore, an influencing factor assessment and dynamic correction mechanism is introduced, which can correct the prediction results based on factors such as market, weather, regulatory adjustments, operations, and traffic, enabling the prediction to adapt to changes in the external environment in a timely manner and improving its practicality. By outputting forecast reports at fixed intervals and collecting actual throughput data in real time for comparative analysis, forecast errors are calculated and the model is optimized to achieve closed-loop optimization, enabling continuous improvement of the model, thereby enhancing the reliability of long-term forecasts and the scientific nature of port management.
[0038] This invention mines key features through historical data correlation analysis, making the model input more targeted and representative. It utilizes the advantages of dynamic prediction models that combine time series and multiple regression, and combines business scenarios and external factors for correction, which greatly improves prediction accuracy and effectively adapts to the complex and ever-changing operating environment of the terminal. Through prediction result output and dynamic optimization mechanism, it can continuously follow up on actual operation data, iterate and optimize the model, and make the prediction results always close to the actual throughput. Attached Figure Description
[0039] Figure 1 A flowchart of a terminal cargo container throughput prediction method provided in an embodiment of the present invention;
[0040] Figure 2 This is a system block diagram of a terminal cargo throughput prediction system provided in an embodiment of the present invention. Detailed Implementation
[0041] The technical solutions of this invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described technical solutions are only a part of this invention, and not all of it. All other technical solutions obtained by those skilled in the art based on the technical solutions of this invention without inventive effort are within the scope of protection of this invention.
[0042] like Figure 1 The diagram shown is a flowchart of a terminal cargo throughput prediction method provided by an embodiment of the present invention. Figure 1 The execution entity of the method shown can be a software and / or hardware device. The execution entity of this application can include, but is not limited to, at least one of the following: user equipment, network equipment, etc. User equipment can include, but is not limited to, computers, smartphones, personal digital assistants (PDAs), and the aforementioned electronic devices. Network equipment can include, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers. Cloud computing is a type of distributed computing, consisting of a super virtual computer composed of a group of loosely coupled computers. This embodiment does not limit this. Steps 1 to 5 are included, as follows:
[0043] Step 1: Collect multi-dimensional basic operation data from the terminal tallying system, preprocess the multi-dimensional operation data to obtain the first operation data, construct a standardized historical database, and classify and store the first operation data according to the dimensions.
[0044] In practical application, this invention involves a port collecting multi-dimensional basic operational data through its tallying system. The collected data undergoes preprocessing, including outlier removal, missing value imputation, unit standardization, and format standardization, resulting in complete and consistent initial operational data. Subsequently, the port categorizes and stores this data according to different dimensions; for example, vessel information, cargo information, and operational information are stored in separate data tables, with associated indexes established for subsequent multi-dimensional analysis and historical trend mining. In this way, the port can establish a standardized, structured historical database, providing a reliable data foundation for subsequent operational pattern analysis, correlation feature extraction, and throughput prediction. It also facilitates data querying and visualization analysis for management personnel, enabling refined management.
[0045] Step 2: Based on the first operation data, construct an operation association feature mining strategy, analyze the operation volume pattern characteristics of different ship types and corresponding cargo types, extract multi-dimensional association features based on the operation volume pattern characteristics, construct a multi-dimensional feature set, and quantify the influence weight of each dimension association feature on the throughput data.
[0046] In a practical application at a port, the management department, based on the compiled initial operational data, constructed an operational correlation feature mining strategy to analyze the operational volume patterns of different ship types corresponding to different cargo types. For example, by statistically analyzing the loading and unloading volumes of containers, bulk cargo, and dangerous goods for different ship types, such as Panamax, Handysize, and Suezmax, over the past year, it was found that Handysize vessels mainly handled bulk cargo, Suezmax vessels mainly handled containers, while Panamax vessels had a more balanced operational volume distribution. Based on these patterns, multi-dimensional correlation features were further extracted, including the correlation between cargo throughput and equipment utilization, the correlation between the number of ships arriving and cargo throughput, and the correlation between the number of ships arriving and equipment utilization. Through statistical analysis and regression methods, the port quantified the impact weights of each dimension of correlation features on overall throughput. For instance, it was found that the correlation feature between cargo throughput and equipment utilization had the highest weight, indicating that equipment scheduling has the most significant impact on operational efficiency at high throughput. The establishment of this multi-dimensional feature set not only helps port managers identify key influencing factors but also provides reliable data support for subsequent dynamic prediction and optimized scheduling.
[0047] Step 3: Construct a dynamic prediction model by fusing a time series model and a multivariate regression model with a multidimensional feature set to predict the throughput data of the terminal containers;
[0048] In a practical application at a large port, this invention utilizes a pre-constructed multi-dimensional feature set, taking historical ship, cargo, operational, and environmental data as input, and integrates time-series and multiple regression models to construct a dynamic prediction model. Specifically, the ARIMA model is first used to analyze historical monthly and annual cargo throughput, capturing long-term growth trends, short-term fluctuations, and seasonal patterns to generate basic time-series prediction features. Subsequently, using multi-dimensional correlation features, such as the correlation between cargo throughput and equipment utilization, and the correlation between the number of ships arriving and cargo throughput, as independent variables, a multiple regression sub-model is constructed to predict the first cargo throughput. The regression results are then fused and calibrated with the time-series prediction features to output a second predicted cargo throughput value, providing a scientific basis for port operation scheduling, equipment configuration, and resource optimization.
[0049] Step 4: Correct and judge the output results based on the dynamic prediction model of influencing factors;
[0050] In a practical application at a large coastal container port, the port management department comprehensively collected and organized data on ship operations, cargo throughput, and operational environment over the past five years, forming a multi-dimensional feature set. Based on these features, the management department used the ARIMA time series model to analyze historical monthly and annual cargo throughput, extracting long-term trends, short-term fluctuations, and seasonal patterns to generate basic time series prediction features. Simultaneously, using related features such as the number of ships arriving, ship type, cargo throughput, and equipment utilization rate as independent variables, a multivariate regression sub-model was constructed to predict the first cargo throughput. Subsequently, the regression model output was fused and calibrated with the time series prediction features to obtain the second predicted cargo throughput value. For example, in the actual prediction for January 2025, this method predicted the port's container throughput to be approximately 155,200, highly consistent with subsequent actual operations. This dynamic prediction model not only quantifies the operational needs of different ship types and cargo types in advance but also provides a scientific basis for port equipment scheduling, berth arrangement, and resource optimization, significantly improving operational efficiency and decision-making accuracy.
[0051] Step 5: Output the terminal cargo container throughput forecast report according to a fixed period, collect the actual throughput data in real time, compare and analyze the actual throughput data with the predicted terminal cargo container throughput data to calculate the prediction error rate, and optimize the dynamic prediction model based on the prediction error rate.
[0052] In a practical application at a coastal container port, this invention involves the port management department generating monthly forecasts of terminal container throughput. These reports include predicted monthly and quarterly throughput data, the distribution of workload for different cargo types and vessel types, and equipment utilization rates. Simultaneously, the port's tallying system collects real-time data on actual monthly throughput, including the number of containers, their weight, and the proportion of dangerous goods. By comparing the actual throughput with the forecasts, the management department calculates the forecast error rate. For example, if the actual throughput for a certain month is found to be 2% higher than the forecast, the deviation for some cargo types may reach 5%. Based on these error analysis results, the dynamic forecasting model adjusts the time series parameters and regression model coefficients, optimizing the weight allocation of correlated features to improve the forecast accuracy for the next period. This mechanism enables the port to continuously optimize the forecasting model under changing operational environments, achieving more precise operational scheduling and resource allocation, and improving overall operational efficiency.
[0053] It is worth noting that all the data mentioned above, as well as all data thereafter, have been dimensionless.
[0054] It should be noted that, at fixed intervals, such as weekly or monthly, terminal cargo throughput forecast reports are output. These reports include predicted throughput values and confidence intervals for different time dimensions, cargo types, and ship types. Simultaneously, actual throughput data during terminal operations is collected in real time and compared with the forecast results. The forecast error rate is calculated, and the parameters of the dynamic forecast model are dynamically adjusted based on error feedback. This includes retraining the model, optimizing feature weights, and continuously iterating to optimize the forecast model. This ensures that the forecasting method maintains high accuracy despite changes in terminal operation scenarios, providing precise data support for terminal resource allocation, such as loading and unloading equipment scheduling, manpower scheduling, yard planning, and operational process optimization.
[0055] The actual throughput data is compared and analyzed with the predicted throughput data of the terminal containers to calculate the prediction error rate. Based on the prediction error rate, the dynamic prediction model is optimized. Specifically, the formula for calculating the prediction error rate can be expressed as: Where γ is the prediction error rate, M is the actual throughput data, and M′ is the predicted throughput data of the terminal container. The prediction error rate is compared with the preset error threshold. If the prediction error rate is greater than or equal to the preset error threshold, the dynamic prediction model needs to be optimized; if the prediction error rate is less than the preset error threshold, the dynamic prediction model does not need to be optimized.
[0056] Preferably, the multi-dimensional basic operational data includes vessel information data, cargo information data, and operational information data; vessel information data includes vessel name, voyage number, vessel type, route, and designed cargo capacity; cargo information data includes cargo type, number of containers, weight, dangerous goods attributes, and storage requirements; operational information data includes operation start and end times, number of team members, and equipment usage records.
[0057] It should be noted that the first operational data is obtained by preprocessing the multi-dimensional operational data. The preprocessing includes removing duplicate and erroneous records through data cleaning, imputing missing values (such as filling based on the average value of the same voyage and the same type of cargo), correcting outliers, and making judgments and adjustments by combining industry standards and terminal operation experience. For example, when the cargo container volume of a single ship exceeds the physical limit, it is necessary to make corrections by comparing industry standards and historical data.
[0058] Preferably, a strategy for mining operational correlation features is constructed based on the first operational data to analyze the operational volume patterns of different ship types and corresponding cargo types. Multi-dimensional correlation features are extracted based on these operational volume patterns to construct a multi-dimensional feature set. The influence weight of each dimension of correlation features on throughput data is quantified. Specifically, the operational volume patterns are as follows:
[0059] Based on the analysis of the first operational data, the operational volume patterns of different ship types and corresponding cargo types are analyzed. The operational volume patterns include the first correlation pattern between cargo information data and operational information data, the second correlation pattern between cargo information data and ship information data, and the third correlation pattern between ship information data and operational information data.
[0060] In the analysis of the first operational data at a port, the research team of this invention explored the patterns of operational volume between different ship types and cargo types. Firstly, in the first correlation between cargo information data and operational information data, by analyzing the correlation between container quantity, cargo weight, and the proportion of dangerous goods with the start and end times of operations and equipment utilization, it was found that the average operational efficiency of dangerous goods significantly decreased during night shifts, while bulk cargo maintained high loading and unloading efficiency under high equipment utilization. Secondly, in the second correlation between cargo information data and ship information data, by matching the ship's designed cargo capacity, ship type, and cargo type and weight, it was concluded that container ships mainly handle standard container cargo, while bulk carriers are more suitable for transporting bulk cargoes such as coal and ore. Furthermore, fluctuations in ship cargo capacity showed a strong positive correlation with the contribution of different cargo types to throughput. Furthermore, in the third correlation pattern of ship information data and operation information data, by combining the time series of the number of ships arriving at the port, ship type, number of operation team personnel, and equipment utilization rate, it was found that when large bulk carriers arrive at the port in a concentrated manner, the port area equipment utilization rate increases significantly, while insufficient team personnel will lead to longer operation time, thereby affecting the overall throughput efficiency.
[0061] Preferably, the specific steps for obtaining the regularity characteristics of workload are as follows:
[0062] The time series of ship information data is used as the first sequence, the time series of cargo information data is used as the second sequence, and the time series of operation information data is used as the third sequence.
[0063] Based on the first sequence of statistics within a fixed time window, ship-related indicators are statistically analyzed; cargo-related indicators are statistically analyzed in the second sequence; and operation-related indicators are statistically analyzed in the third sequence.
[0064] Obtain the time series of ship-related indicators within a preset time period as the first correlation sequence I1 = {s1, ..., s2} t , ..., s T}, where s t Let t be the ship-related indicators at time t, and T be the preset time period length. The time series of cargo-related indicators is used as the second correlation sequence I2 = {c1, ..., c2}. t c T}, where c t Let I3 be the time series of cargo-related indicators and operation-related indicators at time t, which is used as the third correlation sequence I3 = {o1,…,o2}. t ,…,o T}, where o t For the relevant indicators of the operation at time t;
[0065] The first association rule feature F1(c,o) is obtained by performing association analysis on the second and third related sequences. The second association rule feature F2(s,c) is obtained by performing association analysis on the first and second related sequences. The third association rule feature F3(s,o) is obtained by performing association analysis on the first and third related sequences.
[0066] It should be noted that ship-related indicators include the number of ships arriving at port; cargo-related indicators include cargo throughput; and operation-related indicators include equipment utilization rate.
[0067] In the operational data analysis of a container terminal, researchers first acquired time series data on ship information, cargo information, and operations, constructing three basic sequences. For example, ship information data includes the daily number of arriving ships, cargo information data includes cargo throughput, and operations information data includes the utilization rate of loading and unloading equipment. Within a fixed time window (e.g., one month), researchers statistically processed the three sequences to obtain time series data on ship-related indicators, cargo-related indicators, and operations-related indicators. Using January 2023 to December 2024 as a preset time period, the time series of ship-related indicators within this preset time period was obtained as the first relevant sequence I1 = {s1, ..., s2}. t , ..., s T}, where s t Let t be the ship-related indicators at time t, and T be the preset time period length. The time series of cargo-related indicators is used as the second correlation sequence I2 = {c1, ..., c2}. t c T}, where c t Let I3 be the time series of cargo-related indicators and operation-related indicators at time t, which is used as the third correlation sequence I3 = {o1,…,o2}. t ,…,o T}, where o t These are the relevant indicators for the operation at time t.
[0068] Based on the above data, researchers further conducted multi-dimensional correlation analysis. Through correlation analysis of cargo-related indicator sequences and operation-related indicator sequences, the first correlation characteristic was obtained, reflecting the matching relationship between cargo throughput and equipment utilization. For example, if equipment utilization increases simultaneously in months with rapid growth in container throughput, it indicates a strong positive correlation between terminal operation capacity and cargo flow. Further analysis of ship-related indicator sequences and cargo-related indicator sequences yielded the second correlation characteristic, revealing the coupling characteristics between the number of arriving ships and cargo throughput. For instance, in the second half of 2023, the number of arriving ships increased significantly, and cargo throughput also increased simultaneously, indicating that changes in ship size had a significant impact on throughput during this period. Finally, analysis of ship-related indicator sequences and operation-related indicator sequences yielded the third correlation characteristic, measuring the synergistic relationship between the number of arriving ships and equipment utilization. When the number of arriving ships increases, but equipment utilization fails to improve significantly, it indicates an equipment bottleneck, requiring optimization of operation scheduling strategies.
[0069] Preferably, the first association pattern feature F1(c,o) is obtained by performing association analysis on the second and third related sequences. The calculation formula for the association analysis is as follows: In the formula: F1(c,o) is the first correlation feature, c t Let c be the cargo-related index at time t, c′ be the mean of the cargo-related index in the second correlation sequence, and o t Let t be the job-related index at time t, and o′ be the mean of the job-related index in the third correlation sequence.
[0070] The first correlation characteristic indicates a strong positive correlation between cargo throughput and equipment utilization. In other words, as port cargo throughput increases, equipment utilization also rises, reflecting that under high throughput conditions, port equipment resources are fully mobilized, resulting in a significant improvement in operational efficiency. Extracting this characteristic not only verifies the dynamic dependence between cargo and operations but also provides important quantitative evidence for subsequent port equipment scheduling and resource allocation optimization.
[0071] Association analysis was performed on the first and second related sequences to obtain the second association pattern feature F2(s,c). The calculation formula for the association analysis is as follows:
[0072] In the formula: F2(s,c) represents the second correlation feature, and c t Let be the cargo-related index at time t, c′ be the mean of the cargo-related index in the second correlation sequence, and s be the cargo-related index at time t. t Let be the ship-related indicators at time t, and s′ be the ship-related indicators in the first correlation sequence.
[0073] Association analysis was performed on the first and third related sequences to obtain the third association pattern feature F3(s,o). The calculation formula for the association analysis is as follows:
[0074] In the formula: F3(s,o) represents the third correlation feature, c t Let c be the ship-related index at time t, c′ be the mean of the ship-related index in the first correlation sequence, and o t Let t be the job-related index at time t, and o′ be the mean of the job-related index in the third correlation sequence.
[0075] In the operational analysis of a large container port, researchers extracted time series data for three types of indicators—first, second, and third correlation sequences—based on multi-dimensional data from January 2022 to December 2024. Through correlation analysis of cargo-related and operational-related indicators, the first correlation characteristic was calculated. The analysis results show that the first correlation characteristic value is close to 0.85, indicating a significant positive correlation between cargo throughput and equipment utilization. For example, in the third quarter of 2023, cargo throughput increased by 12% quarter-on-quarter, while equipment utilization increased by approximately 10% during the same period, reflecting that during peak cargo flow periods, equipment resources were fully mobilized, and operational efficiency improved simultaneously.
[0076] Furthermore, researchers derived a second correlation characteristic through correlation analysis between the number of ships arriving at the port and the throughput of different cargo types. The results show that the value of this second correlation characteristic is approximately 0.78, indicating a strong positive correlation between the number of ships arriving at the port and the throughput of different cargo types. Taking the first half of 2024 as an example, the average monthly number of ships arriving at the port increased by 15% compared to the same period of the previous year, while the corresponding throughput of different cargo types increased by 13%, demonstrating that port cargo flow is mainly driven by the scale of ship arrivals.
[0077] Finally, through the analysis of the number of ships arriving at the port and the equipment utilization rate, the third correlation characteristic was calculated, with a value of approximately 0.69, indicating a moderate positive correlation between the number of ships arriving at the port and the equipment utilization rate. At the end of 2022, affected by the cold wave, ships arrived at the port in a concentrated manner, with the number of ships surging by 20% in a short period of time. However, due to the lag in equipment scheduling, the equipment utilization rate only increased by 12%, indicating that there is a certain bottleneck in port equipment scheduling.
[0078] By systematically extracting and quantifying the three types of regular characteristics, port managers can not only grasp the dynamic dependence between cargo throughput and operational capacity, but also reveal the transmission mechanism of the scale of ship arrivals on throughput and operational efficiency, providing a scientific basis for optimizing equipment scheduling and improving the efficiency of port resource allocation.
[0079] Preferably, a dynamic prediction model is constructed by fusing a time series model and a multivariate regression model using a multidimensional feature set to predict the throughput data of container terminals. The specific steps are as follows:
[0080] Historical monthly and annual cargo throughput data are obtained, and the ARIMA time series algorithm is used to capture the fluctuation trend of cargo throughput over time. Generate basic time series forecast features; these features include long-term trend characteristics, short-term fluctuation characteristics, and seasonality characteristics.
[0081] Using multidimensional correlation features as independent variables and cargo throughput as dependent variable, a multivariate regression sub-model is constructed to predict the first cargo throughput. The first cargo throughput is then fused and calibrated with the basic time series prediction features to output the second cargo throughput.
[0082] In this embodiment of the invention, in the task of predicting the cargo throughput of a large container terminal, researchers first collected historical monthly and annual total throughput data for the terminal over the past five years. Using the ARIMA time series algorithm, these historical data were modeled to obtain a throughput prediction sequence for the next 12 months. Furthermore, long-term trend features reflecting throughput changes, short-term fluctuation features capturing short-term disturbances, and periodic seasonal features were extracted.
[0083] Meanwhile, based on the previously constructed multi-dimensional correlation features, researchers used the first, second, and third correlation sequences as independent variables and cargo throughput as the dependent variable to construct a multivariate regression sub-model, obtaining the first-stage cargo throughput prediction results. Subsequently, this result was fused and calibrated with the time series prediction features provided by ARIMA to form the final second cargo throughput prediction value.
[0084] In practical applications, when there is an abnormal increase in the number of ships arriving at the port in a certain month, relying solely on time series forecasts may underestimate the throughput. However, because the multivariate regression sub-model incorporates multi-dimensional features such as ships, cargo, and operations, the model can adjust the forecast values in a timely manner, making the final fused forecast results closer to reality and helping the terminal to make advance arrangements for manpower and equipment.
[0085] Preferably, historical monthly and annual cargo throughput data are obtained, and the ARIMA time series algorithm is used to capture the fluctuation trend of cargo throughput over time. The specific steps for generating basic time series prediction features are as follows:
[0086] Obtain historical monthly and annual cargo throughput data, and construct monthly sequences y1 = {y1, ..., y2} respectively. i , ..., y n} and the annual sequence y2={Y1,…,Y j , ..., Y m}, where y i Let Y be the cargo throughput in month i, n be the total number of months, and Y be the total cargo throughput in month i. j Let m be the cargo throughput in month j, and m be the total for the year.
[0087] The monthly and annual series are preprocessed. Based on the stationary series, the order parameters (p, d, q) of the ARIMA model are determined. The difference order d is determined based on the unit root test results. The order p of the autoregressive term is determined by observing the partial autocorrelation function, and the order q of the moving average term is determined by observing the autocorrelation function. If the monthly series exhibits significant seasonality, a seasonal ARIMA model is established: ARIMA(p, d, q)(P, D, Q). 12 Where 12 represents the 12-month seasonal cycle of a year, P is the seasonal autoregressive order, D is the seasonal difference order, and Q is the seasonal moving average order.
[0088] The maximum likelihood estimation method is used to estimate the parameters of the ARIMA model, and the model is fitted using the training dataset to obtain the predicted values. Calculate residual values Obtain the residual sequence, where y i Let the cargo throughput be in month i. The forecast value of cargo throughput for month i is given, and the residual series is tested for white noise. If the residual series fails the test, the parameters (p,d,q) or (p,D,Q) are adjusted until the residual series approximately satisfies the white noise assumption.
[0089] Based on the fitted ARIMA model, predictions are made for the next h time steps to obtain the fluctuation trend sequence. in, To represent the predicted throughput at time point n+h, where h is the prediction step size and H is the total number of time steps;
[0090] It should be noted that the existing monthly throughput data for the port from January 2020 to December 2024: Here, T = December 2024. If h = 1, then... This refers to the projected throughput for January 2025. If h = 6, then... This refers to the projected throughput for June 2025. If it's annual data, That is, the predicted throughput for the next year.
[0091] Extracting predictive features of cargo throughput using the trained ARIMA model: Long-term trend features are derived from the predicted cargo throughput values. The annual throughput is characterized by its overall growth or decline trend over time; short-term volatility is reflected by the variance of the residual series, indicating the intensity of short-term volatility; seasonality is represented by the seasonal difference result y. i -y i-12 Characterizing the periodic fluctuation pattern of different months, y i-12 This represents the cargo throughput for the i-th month of the previous year.
[0092] In this embodiment of the invention, a study on throughput forecasting at a container port collected monthly cargo throughput data from January 2020 to December 2024, constructing monthly and annual series. To establish a forecasting model, these two series were first stationary, and the difference order *d* was determined using a unit root test. The order *p* of the autoregressive term was determined using a partial autocorrelation function, and the order *q* of the moving average term was determined using an autocorrelation function. Because the port's monthly throughput data exhibits significant seasonality, such as peak periods during the Spring Festival, summer vacation, and year-end, the researchers established a seasonal ARIMA model: ARIMA(p,d,q)(P,D,Q). 12 The number 12 represents the 12-month seasonal cycle of a year.
[0093] After parameter selection, the maximum likelihood estimation method was used for parameter estimation, and the model was fitted using the training dataset from 2020–2023 to obtain the predicted values. The residual sequence is then calculated. The white noise hypothesis is tested to ensure that the residuals no longer exhibit significant autocorrelation. If the test fails, the parameters (p,d,q) and (P,D,Q) are adjusted until the residuals meet the white noise requirement. Finally, based on the fitted model, researchers predict the throughput from January 2025 to December 2025; for example, when h=1, the predicted value for January 2025 is obtained. When h=6, the predicted value for June 2025 is obtained.
[0094] When extracting predictive features, researchers defined long-term trend features as the direction of change in the annual throughput forecast, used to characterize the overall growth or decline trend of port throughput; short-term fluctuation features were defined as the variance of the residual sequence, used to measure the short-term fluctuation intensity of port throughput; and seasonal features were characterized by seasonal differencing results to reflect the periodicity of the same month in different years. For example, the model prediction results show that the port's throughput decreases in the first quarter of each year due to the Spring Festival holiday, while it increases in the third quarter due to the peak export season. This periodic fluctuation pattern is clearly reflected in the seasonal features. Using the ARIMA model, not only can the predicted throughput sequence for a future period be obtained, but it can also provide ports with feature information on long-term trends, short-term fluctuations, and seasonal periodic patterns, thereby assisting ports in making more scientific decisions on equipment scheduling and resource allocation.
[0095] Preferably, a multivariate regression sub-model is constructed using multidimensional correlation features as independent variables and cargo throughput as the dependent variable to predict the first cargo throughput. The first cargo throughput is then fused and calibrated with the basic time series prediction features to output the second cargo throughput. The specific steps are as follows: Extract multidimensional correlation features, including the first correlation regularity feature F1(c,o), the second correlation regularity feature F2(s,c), and the third correlation regularity feature F3(s,o), and obtain the cargo throughput time series {G1,…,G...}. t ,…,G T}, G t Let G be the first cargo throughput at time t. T The cargo throughput within a preset time period T;
[0096] Based on multi-dimensional correlation features and cargo throughput time series, a multiple linear regression equation is constructed to output the first cargo throughput. The formula for the multiple linear regression equation is as follows:
[0097] Where: G t Let be the first cargo throughput at time t, β0 be the intercept term, and β1 be the regression coefficient of the first correlation characteristic. Let β be the first correlation feature at time t, and β2 be the regression coefficient of the second correlation feature. β3 represents the regression coefficient of the second correlation feature at time t, and β3 represents the regression coefficient of the third correlation feature. The third correlation feature at time t, ε t The error term at time t;
[0098] The basic time series forecast features are obtained, including long-term trend features, short-term fluctuation features, and seasonality features. The first cargo throughput is then fused and calibrated with the basic time series forecast features. The formula for fusion calibration is as follows:
[0099] In the formula: y t Let ' be the second cargo throughput at time t, s1 be the long-term trend characteristic, s2 be the short-term fluctuation characteristic, and s3 be the seasonal characteristic. G t Let t be the first cargo throughput at time t.
[0100] It should be noted that during the training of the dynamic prediction model, the training set and the test set can be divided in a 7:3 ratio. The model can be trained iteratively using historical data. By adjusting the model parameters, such as the order of the time series model and the coefficients of the regression model, the prediction error can be minimized and the model prediction performance can be optimized.
[0101] In the throughput prediction practice of a port, researchers first extracted multi-dimensional correlation features, including a first correlation feature between cargo and operations, a second correlation feature between ships and cargo, and a third correlation feature between ships and operations. Simultaneously, they obtained the time series of cargo throughput from January 2020 to December 2024. Based on this data, researchers constructed a multiple linear regression sub-model to quantify the impact of the multi-dimensional correlation features on throughput.
[0102] By fitting historical data, the model derives the regression coefficients of each correlation feature on cargo throughput. For example, β1 = 0.42 indicates that the matching degree between cargo and operations has a significant impact on throughput; β2 = 0.35 shows that the number of ships arriving at port also has a significant positive impact on cargo throughput; and β3 = 0.21 reflects a moderate impact of ship and equipment operational efficiency. During model training, historical data is divided into training and test sets in a 7:3 ratio. The regression coefficients are continuously adjusted through iterative training to minimize prediction error.
[0103] After obtaining the first cargo throughput forecast, the researchers further combined the basic forecast features extracted from the ARIMA time series model, including long-term trend features, short-term fluctuation features, and seasonal features, to perform fusion calibration and generate a second cargo throughput forecast.
[0104] Taking February 2025 as an example, the ARIMA model captured a long-term trend of cargo throughput of 1.08, short-term fluctuations of 0.05, and seasonality of 0.02. The first cargo throughput forecast was 152,000, and after fusion calibration, a second cargo throughput forecast of 136,574 was obtained, more accurately reflecting the combined impact of historical trends, short-term fluctuations, and seasonal changes. The dynamic forecasting model can not only quantify the impact of multi-dimensional factors on throughput, but also combine time series trends for calibration and optimization forecasting, providing port management departments with a scientific basis for decision-making in equipment scheduling, resource allocation, and annual planning.
[0105] Preferably, the output results are corrected and judged based on the dynamic prediction model of influencing factors. The specific steps are as follows:
[0106] Obtain the second historical cargo throughput, and construct a first-level correction model by combining it with the actual cargo throughput of similar ships in history. Perform preliminary correction on the second cargo throughput to be corrected and output the first-level predicted throughput.
[0107] An influencing factor assessment model is established based on the influencing factors to output the degree of influence of the factors. Based on the degree of influence of the factors, it is determined whether a secondary correction is needed to the first-level predicted throughput.
[0108] In forecasting the throughput of a large container port, researchers, after obtaining the second cargo throughput from a dynamic forecasting model, employed a correction and discrimination mechanism to calibrate the output to improve forecast accuracy. First, they acquired historical second cargo throughput data from 2020 to 2024, and simultaneously collected actual cargo throughput data from similar vessels during the same period to construct a primary correction model. This model calculates an error correction coefficient by comparing the dynamic forecast value with the actual historical value, thus providing an initial correction to the second cargo throughput to be corrected. Subsequently, researchers constructed an influencing factor assessment model based on factors such as port equipment utilization efficiency, weather conditions, tidal conditions, and port congestion, quantifying the impact of each factor on the throughput forecast results. For example, the assessment showed that an extreme cold wave in January 2025 might cause a 3% decrease in equipment scheduling efficiency, while port congestion would be 2% higher than the average level for the same period. Based on the assessment results, researchers determined that the primary forecast value was biased by environmental factors, and therefore performed a secondary correction on the primary forecast value. By combining primary correction with secondary factor calibration, port management departments can refine cargo throughput forecasts based on dynamic predictions, taking into account historical patterns and current environmental conditions. This provides more reliable decision support for equipment scheduling, resource allocation, and emergency management.
[0109] Preferably, the historical second-highest cargo throughput is obtained, and a first-level correction model is constructed by combining it with the actual cargo throughput of similar historical vessels. This first-level correction model is then used to initially correct the second-highest cargo throughput to be corrected, outputting the first-level predicted throughput. The formula for the first-level correction model is as follows: In the formula: y P1 ′ represents the first-level predicted throughput, y P Let ' be the second cargo throughput to be corrected at time P, and y be the second cargo throughput to be corrected at time P. p ' represents the second-highest historical cargo throughput at time p, and y represents the second-highest throughput at time p. p Let p be the historical cargo throughput.
[0110] It should be noted that, This represents the total historical second-highest cargo throughput before time P, i.e., from time P-1 onwards. This represents the total historical cargo throughput up to time P, i.e., from time P-1.
[0111] In the cargo throughput forecasting practice of a container port, the dynamic forecasting model obtained a predicted value of 155,200 for the second cargo throughput in January 2025. To improve the forecast accuracy, the researchers introduced a first-level correction model, combining the historical second cargo throughput from January 2020 to December 2024 with the historical actual throughput, to calculate the average absolute value of the historical deviation, which was used to initially correct the second cargo throughput to be corrected.
[0112] Where T = 60 represents the number of months from January 2020 to December 2024, and the total number of months is... This is the second-highest cumulative cargo throughput in history. This represents the cumulative historical actual cargo throughput. Assuming the total historical second-largest cargo throughput over these five years is 9,000,000 and the total historical actual throughput is 8,900,000, then the average deviation is (9,000,000 - 8,900,000) / 60 = 1667.
[0113] Adding this deviation to the forecast value to be corrected in January 2025, we get a first-level forecast throughput of 156,867. Through this process, the first-level correction model effectively uses the differences in historical throughput of similar ships and cargo to make preliminary corrections to the dynamic forecast results, making the forecast value closer to the actual level, and providing a reliable basis for subsequent second-level corrections and comprehensive forecasts based on influencing factors.
[0114] Preferably, an influencing factor assessment model is established based on the influencing factors, outputting the degree of influence of each factor. The degree of influence is then used to determine whether a secondary correction to the first-level predicted throughput is needed. The formula for the factor assessment model is: f0 = ω1b1 + ... + ω Q b Q In the formula: f0 represents the degree of influence of the factor, ω1 represents the weight of the first influencing factor, b1 represents the weight of the first influencing factor, and ω Q Let b be the weight of the Qth influencing factor. Q Let Q be the Qth influencing factor, where Q is the number of influencing factors;
[0115] The influence of a factor is compared with a preset threshold. If the influence of a factor is greater than or equal to the preset threshold, a second correction is triggered; if the influence of a factor is less than the preset threshold, a second correction is not triggered.
[0116] When forecasting the cargo throughput of a large container port in January 2025, researchers, after obtaining a primary forecast throughput of 156,867, introduced an influencing factor assessment model to perform a secondary correction and judgment on the forecast result. The model calculates the overall impact of various factors by quantifying their degree. If the port faces a cold wave weather warning in January 2025, the meteorological impact factor b2 = 0.8, the market demand index is high (b1 = 0.6), the impact of the terminal equipment maintenance plan is limited (b4 = 0.2), and inland transportation congestion is slight (b5 = 0.1). Assuming the weights of each factor are ω1 = 0.3, ω2 = 0.25, ω4 = 0.2, and ω5 = 0.15, the overall impact is calculated as: f0 = 0.3 × 0.6 + 0.25 × 0.8 + 0.2 × 0.2 + 0.15 × 0.1 = 0.435. The preset threshold is 0.4. Since f0 = 0.435 ≥ 0.4, a secondary correction is triggered. Subsequently, based on the assessment results, the first-level predicted throughput of 156,867 was calibrated a second time to obtain the final predicted value after taking into account environmental and operational impacts.
[0117] Through this process, port management departments can scientifically judge and correct forecast results based on dynamic forecasts and by combining multiple factors such as market, weather, regulations, operations, and traffic, thereby providing a more reliable basis for decision-making in equipment scheduling, berth allocation, and logistics coordination.
[0118] It should be noted that the influencing factors include market factors, meteorological factors, regulatory adjustment factors, operational factors, and traffic factors.
[0119] Market influencing factors include demand index for specific cargo types, changes in import and export trade volume, freight rate index, and order volume in the industrial chain; meteorological influencing factors include weather warning levels, rainfall, wind speed, visibility, and tidal anomalies; regulatory adjustment influencing factors include changes in customs clearance regulations, quarantine / detention, port flow restrictions, and restrictions on nighttime operations; operational influencing factors include dock strikes, equipment maintenance plans, and temporary closure of berths; and transportation influencing factors include delays in upstream cargo supply, inland transportation congestion, and route adjustments.
[0120] A terminal cargo throughput prediction system includes a multi-dimensional operation data acquisition module, a multi-dimensional correlation feature mining module, a cargo throughput prediction module, a correction and discrimination module, and an error optimization module. The multi-dimensional operation data acquisition module is connected to the multi-dimensional correlation feature mining module; the multi-dimensional correlation feature mining module is connected to the cargo throughput prediction module; the cargo throughput prediction module is connected to the correction and discrimination module; and the correction and discrimination module is connected to the error optimization module.
[0121] The multi-dimensional operation data acquisition module is used to collect multi-dimensional basic operation data of the terminal tallying operation system, preprocess the multi-dimensional operation data to obtain the first operation data, build a standardized historical database, and classify and store the first operation data according to the dimensions.
[0122] The multi-dimensional correlation feature mining module is used to build and run correlation feature mining strategies based on the first operation data, analyze the operation volume patterns of different ship types and corresponding cargo types, extract multi-dimensional correlation features based on the operation volume patterns, construct a multi-dimensional feature set, and quantify the influence weight of each dimension correlation feature on throughput data.
[0123] The cargo throughput prediction module is used to construct a dynamic prediction model by fusing time series models and multivariate regression models with multidimensional feature sets to predict the throughput data of cargo containers at the terminal.
[0124] The correction and discrimination module is used to correct and discriminate the output results based on the dynamic prediction model of influencing factors;
[0125] The error optimization module is used to output a throughput forecast report for terminal containers at fixed intervals, collect actual throughput data in real time, compare and analyze the actual throughput data with the predicted throughput data of terminal containers to calculate the prediction error rate, and optimize the dynamic prediction model based on the prediction error rate.
[0126] like Figure 2 The diagram shown is a system block diagram of a terminal cargo container throughput prediction system according to an embodiment of the present invention, which can be used to execute... Figure 1 The steps in the method embodiments shown are implemented in a similar manner and have similar technical effects, and will not be repeated here.
[0127] A readable storage medium storing a computer program, which, when executed by a processor, is used to implement the steps of a terminal cargo container throughput prediction method as described in any of the above claims.
[0128] The readable storage medium can be a computer storage medium or a communication medium. A communication medium includes any medium that facilitates the transfer of computer programs from one location to another. A computer storage medium can be any available medium accessible to a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application-Specific Integrated Circuit (ASIC). Alternatively, the ASIC can be located in a user device. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0129] The present invention also provides a program product including executable instructions stored in a readable storage medium. At least one processor of the device can read the executable instructions from the readable storage medium, and the at least one processor executes the executable instructions to cause the device to implement the methods provided in the various embodiments described above.
[0130] In the embodiments of the above-described device, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules within the processor.
[0131] Through the above embodiments, this invention collects multi-dimensional basic operational data on ships, cargo, operations, and the environment, preprocesses and standardizes the data to construct a historical database, ensuring data integrity and consistency and providing a reliable data foundation for subsequent analysis. The operational correlation feature mining strategy built based on the first operational data can deeply analyze the operational volume patterns of different ship types and corresponding cargo types, extract multi-dimensional correlation features, and quantify their impact on throughput. This allows key patterns hidden in historical operational data to be scientifically identified, providing a basis for optimal port resource allocation and operational scheduling. By integrating time series models with multiple regression models to construct a dynamic prediction model, it can not only capture long-term trends, seasonal fluctuations, and short-term fluctuations in throughput but also quantify the impact of various features on throughput, thereby improving prediction accuracy and stability. Furthermore, an influencing factor assessment and dynamic correction mechanism is introduced, which can correct the prediction results based on factors such as market, weather, regulatory adjustments, operations, and traffic, enabling the prediction to adapt to changes in the external environment in a timely manner and improving its practicality. By outputting forecast reports at fixed intervals and collecting actual throughput data in real time for comparative analysis, forecast errors are calculated and the model is optimized to achieve closed-loop optimization, enabling continuous improvement of the model, thereby enhancing the reliability of long-term forecasts and the scientific nature of port management.
[0132] This invention mines key features through historical data correlation analysis, making the model input more targeted and representative. It utilizes the advantages of dynamic prediction models that combine time series and multiple regression, and combines business scenarios and external factors for correction, which greatly improves prediction accuracy and effectively adapts to the complex and ever-changing operating environment of the terminal. Through prediction result output and dynamic optimization mechanism, it can continuously follow up on actual operation data, iterate and optimize the model, and make the prediction results always close to the actual throughput.
[0133] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
[0134] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method of predicting throughput of a terminal container, characterized by, The method comprises the following steps: Collecting multi-dimensional basic operation data of the port tally operation system, pre-processing the multi-dimensional operation data to obtain first operation data, constructing a standardized historical database, and classifying and storing the first operation data according to dimensions; Based on the first operation data, an operation correlation feature mining strategy is constructed, the operation amount law characteristics of different ship types corresponding to cargo types are analyzed, multi-dimensional correlation features are extracted based on the operation amount law characteristics, a multi-dimensional feature set is constructed, and the influence weight of each dimension correlation feature on the throughput data is quantified; The operation amount law characteristics are as follows: Based on the first operation data, the operation amount law characteristics of different ship types corresponding to cargo types are analyzed, and the operation amount law characteristics include a first correlation law characteristic of cargo information data and operation information data, a second correlation law characteristic of cargo information data and ship information data, and a third correlation law characteristic of ship information data and operation information data; The specific steps of obtaining the operation amount law characteristics are as follows: Respectively, the time sequence of the ship information data is obtained as a first sequence, the time sequence of the cargo information data is obtained as a second sequence, and the time sequence of the operation information data is obtained as a third sequence; Based on the first sequence in the fixed time window, ship-related indexes are counted, cargo-related indexes are counted based on the second sequence, and operation-related indexes are counted based on the third sequence; Obtain the time series of the ship-related indicators in a preset time period as a first related sequence I1={s1, …, s t ,…,s T}, wherein s t is a ship-related indicator at time t, T is the length of the preset time period, the time series of the cargo-related indicators as a second related sequence I2={c1, …, c t ,…,c T}, wherein c t is a cargo-related indicator at time t, and the time series of the operation-related indicators as a third related sequence I3={o1, …, o t ,…,o T}, wherein o t is an operation-related indicator at time t. Through correlation analysis on the second and third related sequences, the first correlation law characteristic F1(c,o) is obtained, correlation analysis is performed on the first and second related sequences to obtain the second correlation law characteristic F2(s,c), and correlation analysis is performed on the first and third related sequences to obtain the third correlation law characteristic F3(s,o); A dynamic prediction model is constructed by fusing a time sequence model and a multiple regression model through the multi-dimensional feature set, and the throughput data of the port container is predicted, and the specific steps are as follows: Obtain historical monthly cargo throughput and annual cargo throughput, and capture the fluctuation trend sequence of cargo throughput over time using an ARIMA time series algorithm Generate a base time series prediction feature; the base time series prediction feature includes a long-term trend feature, a short-term fluctuation feature, and a seasonal feature; Taking the multi-dimensional correlation features as independent variables and the cargo throughput as dependent variables, a multiple regression sub-model is constructed to predict the first cargo throughput, and the first cargo throughput is fused and calibrated with the basic time sequence prediction features to output the second cargo throughput; The output result is modified and discriminated according to the influence factor dynamic prediction model; A throughput prediction report of the port container is output according to a fixed period, actual throughput data is collected in real time, the actual throughput data is compared and analyzed with the predicted throughput data of the port container to calculate a prediction error rate, and the dynamic prediction model is optimized according to the prediction error rate.
2. The method of claim 1, wherein, Taking the multi-dimensional correlation features as independent variables and the cargo throughput as dependent variables, a multiple regression sub-model is constructed to predict the first cargo throughput, and the first cargo throughput is fused and calibrated with the basic time sequence prediction features to output the second cargo throughput, and the specific steps are as follows: Extracting multi-dimensional correlation features, including a first correlation law feature F1(c, o), a second correlation law feature F2(s, c), and a third correlation law feature F3(s, o), to obtain a cargo throughput time sequence {G1, …, G t ,…,G T}, G t is the first cargo throughput at time t, and G T is the cargo throughput in a preset time period T; A multiple linear regression equation is constructed based on the multi-dimensional correlation features and the cargo throughput time sequence to output the first cargo throughput; The basic time sequence prediction features including long-term trend features, short-term fluctuation features and seasonal features are obtained, and the first cargo throughput is fused and calibrated with the basic time sequence prediction features to obtain the second cargo throughput.
3. The method of claim 1, wherein, The output result is modified and discriminated according to the influence factor dynamic prediction model, and the specific steps are as follows: The historical second cargo throughput is acquired, a first correction model is constructed in combination with the real cargo throughput of the historical same type of ship, the second cargo throughput to be corrected is preliminarily corrected to output a first predicted throughput; An influence factor evaluation model is established based on the influence factors to output the influence degree of the factors, and whether the first predicted throughput needs to be secondarily corrected is judged according to the influence degree of the factors.
4. The method of claim 1, wherein, The multi-dimensional basic operation data includes ship information data, cargo information data and operation information data; the ship information data includes ship name, voyage, ship type, belonging route and designed cargo carrying capacity; the cargo information data includes cargo type, container quantity, weight, dangerous property and storage requirement; and the operation information data includes operation start and end time, crew quantity and equipment use record.
5. The method of claim 1, wherein, The prediction error rate is calculated by comparing the actual throughput data with the predicted throughput data of the wharf container, and the dynamic prediction model is optimized according to the prediction error rate, specifically: the calculation formula of the prediction error rate is represented as: Wherein, γ is the prediction error rate, M is the actual throughput data, and M' is the predicted throughput data of the wharf container; the prediction error rate is compared with the preset error threshold value, if the prediction error rate is greater than or equal to the preset error threshold value, the dynamic prediction model needs to be optimized; if the prediction error rate is less than the preset error threshold value, the dynamic prediction model does not need to be optimized.
6. A terminal container throughput prediction system for use in a terminal container throughput prediction method as claimed in any one of the claims 1-5, characterized in that, The system comprises a multi-dimensional operation data acquisition module, a multi-dimensional correlation feature mining module, a container throughput prediction module, a correction discrimination module and an error optimization module. The multi-dimensional operation data acquisition module is used for acquiring the multi-dimensional basic operation data of the terminal cargo handling operation system, pre-processing the multi-dimensional operation data to obtain first operation data, constructing a standardized historical database, and classifying and storing the first operation data according to dimensions; The multi-dimensional correlation feature mining module is used for constructing a running correlation feature mining strategy based on the first operation data, analyzing the operation amount law characteristics of different ship types corresponding to cargo types, extracting multi-dimensional correlation features based on the operation amount law characteristics, constructing a multi-dimensional feature set, and quantifying the influence weight of each dimensional correlation feature on the throughput data; The container throughput prediction module is used for constructing a dynamic prediction model by fusing a time series model and a multiple regression model through the multi-dimensional feature set, and predicting the throughput data of the terminal container; The correction discrimination module is used for correcting and discriminating the output result according to the influence dynamic prediction model; The error optimization module is used for outputting a terminal container throughput prediction report at a fixed period, collecting actual throughput data in real time, comparing and analyzing the actual throughput data with the predicted throughput data of the terminal container, calculating a prediction error rate, and optimizing the dynamic prediction model according to the prediction error rate.
7. A readable storage medium, in which a computer program is stored, characterized in that, The computer program is executed by a processor to implement the steps of the terminal container throughput prediction method, system and storage medium of any one of claims 1-5.
Citation Information
Patent Citations
Method and device for predicting throughput of container terminal, terminal equipment and medium
CN115049116A
Port container throughput estimation method and system
CN116739444A