Multi-cargo ship cargo capacity prediction method based on multi-source heterogeneous data and machine learning model

By integrating multi-source heterogeneous data and machine learning models, and combining ship segment and port entry/exit data, prediction models for different types of cargo are constructed. This solves the problems of data reliability and accuracy in ship cargo volume prediction methods and achieves efficient cargo volume prediction.

CN122020375APending Publication Date: 2026-05-12COSCO SHIPPING TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
COSCO SHIPPING TECH CO LTD
Filing Date
2026-01-22
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing methods for estimating ship cargo capacity have poor data reliability and low prediction accuracy. Traditional methods are complex to operate and cannot take into account the impact of different ship sizes.

Method used

By employing multi-source heterogeneous data and machine learning models, and integrating ship segment data and port entry/exit data, prediction models corresponding to different cargo types are constructed through data correction, anomaly detection, and missing value imputation. The K-nearest neighbor algorithm is used to predict cargo load.

Benefits of technology

This improves the reliability and accuracy of ship cargo volume forecasting data, provides more accurate cargo volume forecasting results, and provides data support for port throughput estimation and impact assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020375A_ABST
    Figure CN122020375A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of ship cargo capacity prediction, and particularly relates to a multi-cargo-type ship cargo capacity prediction method based on multi-source heterogeneous data and a machine learning model, and the method comprises the steps: fusing ship leg data and port entering and leaving data, matching the types of loaded cargoes in the ship port entering and leaving data into the ship leg data, and carrying out the matching of the types of loaded cargoes in the ship port entering and leaving data; different data sets are divided according to different types of loaded goods; the draught and cargo capacity of the data in the data set is corrected; removing abnormal data; supplementing missing values of the data in the data set; preprocessing data in the data set corresponding to the types of the loaded goods to obtain input data corresponding to different types of the loaded goods; and constructing a plurality of ship cargo capacity prediction models corresponding to different loaded cargo types, and performing evaluation to obtain an optimal ship cargo capacity prediction model corresponding to different loaded cargo types. According to the method provided by the invention, the problems of poor data reliability and low prediction result accuracy of an existing ship cargo capacity prediction method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of ship cargo volume prediction technology, specifically relating to a method for predicting the cargo volume of multi-cargo ships based on multi-source heterogeneous data and machine learning models. Background Technology

[0002] In the shipping industry, ship load information directly affects the estimation of port throughput and the assessment of port influence. However, due to issues such as lack of transparency and limited data sources within the industry, it is difficult to directly infer and obtain ship load data from industry data. Furthermore, traditional methods for estimating load are inaccurate and complex to operate.

[0003] Current traditional methods for estimating cargo load mainly focus on the following aspects:

[0004] (1) Equipment measurement method: This method mainly includes the draft survey method and the instrument measurement method. The draft survey method mainly relies on the crew or managers to read the six drafts of the ship and make an estimate in conjunction with the "Ship Load and Draft Comparison Table". The instrument measurement method mainly uses equipment such as ultrasonic, pressure and laser sensors to detect the current draft of the ship, and then uses computer analysis to obtain the ship's load.

[0005] (2) Empirical formula method: This method is mainly derived and calculated based on the relevant empirical formula in the Method to estimate cargo tonnemiles on page 361 of the 2020 IMO report. The ship's load capacity is calculated by using relevant parameters of the ship, including design draft, instantaneous draft, and design speed.

[0006] The shortcomings of these existing methods are that the equipment measurement method is complicated to operate and has poor real-time measurement capability, while the empirical formula method is not accurate in estimating the ship's load because it is difficult to take into account the impact of different ship sizes. Summary of the Invention

[0007] This invention addresses the problems of poor data reliability and low prediction accuracy in existing ship cargo volume prediction methods by providing a multi-cargo ship cargo volume prediction method based on multi-source heterogeneous data and machine learning models.

[0008] The technical solution claimed by this invention is as follows:

[0009] 1. A method for predicting cargo load of multi-cargo vessels based on multi-source heterogeneous data and machine learning models, characterized by comprising the following steps:

[0010] S1: Integrate ship segment data and port entry / exit data, and match the cargo types loaded in the port entry / exit data to the ship segment data; find ships loaded with crude oil through the tank volume field, and add or update the cargo type as crude oil in the ship segment data.

[0011] S2: The dataset is divided according to the types of cargo loaded in the ship segment data in S1, with different types of cargo corresponding to different datasets;

[0012] S3: Correct the draft and cargo load of the data in the dataset obtained in S2, including: using nearby draft data to correct data where the water output did not change before and after berthing operations, correcting data with excessive empty draft, and correcting data where draft and cargo load are not positively correlated.

[0013] S4: The data in the dataset obtained after processing S3 is checked for a positive correlation between ship draft and cargo capacity based on ship structure theory. If not, it is judged as abnormal data and removed from the dataset.

[0014] S5: For the dataset obtained from S4, use the random forest model to fill in the missing values ​​in the dataset;

[0015] S6: Preprocess the data in the dataset corresponding to the type of loaded goods to obtain the input data corresponding to different types of loaded goods;

[0016] S7: Based on different models, construct multiple ship cargo volume prediction models corresponding to different types of loaded cargo, train the prediction models using input data corresponding to the types of loaded cargo, and evaluate the model output data to obtain the optimal ship cargo volume prediction model corresponding to different types of loaded cargo.

[0017] 2. The method for predicting cargo volume of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 1, characterized in that: the vessel segment data is segmented and extracted from AIS data; in S1, for all vessel data in the vessel segment data where the liquid tank volume field is not empty, the cargo type is classified as crude oil; in S2, the data in the dataset is vessel segment data with a cargo type field.

[0018] 3. The method for predicting the cargo load of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 2, characterized in that, in S3, adjacent draft data is used to correct data whose outflow has not changed before and after berthing operations, specifically: extract all vessel segment data with berthing status from the dataset obtained in S2, each data point is taken as the start of a voyage, and the data is traversed sequentially. For vessel segment data where the loading and unloading status has changed, if the starting draft and ending draft have not changed, the data of the next adjacent vessel segment is traversed line by line until data with a change in starting draft is found and replaced.

[0019] 4. The method for predicting cargo load of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 2, characterized in that, in S3, the correction of data with excessive empty draft is specifically as follows: For vessel segment data where the cargo load is 0 but the draft data is too large and deviates from the draft corresponding to the empty vessel, the solution includes the following steps:

[0020] Step 1: Query all ship segment data with a cargo load of 0 in the dataset obtained from S2;

[0021] Step 2: Calculate the absolute value of the actual draft and the minimum midship draft in the ship segment data, and take 5% of the actual draft as the allowable draft deviation based on empirical observations;

[0022] Step 3: Compare the absolute value with the draft deviation. If the absolute value is greater than the draft deviation, it means that the actual draft in the empty data exceeds the maximum allowable empty draft of the corresponding vessel. Generate a random number from the interval [minimum midship draft, minimum midship draft + draft deviation] to replace the actual draft data in the vessel segment data. If the absolute value is less than or equal to the draft deviation, it means that the actual draft in the empty data is within the maximum allowable empty draft of the corresponding vessel. Then retain the draft data in the vessel segment data.

[0023] 5. The method for predicting the cargo load of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 2, characterized in that the formula for calculating the minimum midship draft is:

[0024] (1)

[0025] in: Indicates the minimum midship draft, Indicates the length of the ship.

[0026] 6. The method for predicting cargo load of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 5, characterized in that, in S3, the correction of data where draft and cargo load are not positively correlated is specifically as follows: based on the actual draft, if theoretically the vessel is empty, but the actual data does not, then the actual draft is corrected to 0; if theoretically the vessel is not empty, but in the actual data, there is no positive correlation between the actual data and the draft, then the following processing is performed:

[0027] Step 1: Calculate the absolute value of the actual draft and the minimum midship draft in the ship segment data of all datasets. Take 5% of the actual draft as the allowable draft deviation.

[0028] Step 2: Compare the absolute value with the draft deviation. If the absolute value is less than or equal to the draft deviation, it means the actual draft is close to the minimum midship draft, so the cargo load in the ship segment data is corrected to 0. If the absolute value is greater than the draft deviation, it means the actual draft is not close to the minimum midship draft. In this case, the ship is not empty and further processing is required. The specific processing procedure is as follows:

[0029] Calculate the expected cargo capacity based on formula (2):

[0030] (2)

[0031] Step 3: Calculate the percentage deviation between the expected cargo volume and the actual cargo volume, using the following formula:

[0032] (3)

[0033] Step 4: If the deviation percentage is greater than 5%, replace the cargo volume in the ship segment data with the expected cargo volume; otherwise, keep the cargo volume data in the ship segment data unchanged.

[0034] 7. The method for predicting the cargo load of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 6, characterized in that the method for judging abnormal data in S4 is as follows: with draft as the x-axis and cargo load as the y-axis, a schematic diagram of the relationship between the ship's cargo load and draft is drawn. Point A represents the ship's design draft and maximum cargo load, points B and C represent 90% and 110% of the design draft, respectively, and the y-axis represents the ship's maximum cargo load. Point D represents the minimum empty draft according to the ship's safety design requirements, and point E represents the upper limit of the empty draft. Data within the range of quadrilateral BCDE is considered normal data, while data outside this range is considered abnormal data.

[0035] 8. The method for predicting cargo volume of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 7, characterized in that the specific method for handling missing values ​​in S5 is as follows: In the dataset obtained by processing in S4, the field with missing values ​​is the change in deadweight tons per centimeter of the vessel's draft. Models such as RF, GBDT, or XGBoost, which are used to view the importance of features, are used. The maximum deadweight tons, length, width, height, and actual draft are used as inputs, and the output of the model with the highest R² and the lowest MAE is selected as the filling result.

[0036] 9. The method for predicting cargo volume of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 8, wherein the preprocessing in S6 is logarithmic.

[0037] 10. The method for predicting the cargo load of ships of various cargo types based on multi-source heterogeneous data and machine learning models according to claim 9, characterized in that the model in S7 includes: K-nearest neighbors, decision tree, random forest, extreme gradient boosting, gradient boosting decision tree, or additional tree model; the evaluation of the model output data in S7 to obtain the optimal ship cargo load prediction model corresponding to different types of loaded cargo is specifically as follows:

[0038] The hyperparameters of the Bayesian search optimization model are used. After the outputs of different models are inversely normalized and inversely logarithmicized, the accuracy is calculated for samples with actual cargo loads of 0 and non-zero using formulas (4)-(6), and the optimal model is saved.

[0039] (4)

[0040] (5)

[0041] (6)

[0042] in: Calculations are performed for samples where y_real is non-zero; Calculations are performed for samples where y_real is 0; Calculate the model's accuracy for all samples;

[0043] Tests revealed that the K-nearest neighbor algorithm provides higher cargo load prediction accuracy and better generalization performance. The specific calculation process of the K-nearest neighbor algorithm is as follows:

[0044] (1) Calculate the Euclidean distance using the following formula:

[0045] (7)

[0046] in: Indicates the sample to be predicted; The training set samples are represented by n; the feature dimension is represented by k; and the k-th feature is represented by k.

[0047] (2) The predicted value is calculated based on the weighted average, and the formula is as follows:

[0048] (8)

[0049] Where: K represents the number of nearest neighbors; This represents the target value of the i-th neighbor; This represents the distance between the point to be predicted and its i-th neighbor.

[0050] Beneficial effects

[0051] This invention provides a method for predicting the cargo load of multiple types of ships based on multi-source heterogeneous data and machine learning models. It integrates ship segment data and port entry / exit data, matching the types of cargo loaded in the port entry / exit data to the ship segment data. The dataset is then categorized according to different cargo types and further processed: draft and cargo values ​​are corrected, outliers are removed, and missing values ​​are filled. By integrating multi-source data and processing it, the overall quality of the dataset is improved, thereby enhancing the reliability of ship cargo load prediction and addressing the problem of poor data reliability in existing ship cargo load estimation methods. Different prediction models are established for different cargo types, and various normalization and logarithmic transformation methods are used to process the dataset. Finally, the prediction accuracy of different cargo type prediction models under different conditions is tested to obtain the optimal ship cargo load prediction model for each cargo type. Testing shows that the obtained optimal ship cargo load prediction model has high prediction accuracy, solving the problem of low prediction accuracy in existing ship cargo load estimation methods.

[0052] The method provided in this application develops new evaluation metrics and tests the constructed dataset with different data preprocessing methods for different types of loaded cargo based on a machine learning model. The accuracy of the prediction method in this application is compared with traditional methods. Furthermore, the method provided in this application can accurately predict the cargo capacity of ships transporting different types of cargo, thus providing relevant data support for port throughput estimation and further providing an evaluation basis for the port's influence index. In addition, the model output provided in this application can also be combined with the draft and cargo capacity correction processing in the data processing steps to perform secondary correction for abnormal draft in business data, providing support for other research involving draft. Attached Figure Description

[0053] Figure 1This is a flowchart of a method for predicting the cargo volume of multiple types of ships based on multi-source heterogeneous data and machine learning models, according to an embodiment of the present invention.

[0054] Figure 2 This is an example diagram illustrating the theoretical relationship between ship cargo capacity and draft in an embodiment of the present invention.

[0055] Figure 3 This is a schematic diagram of abnormal data marking in an embodiment of the present invention.

[0056] Figure 4 This is a schematic diagram showing the distribution of filled TPCI data in an embodiment of the present invention; wherein: the red part is the filled data, and the blue part is the original data. Specific implementation methods

[0057] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions will be further described clearly and completely below with reference to the accompanying drawings.

[0058] This invention provides a method for predicting the cargo load of multiple cargo types of ships based on multi-source heterogeneous data and machine learning models, such as... Figure 1 As shown, it includes the following steps:

[0059] S1: Integrate ship segment data and port entry / exit data, and match the types of cargo loaded in the port entry / exit data to the ship segment data; find ships loaded with crude oil through the tank volume field, and add or update the cargo type to crude oil in the ship segment data; since it is necessary to build a prediction model for ships transporting different types of cargo (cargo types), and there is no cargo type information in the ship segment data, only in the port entry / exit data, it is necessary to statistically classify the cargo name and cargo type based on the relevant fields in the port entry / exit data. The specific results are shown in Table 1. Furthermore, since the liquid (tank volume) field in the data of ships transporting crude oil is not empty, to avoid omissions, all ship segment data with non-empty liquid fields are identified. The cargo type in the ship segment data that is not classified as crude oil through the above mapping relationship is then set to crude oil type. In a specific embodiment of this invention, data queried based on MMSI (unique ship identification number) is classified as crude oil type. Specifically, this involves: identifying ships carrying crude oil, finding the corresponding MMSI based on the liquid field, and then setting the cargo type in the ship segment data corresponding to these ships to crude oil type based on the MMSI. The query refers to filtering a specified field based on constraints to find specific data, similar to the column filtering function in Excel, and specifically implemented using Python or SQL code. The ship segment data is extracted from AIS data by segmenting it. Specifically, the ship's navigation status is determined in the AIS data using fields such as speed, longitude, and latitude (for example, if the latitude and longitude of multiple data points do not change significantly and the speed is in the range of 1 to 2 knots, it is considered to be anchored; if the speed is in the range of 0 to 1 knots, it is considered to be berthed). In this way, the AIS data can be divided into three states: berthed, anchored, and sailing. Each segment represents only one state and the length is not fixed.

[0060] In a specific embodiment of the present invention, based on the cargo type record information in the ship entry and exit port data, the types of cargo loaded in the collected data can be divided into the following types. According to the relevant types in the sub-column and the corresponding code number, the overall data is divided into the corresponding types (crude oil, LNG, LPG, coal, grain, ore, dry bulk). Some types are shown in Table 1.

[0061] Table 1. Statistics on Cargo Types for Ship Arrivals and Departures (Partial)

[0062]

[0063] S2: The dataset is divided according to the types of cargo loaded in the ship segment data in S1, and different types of cargo loaded correspond to different datasets; the data in the dataset is ship segment data with a cargo loading type field.

[0064] S3: Correct the draft and cargo load data in the dataset obtained in S2, including: correcting data where the water level did not change before and after berthing operations using nearby draft data, correcting data with excessive empty draft, and correcting data where draft and cargo load are not positively correlated. Theoretically, the heavier the cargo loaded on a ship, the greater its draft, i.e., there is a positive correlation between the two. After observing the distribution of draft and cargo load, the following problems were found and their solutions are proposed:

[0065] ① Regarding the ship segment data where dynamic_type is 5 (dynamic_type represents the ship's navigation state; 1 for sailing, 0 for anchoring, and 5 for berthing; only in the berthing state (5) will the ship engage in loading and unloading, and the draft corresponding to the cargo load is more accurate in this case), the starting draft (start_draught) and ending draft (end_draught) should change. However, if the starting draft (start_draught) and ending draft (end_draught) do not change, then there is a problem with the draft data in the ship segment data. To address this issue, the specific approach is as follows: From the data described in S2... Data with dynamic_type 5 is extracted from the centralized ship segment data. Each data point is treated as the start of a voyage, and the process is repeated sequentially. For ship segment data where the loading / unloading status changes, if the starting and ending drafts remain unchanged, the process iterates through the next adjacent segment (ship segment data) line by line, finding and replacing data with changed start_draught. "Line by line" means starting from the line where the starting and ending drafts remain unchanged and recursively searching downwards to find data that meets the conditions. "Replacement" means replacing the ending draft in the ship segment data where the starting and ending drafts remain unchanged with the starting draft of the found data.

[0066] ② Abnormal draft under no-load conditions, i.e., the cargo load is 0, but the draft is too high, deviating from the draft corresponding to the ship when it is empty; according to relevant literature on ship structural principles, the draft of the ship's lowest midship section is taken as the draft corresponding to the no-load condition. The specific calculation method is as follows:

[0067] (1)

[0068] in: Indicates the minimum midship draft, Indicates the length of the ship.

[0069] The specific solutions to the problem of abnormal draft under no-load conditions are as follows:

[0070] Step 1: Query all ship segment data with a cargo load of 0 in the dataset obtained from S2;

[0071] Step 2: Calculate the absolute value 'a' of the actual draft and the minimum midship draft in the ship segment data, and based on empirical observation, take 5% of the actual draft of the ship as the allowable draft deviation 'b'.

[0072] Step 3: Compare the size of a and b. If a is greater than b, it means that the actual draft in the empty data exceeds the maximum allowable empty draft of the corresponding ship. Generate random numbers in the interval [minimum midship draft, minimum midship draft + b] to replace the actual draft data in the ship segment data (random seed is set to 42). The purpose is to allow the model to maintain good generalization performance when learning empty data.

[0073] Step 4: Compare the size of a and b. If a is less than or equal to b, it means that the actual draft in the no-load data is within the maximum allowable no-load draft of the corresponding ship. In this case, the actual draft data in the ship segment data is retained.

[0074] ③ Cargo load anomaly: 1. Based on the actual draft, the ship is theoretically empty, but the actual data shows otherwise; the actual draft needs to be corrected to 0. 2. Based on the actual draft, the ship is theoretically not empty, but in the actual data, within each MMSI (MMSI refers to the ship's unique identification number, which is internationally recognized and unique) (grouped by MMSI, each MMSI corresponds to multiple voyage segments; correction of cargo load anomalies requires processing each ship individually), there is no positive correlation between the actual data and the draft. The specific solutions for these two problems are as follows:

[0075] Step 1: Calculate the absolute value c of the actual draft and minimum midship draft in all ship segment data in the dataset, and take 5% of the actual draft as the allowable draft deviation b.

[0076] Step 2: Compare the size of c and b. If c is less than or equal to b, it means that the actual draft is close to the minimum midship draft. In this case, the cargo load is likely to be 0. Therefore, the cargo load in the ship segment data is corrected to 0.

[0077] Step 3: Compare the values ​​of c and b. If c is greater than b, it means that the actual draft is not close to the minimum midship draft. In this case, the ship is likely not empty, so further processing is required. The specific processing procedure is as follows: Calculate the expected cargo capacity based on the following formula:

[0078] (2)

[0079] Step 4: Calculate the percentage deviation between the expected cargo volume and the actual cargo volume, using the following formula:

[0080] (3)

[0081] Step 5: If the deviation percentage realcargo_error is greater than 5%, replace the cargo volume in the ship segment data with the expected cargo volume; otherwise, keep the cargo volume data in the ship segment data unchanged.

[0082] In a specific embodiment of the present invention, the corrected relevant information is recorded as shown in Table 2, based on the correction logic proposed according to the correlation between draft and cargo capacity.

[0083] Table 2. Statistical data on revised draft and cargo capacity.

[0084]

[0085] S4: The data in the dataset obtained after S3 processing is used to check whether the ship's draft and cargo capacity are positively correlated based on ship structure theory. If not, it is judged as abnormal data and removed from the dataset. In a specific embodiment of the present invention, the accuracy of the draft data directly affects the predictive ability of the model. After the correlation processing in S3, in order to ensure that the draft and cargo capacity data remain logically accurate and positively correlated, anomalies in draft and cargo capacity are identified based on scatter plots and ship structure theory. Figure 2 This diagram illustrates the relationship between a ship's cargo capacity and draft, where the x-axis represents draft and the y-axis represents cargo capacity. Point A represents the ship's design draft and maximum cargo weight (DWT). Points B and C represent 90% and 110% of the design draft, respectively (i.e., x = 10%, but x can be adjusted as needed). The y-axis represents the ship's DWT. Point D represents the minimum unloaded draft according to the ship's safety design requirements, and point E represents the upper limit of the unloaded draft, which can be expressed as (minimum midship draft ratio + 0.2) * design draft. The relationship varies depending on the type of cargo being loaded. The basic requirements for draft when navigating empty are as follows: For vessels with a length of 150 meters or less, when navigating empty, the minimum forward draft must be greater than or equal to 0.025 times their length, and the minimum midship draft must be greater than or equal to 0.02 times their length + 2. For vessels with a length greater than 150 meters, when navigating empty, the minimum forward draft must be greater than or equal to 0.12 times their length + 2, and the minimum midship draft must be greater than or equal to 0.02 times their length + 2. Data within the range of quadrilaterals B, C, D, and E above are considered normal data, while data outside this range are considered abnormal data.

[0086] In a specific embodiment of the present invention, the result of abnormal data identification is as follows: Figure 3 As shown in the figure; the statistics of normal data and abnormal data are shown in Table 3.

[0087] Table 3. Statistics on Anomaly Identification

[0088]

[0089] S5: For the dataset obtained from S4, a random forest model is used to fill in the missing values ​​in the dataset. The main field with missing values ​​in the dataset obtained after S4 is tpci (the change in deadweight tons per centimeter of ship draft). This can reflect the ship size to a certain extent and thus determine the possible cargo capacity. Therefore, we consider filling this feature. We use models such as RF, GBDT, and XGBoost, which can be used to check the importance of features. We take the maximum deadweight tons (dwt), length (length), width (width), height (height), and design draft (draught) as inputs, and select the output of the model with the highest R² and the lowest MAE as the filling result.

[0090] In a specific embodiment of the present invention, three models, RF, GBDT, and XGBoost, are used to fill in the missing data. The feature importance between the target TPCI and the input features is shown in Table 4.

[0091] Table 4. Feature Importance of TPCI Filler in Different Models

[0092]

[0093] The evaluation results of the three models, RF, GBDT, and XGBoost, are shown in Table 5.

[0094] Table 5. Evaluation results of TPCI filling for different models

[0095]

[0096] The distribution of the ppc data after padding is as follows Figure 4 As shown.

[0097] S6: Preprocess the data in the dataset corresponding to different types of loaded goods to obtain input data corresponding to different types of loaded goods; in a specific embodiment of the present invention, firstly, normalization and logarithmic transformation methods are used to form new input data according to different types of loaded goods in the dataset. The normalization formula is as follows:

[0098]

[0099] in: It is the original feature data. It is the minimum value among the feature data of each dimension. It is the maximum value in the feature data of each dimension. These are the normalized feature data.

[0100] The formula for logarithmic transformation is as follows:

[0101]

[0102] in: This is the raw cargo volume data. This is the logarithmic cargo volume data.

[0103] This invention employs logarithmic transformation to process cargo volume data because ship cargo volume data follows a right-skewed distribution, meaning there are no negative values. Compared to normalized data preprocessing methods, logarithmic transformation can better correct the data distribution, making it approximately normal, while also compressing the data scale, resulting in a more stable model output. Especially for samples with a cargo volume of 0, using the 1+y form to transform into logarithmically avoids accuracy loss and helps the model output with smaller bias.

[0104] S7: Based on K-nearest neighbors, decision tree, random forest, extreme gradient boosting, gradient boosting decision tree, or extreme random model, construct multiple ship cargo volume prediction models corresponding to different types of loaded cargo, train the prediction models using input data corresponding to the types of loaded cargo, and evaluate the model output data to obtain the optimal ship cargo volume prediction model corresponding to different types of loaded cargo; In a specific embodiment of the present invention, based on K-nearest neighbors, decision tree, random forest, extreme gradient boosting, gradient boosting decision tree, or extreme random model, construct multiple ship cargo volume prediction models corresponding to different types of loaded cargo, and perform prediction after different data preprocessing methods. For the different prediction models established, use Bayesian search to optimize the hyperparameters of the model, and after inverse normalization and inverse logarithmicization of the model output, use the following formulas (4)-(6) to obtain the indicators to calculate the accuracy of the actual cargo volume for samples with 0 and non-0 cargo volumes and the model output, and save the optimal model; where formulas (4)-(6) are as follows:

[0105] (4)

[0106] (5)

[0107] (6)

[0108] in: Calculations are performed for samples where y_real is non-zero; Calculations are performed for samples where y_real is 0; Calculate the model's accuracy for all samples;

[0109] Tests revealed that the K-nearest neighbor algorithm provides higher cargo load prediction accuracy and better generalization performance. The specific calculation process of the K-nearest neighbor algorithm is as follows:

[0110] Step 1: Calculate the Euclidean distance using the following formula:

[0111] (7)

[0112] in: Indicates the sample to be predicted. Let n represent the training set sample, n represent the number of feature dimensions, and k represent the k-th feature.

[0113] Step 2: Calculate the predicted value based on the weighted average, using the following formula:

[0114] (8)

[0115] Where: K represents the number of nearest neighbors; This represents the target value of the i-th neighbor; This represents the distance between the point to be predicted and its i-th neighbor.

[0116] In a specific embodiment of the present invention, taking a ship transporting ore cargo as an example, firstly, grouped cross-validation is used to evaluate the training data, normalization is used for the input features, and normalization and logarithmic transformation are used for the output data cargo volume, respectively. The evaluation results of different data preprocessing models on the training set are shown in Table 6.

[0117] Table 6. Evaluation results of different data preprocessing models on the training set

[0118]

[0119] The optimal model was further selected and compared with traditional methods using validation data. The overall results are shown in Table 7.

[0120] Table 7. Evaluation Comparison of Traditional Formula and KNN Model on Training and Validation Sets

[0121]

[0122] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.

Claims

1. A method for predicting cargo load of multi-cargo type vessels based on multi-source heterogeneous data and machine learning models, characterized in that, Includes the following steps: S1: Integrate ship segment data and port entry / exit data, and match the cargo types loaded in the port entry / exit data to the ship segment data; find ships loaded with crude oil through the tank volume field, and add or update the cargo type as crude oil in the ship segment data. S2: The dataset is divided according to the types of cargo loaded in the ship segment data in S1, with different types of cargo corresponding to different datasets; S3: Correct the draft and cargo load of the data in the dataset obtained in S2, including: using nearby draft data to correct data where the water output did not change before and after berthing operations, correcting data with excessive empty draft, and correcting data where draft and cargo load are not positively correlated. S4: The data in the dataset obtained after processing S3 is checked for a positive correlation between ship draft and cargo capacity based on ship structure theory. If not, it is judged as abnormal data and removed from the dataset. S5: For the dataset obtained from S4, use the random forest model to fill in the missing values ​​in the dataset; S6: Preprocess the data in the dataset corresponding to the type of loaded goods to obtain the input data corresponding to different types of loaded goods; S7: Based on different models, construct multiple ship cargo volume prediction models corresponding to different types of loaded cargo, train the prediction models using input data corresponding to the types of loaded cargo, and evaluate the model output data to obtain the optimal ship cargo volume prediction model corresponding to different types of loaded cargo.

2. The method for predicting cargo load of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 1, characterized in that, The ship segment data is segmented and extracted from AIS data; in S1, for all ship data in the ship segment data where the liquid tank volume field is not empty, the cargo type is classified as crude oil; in S2, the data in the dataset is ship segment data with a cargo type field.

3. The method for predicting cargo volume of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 2, characterized in that, In S3, adjacent draft data is used to correct data where the water level did not change before and after berthing operations. Specifically, the data of all ship segments with berthing status are extracted from the dataset obtained in S2. Each data point is taken as the start of a voyage and is traversed sequentially. For ship segments with changes in loading and unloading status, if the starting and ending drafts have not changed, the data of the next adjacent ship segment is traversed line by line until the data with the change in starting draft is found and replaced.

4. The method for predicting cargo volume of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 2, characterized in that, In S3, the correction for excessive draft when the ship is empty is as follows: For ship segments with zero cargo load but excessive draft data that deviates from the ship's draft when empty, the solution includes the following steps: Step 1: Query all ship segment data with a cargo load of 0 in the dataset obtained from S2; Step 2: Calculate the absolute value of the actual draft and the minimum midship draft in the ship segment data, and take 5% of the actual draft as the allowable draft deviation based on empirical observations; Step 3: Compare the absolute value with the draft deviation. If the absolute value is greater than the draft deviation, it means that the actual draft in the empty data exceeds the maximum allowable empty draft of the corresponding vessel. Generate a random number from the interval [minimum midship draft, minimum midship draft + draft deviation] to replace the actual draft data in the vessel segment data. If the absolute value is less than or equal to the draft deviation, it means that the actual draft in the empty data is within the maximum allowable empty draft of the corresponding vessel. Then retain the draft data in the vessel segment data.

5. The method for predicting cargo volume of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 2, characterized in that, The formula for calculating the minimum draft of a ship is: (1) in: Indicates the minimum midship draft, Indicates the length of the ship.

6. The method for predicting cargo volume of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 5, characterized in that, In S3, data where draft and cargo load are not positively correlated are corrected as follows: Based on actual draft conditions, if the ship is theoretically empty but the actual data does not, then the actual draft is corrected to 0; if the ship is theoretically not empty, but the actual data does not show a positive correlation with draft, then the following processing is performed: Step 1: Calculate the absolute value of the actual draft and the minimum midship draft in the ship segment data of all datasets. Take 5% of the actual draft as the allowable draft deviation. Step 2: Compare the absolute value with the draft deviation. If the absolute value is less than or equal to the draft deviation, it means the actual draft is close to the minimum midship draft, so the cargo load in the ship segment data is corrected to 0. If the absolute value is greater than the draft deviation, it means the actual draft is not close to the minimum midship draft. In this case, the ship is not empty and further processing is required. The specific processing procedure is as follows: Calculate the expected cargo capacity based on formula (2): (2) Step 3: Calculate the percentage deviation between the expected cargo volume and the actual cargo volume, using the following formula: (3) Step 4: If the deviation percentage is greater than 5%, replace the cargo volume in the ship segment data with the expected cargo volume; otherwise, keep the cargo volume data in the ship segment data unchanged.

7. The method for predicting cargo volume of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 6, characterized in that, The method for judging abnormal data in S4 is as follows: With draft as the x-axis and cargo capacity as the y-axis, draw a diagram showing the relationship between the ship's cargo capacity and draft. Point A represents the ship's design draft and maximum cargo capacity. Points B and C represent 90% and 110% of the design draft, respectively. The y-axis represents the ship's maximum cargo capacity. Point D represents the minimum empty draft according to the ship's safety design requirements. Point E represents the upper limit of the empty draft. Data within the range of quadrilateral BCDE is considered normal data, while data outside this range is considered abnormal data.

8. The method for predicting cargo volume of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 7, characterized in that, The specific method for handling missing values ​​in S5 is as follows: In the dataset obtained from S4, the field with missing values ​​is the change in deadweight tons per centimeter of ship draft. Models such as RF, GBDT, or XGBoost, which are used to examine the importance of features, are used. The maximum deadweight tons, length, width, height, and actual draft are used as inputs. The output of the model with the highest R² and the lowest MAE is selected as the filling result.

9. The method for predicting cargo volume of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 8, characterized in that, The preprocessing described in S6 is logarithmic transformation.

10. The method for predicting cargo volume of multi-cargo vessels based on multi-source heterogeneous data and machine learning models according to claim 9, characterized in that, The models described in S7 include: K-nearest neighbors, decision trees, random forests, extreme gradient boosting, gradient boosting decision trees, or additional tree models; S7 evaluates the model output data to obtain the optimal ship cargo volume prediction model corresponding to different types of loaded cargo, specifically as follows: The hyperparameters of the Bayesian search optimization model are used. After the outputs of different models are inversely normalized and inversely logarithmicized, the accuracy is calculated for samples with actual cargo loads of 0 and non-zero using formulas (4)-(6), and the optimal model is saved. (4) (5) (6) in: Calculations are performed for samples where y_real is non-zero; Calculations are performed for samples where y_real is 0; Calculate the model's accuracy for all samples; Tests revealed that the K-nearest neighbor algorithm provides higher cargo load prediction accuracy and better generalization performance. The specific calculation process of the K-nearest neighbor algorithm is as follows: (1) Calculate the Euclidean distance using the following formula: (7) in: Indicates the sample to be predicted; The training set samples are represented by n; the feature dimension is represented by k; and the k-th feature is represented by k. (2) The predicted value is calculated based on the weighted average, and the formula is as follows: (8) Where: K represents the number of nearest neighbors; This represents the target value of the i-th neighbor; This represents the distance between the point to be predicted and its i-th neighbor.