Ship classification and prediction method and system based on cosine similarity optimization
By combining cosine similarity optimization and random forest model, the sample imbalance problem in ship classification was solved, high-precision ship classification and fuel consumption prediction were achieved, and operational efficiency and automation were improved.
Patent Information
- Application Number
- CN202511035625.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-25
AI Technical Summary
The existing ship classification method suffers from sample imbalance, resulting in insufficient number of model training samples, insufficient classification model accuracy and generalization ability, and inability to effectively utilize the similarities between ships, affecting the accuracy of fuel consumption prediction and operational efficiency.
The cosine similarity optimization method is used to reclassify ship categories, and the categories with insufficient sample numbers are merged into the multi-sample category with the highest similarity. The random forest model is then combined for training to construct a ship classification prediction model. The data is preprocessed and standardized to reduce noise interference.
It improves the accuracy and generalization ability of ship classification, reduces classification errors, achieves the accuracy of fuel consumption prediction and full process automation, reduces human errors, and improves operational efficiency.
Smart Images

Figure CN120804915A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of maritime management and artificial intelligence, in particular to a ship classification and prediction method and system based on cosine similarity optimization. BACKGROUND
[0002] In modern ship management and fuel consumption prediction, different types of ship classification have a crucial impact on the accuracy and operational efficiency of the prediction model. The fuel consumption of a ship is closely related to the ship's structural size, performance parameters and operating state. Therefore, scientific classification of individual ships can not only improve the accuracy of the fuel prediction model, but also provide important support for the operational optimization, fuel cost control and environmental protection of ship companies.
[0003] The ship classification method of a ship depends mainly on expert experience, and individual ships are classified into different ship groups based on their size, design purpose, load capacity and other characteristics. These classifications are usually based on subjective judgments, such as the tonnage, draft, speed or engine type of the ship. Then, with the diversification of ship types, relying solely on manual experience for classification gradually faces the following problems: 1) sample imbalance problem: some ship types contain a very small number of even a single ship, which leads to insufficient number of training samples for the model, and the classification model cannot effectively learn the characteristics of these classes, resulting in a decrease in model accuracy; 2) insufficient model generalization ability: traditional models perform poorly when dealing with sample imbalance, and prediction of single ship classes is prone to bias. This bias not only affects fuel consumption prediction, but also may lead to errors in management and operational decision-making of ship companies; 3) inability to effectively utilize similarity: different ships may have certain similarities, such as ships of similar size but different purposes, and if these ships can be reclassified based on similarity, it will help improve the performance of the classification model. However, traditional methods are difficult to fully capture these similarities.
[0004] Therefore, there is an urgent need for a ship classification and mapping method that combines cosine similarity and machine learning to solve the sample imbalance problem and achieve accurate ship classification and management. SUMMARY
[0005] To solve the problems of sample imbalance, insufficient model generalization ability and inability to effectively utilize similarity in the existing ship classification process, the present application provides a ship classification and prediction method based on cosine similarity optimization, which effectively solves the sample imbalance problem in ship management and fuel consumption prediction, and can effectively filter out noise interference in the data, ensuring the stability and reliability of the final classification results, and effectively improving the accuracy and generalization ability of the ship classification prediction model. The present application also relates to a ship classification and prediction system based on cosine similarity optimization.
[0006] The technical scheme of the present application is as follows:
[0007] A ship classification and prediction method based on cosine similarity optimization, characterized by comprising the following steps:
[0008] Ship feature acquisition and preliminary classification step: collect multi-dimensional original feature data of ships covering physical characteristics of ships and performance parameters reflecting power level and running efficiency of ships, pre-process the collected multi-dimensional original feature data, preliminarily classify ships based on the pre-processed multi-dimensional original feature data and combined with experience rules to obtain a preliminary classification result containing multiple ship categories; then count the sample number of each ship category in the preliminary classification result, mark the ship category with a sample number less than a preset number threshold as an optimization category, and mark other ship categories except the optimization category as reference categories;
[0009] Ship category optimization step: convert each feature parameter in the pre-processed multi-dimensional original feature data into a feature vector, calculate the cosine similarity between each ship in the optimization category and all ships in the reference categories based on the feature vector and using the cosine similarity method, and then generate a cosine similarity matrix; find out the target ship with the highest similarity to each ship in the optimization category in the cosine similarity matrix, and update the category label of each ship in the optimization category to the ship category to which the corresponding target ship belongs; when all ships in the optimization category are reclassified, delete the optimization category, and then obtain an optimized ship classification result;
[0010] Model construction and training step: use the pre-processed multi-dimensional original feature data and the optimized ship classification result as a data set, divide the data set into a training set and a test set according to a preset ratio, input the training set into a random forest model for training, obtain a trained random forest model and test it through the test set, verify the model performance, and obtain a final trained and tested ship classification prediction model;
[0011] Model category prediction step: use the ship classification prediction model to predict the category of a new ship to be detected, and predict the ship category to which the new ship belongs.
[0012] Preferably, in the ship feature acquisition and preliminary classification step, the pre-processing includes missing value processing and outlier removal processing of the multi-dimensional original feature data to obtain processed multi-dimensional feature data; and a standardization processor is used to standardize the processed multi-dimensional feature data to obtain standardized feature data.
[0013] Preferably, in the model construction and training step, the preset proportion is that the training set accounts for 70-80% and the test set accounts for 20-30%; the random forest model is an ensemble learning model based on decision trees, which is composed of 100 independent decision trees, each of which is trained by randomly sampling with replacement from the training set to generate a sample subset and randomly selecting a feature subset in each node splitting process; at the same time, the maximum depth of each decision tree is set to 20; after all the decision trees are trained, the final prediction result is determined by ensemble voting.
[0014] Preferably, in the model construction and training step, the test set is used for test analysis to verify the model performance, which specifically includes:
[0015] The performance of the trained random forest model is evaluated by using the test set and evaluation indicators, the evaluation indicators include detection speed, detection accuracy and complexity, the detection accuracy includes accuracy, recall rate and average precision, the detection speed uses the number of samples processed per second, and the complexity uses the number of parameters and floating point operations.
[0016] Preferably, the calculation step of the accuracy rate specifically includes: inputting each ship data in the test set into the trained random forest model to output the prediction result; and based on the prediction result and the true label manually labeled by each ship data in the test set, the true positive, true negative, false positive and false negative are counted, and the accuracy rate is calculated according to the true positive, true negative, false positive and false negative.
[0017] Preferably, in the ship feature acquisition and preliminary classification step, the physical properties in the multi-dimensional original feature data include deadweight ton, total tonnage of the ship, net tonnage of the ship, ship length, ship width, designed draft, displacement, depth, cylinder diameter of the main engine and cylinder number of the main engine, the performance parameters reflecting the power level of the ship in the multi-dimensional original feature data include main engine stroke type, maximum power, braking power, operating power, total horsepower of the main engine and total power of the main engine, and the performance parameters reflecting the running efficiency of the ship in the multi-dimensional original feature data include designed speed, maximum speed and operating speed.
[0018] A ship classification and prediction system based on cosine similarity optimization, characterized by sequentially connected ship feature acquisition and preliminary classification module, ship category optimization module, model construction and training module, and model category prediction module,
[0019] The ship feature collection and preliminary classification module collects multi-dimensional original feature data of the ship, which covers physical characteristics of the ship and performance parameters reflecting power level and running efficiency of the ship, pre-processes the collected multi-dimensional original feature data, preliminarily classifies the ship based on the pre-processed multi-dimensional original feature data and in combination with experience rules, and obtains a preliminary classification result containing multiple ship categories; then, the sample number of each ship category in the preliminary classification result is counted, a ship category with a sample number less than a preset number threshold is recorded as an optimized category, and other ship categories except the optimized category are recorded as reference categories;
[0020] The ship category optimization module converts each feature parameter in the pre-processed multi-dimensional original feature data into a feature vector, calculates the cosine similarity between each ship in the optimized category and all ships in the reference categories based on the feature vector and by using a cosine similarity method, and further generates a cosine similarity matrix; and finds out a target ship with the highest similarity to each ship in the optimized category in the cosine similarity matrix, and updates the category label of each ship in the optimized category to the ship category to which the corresponding target ship belongs; when all ships in the optimized category are reclassified, the optimized category is deleted, and an optimized ship classification result is obtained;
[0021] The model construction and training module takes the pre-processed multi-dimensional original feature data and the optimized ship classification result as a data set, divides the data set into a training set and a test set according to a preset proportion, inputs the training set into a random forest model for training, obtains a trained random forest model, and tests and analyzes the trained random forest model by using the test set to verify the performance of the model, and obtains a final trained and tested ship classification prediction model;
[0022] The model category prediction module uses the ship classification prediction model to predict the category of a new ship to be detected, and predicts the ship category to which the new ship belongs.
[0023] Preferably, in the ship feature collection and preliminary classification module, the pre-processing includes missing value processing and outlier removal processing on the multi-dimensional original feature data to obtain processed multi-dimensional feature data; and a standardization processor is used to standardize the processed multi-dimensional feature data to obtain standardized feature data.
[0024] Preferably, in the model construction and training module, the test set is used for test analysis to verify the performance of the model, and specifically includes:
[0025] The performance of the trained random forest model is evaluated by using a test set and evaluation indexes, the evaluation indexes include detection speed, detection accuracy and complexity, the detection accuracy includes accuracy, recall rate and average precision, the detection speed uses the number of samples processed per second, and the complexity uses the number of parameters and the number of floating point operations.
[0026] Preferably, in the model construction and training module, when the model performance is verified by test analysis through the test set, the calculation steps of the accuracy rate specifically include: inputting each ship data in the test set into the trained random forest model to output a prediction result; and based on the prediction result and the true label manually labeled for each ship data in the test set, counting true positives, true negatives, false positives and false negatives, and calculating the accuracy rate according to the true positives, the true negatives, the false positives and the false negatives.
[0027] The technical effects of the present application are as follows:
[0028] The application provides a ship classification and prediction method based on cosine similarity optimization, which is used to solve the sample imbalance problem in ship management and fuel consumption prediction, and improve the accuracy and generalization ability of the classification model. First, the ship feature collection and preliminary classification step collects multi-dimensional original feature data of the ship (including physical characteristics, power level performance parameters and operation efficiency performance parameters) and performs pretreatment, and then combines experience rules to preliminarily classify the ship to obtain a preliminary classification result containing multiple ship categories. The existing knowledge system is used to quickly establish a classification framework, which helps to quickly identify common ship categories, improve the initial classification efficiency, and provide a reliable starting point for subsequent optimization. At the same time, it reduces the misclassification risk caused by complete dependence on algorithms and improves the system stability. Then, the sample number of each ship category in the preliminary classification result is counted, and the ship category with a sample number less than the preset number threshold is recorded as the to-be-optimized category, and the other ship categories except the to-be-optimized category are recorded as the reference category. That is, the sample amount of the to-be-optimized category is insufficient, and the sample amount of the reference category is sufficient. The system can dynamically adjust the classification strategy to effectively avoid the classification error or model overfitting problem caused by the small sample. Based on the experience rule preliminary classification, the preliminary classification of the ship is quickly realized, and the classification framework is clear. At the same time, the to-be-optimized category (small sample category) is identified through sample amount screening, the weak link (small sample is easy to cause classification deviation) in the classification is accurately positioned, and the clear target is provided for subsequent optimization. Then, the ship category optimization step is used to reclassify the categories by introducing the cosine similarity method for the categories with insufficient sample number after preliminary classification. The similarity between ships is used to classify ships with similar features into the same category, so that the optimized ship classification result is more consistent with the actual distribution of ship features, and can accurately reflect the internal correlation between different ships. It is ensured that even a small number of sample categories can also get reasonable classification results, which not only improves the comprehensiveness of classification, but also reduces the model bias caused by sample sparseness, effectively solves the sample imbalance problem in ship management and fuel consumption prediction, and reclassifies single-sample or small-sample ships into similar multi-sample categories through cosine similarity, i.e. secondary classification, to ensure the balance of sample number in the classification label and improve the subsequent model training effect.The model construction and training step is to construct a ship classification prediction model based on the multi-dimensional original feature data after preprocessing and the optimized ship classification result, train the model based on the optimized ship classification result, avoid the small sample label error from being transmitted to the model, improve the learning accuracy of the model on the ship class, and the random forest model has strong adaptability to the multi-dimensional features of the ship (such as physical characteristics, power parameters, and operation efficiency parameters), can effectively process high-dimensional data, has good generalization ability, anti-overfitting ability and robustness to noise, the model is trained based on the optimized ship classification result (the small sample class has been corrected), and the performance is verified through the test set, so that the reliability of the model in distinguishing different ship classes can be ensured, and finally a reusable prediction model that is free from the dependence on “experience rules” is formed, thereby improving the accuracy of the classification mapping model, efficiently predicting the type of a new ship in practical application, and realizing the upgrade from “manual / rule classification” to “model automatic classification”; in addition, the random forest model can effectively filter out noise interference in the data through comprehensive analysis of the results of a large number of decision trees, thereby ensuring the stability and reliability of the final classification result. The model class prediction step is to use the trained and tested ship classification prediction model to predict the class of a new ship to be detected, predict the ship class to which the new ship belongs, realize the rapid and automatic classification of the new ship without manual experience-based judgment, improve the classification prediction efficiency (especially suitable for batch new ship detection scenarios), and based on the optimized data and the trained model, the accuracy and consistency of the prediction result are higher (compared with experience rules or unoptimized models), and human error is reduced.
[0029] The ship classification and prediction method based on cosine similarity optimization disclosed in the application can also be referred to as a ship type classification and mapping method based on cosine similarity calculation and random forest ensemble learning, which relates to preliminary classification through experience rules and reclassification (secondary classification) of small sample classes in the preliminary classification through cosine similarity optimization, obtains more reliable ship classification results, and is used as a basis to train a ship classification prediction model by combining random forest ensemble learning technology, and finally realizes accurate mapping of ship types and efficient prediction of new ship classes.
[0030] The present application realizes the following overall technical effects by introducing a small sample optimization mechanism based on cosine similarity combined with a random forest model: 1) balanced sample quantity: traditional ship classification relies heavily on expert experience, and individual ships are divided into several categories. However, the number of samples in some categories is extremely small (such as only one ship), and this sample imbalance affects model training and generalization, easily leading to model overfitting or prediction bias for specific categories. The present application reclassifies single-sample or low-sample ships into similar multi-sample categories through cosine similarity, ensuring balanced sample quantity in the classification label, effectively improving model training results and effectively solving the classification error problem in small sample categories. 2) improved classification accuracy: multi-dimensional feature parameters (such as deadweight tonnage, total tonnage, ship length, etc.) are used to construct feature vectors, and similar ships are classified into the same category based on their similarity, making the classification results more consistent with the actual distribution of ship features. The cosine similarity is used to quantify the similarity between ships, i.e., the cosine similarity is introduced to reclassify categories, which can accurately reflect the internal relevance between different ships. This method can accurately measure the similarity between different ships while maintaining high-dimensional features, thereby improving classification accuracy. 3) small classification error and enhanced robustness: the random forest model consists of multiple decision trees and can handle complex nonlinear relationships and high-dimensional data. Through learning on the training set, the model can effectively capture patterns in the data and reduce classification errors. At the same time, due to the complexity of the marine environment, data inevitably contains various outliers and noise. The random forest model performs well in the presence of noise data due to its ensemble learning characteristics, maintaining high prediction accuracy even in the presence of outliers or noise. In addition, the cosine similarity method itself has some resistance to noise in the data. 4) improved accuracy: based on the preliminary classification, the ship categories are reclassified using cosine similarity to ensure that even small sample categories can be reasonably classified. Combined with the random forest model for final classification, the overall classification accuracy is effectively improved. 5) improve the accuracy of the fuel consumption prediction model: accurate ship classification is important for fuel consumption prediction. The ship classification and prediction model established by the present application can provide accurate basic data for fuel consumption prediction, helping ship companies optimize their operations, reduce fuel costs, and achieve more environmentally friendly operations. 6) full-process automation: from multi-dimensional raw feature data acquisition, preliminary classification, category optimization based on cosine similarity to random forest model training and testing, the entire process is highly automated, improving work efficiency, reducing the need for human intervention, and reducing the risk of human error.In summary, the present application successfully solves the key challenges in traditional ship classification through a series of innovative technical means, including a small sample optimization mechanism based on cosine similarity and the application of random forest model, and can be widely applied in maritime traffic management, ship monitoring, shipping data analysis and other fields, with high popularization value and technical prospect.
[0031] Further, the preprocessing includes missing value processing and outlier removal processing on the multi-dimensional original feature data to obtain processed multi-dimensional feature data, and using a standardization processor to standardize the processed multi-dimensional feature data to obtain standardized feature data. The missing value processing ensures the integrity and consistency of the data set, avoids model training bias caused by data missing, improves the stability and accuracy of the subsequent classification algorithm, and ensures that all ships can be classified based on complete data. The outlier removal eliminates possible errors or extreme values, reduces the influence of noise on the classification result, enhances the robustness of the model, and makes it better adapt to data changes within the normal range. Through standardization processing, the influence of different feature magnitudes is eliminated, which can ensure that each feature contributes relatively evenly to model training, avoid the dominant role of high magnitude features, and improve the learning effect of the model.
[0032] Further, the performance of the trained random forest model is evaluated by using the test set through evaluation indicators, including detection speed, detection accuracy and complexity. The detection accuracy includes accuracy, recall rate and average accuracy. The detection speed uses the number of samples processed per second, and the complexity uses the number of parameters and floating point operations. Through performance evaluation indicators, the advantages and disadvantages of the model are identified, the model performance is comprehensively measured, and the efficiency and reliability of the model in different application scenarios are ensured.
[0033] Further, each ship data in the test set is input into the trained random forest model to output a prediction result. Based on the prediction result and the true label manually labeled by each ship data in the test set, the true positive, true negative, false positive and false negative are counted, and the accuracy is calculated according to the true positive, true negative, false positive and false negative. By comparing the prediction result in the test set with the true label, the true positive, true negative, false positive and false negative are counted, and the accuracy is accurately calculated, which can more comprehensively evaluate the overall performance of the model, ensure its reliability in actual application, and provide solid data support for model optimization.
[0034] The application also relates to a ship classification and prediction system based on cosine similarity optimization, which corresponds to the ship classification and prediction method based on cosine similarity optimization described above and can be understood as a system for implementing the ship classification and prediction method based on cosine similarity optimization described above, and comprises a ship feature acquisition and preliminary classification module, a ship category optimization module, a model construction and training module and a model category prediction module which are sequentially connected and work cooperatively. The preliminary classification is performed by using multi-dimensional feature data and combining expert experience, and a small category with insufficient sample quantity is identified as an object to be optimized to provide a basis for subsequent processing. Then, the ship categories are redivided by using the cosine similarity method, a cosine similarity matrix is generated by calculating the similarity between each ship and all other ships, and the ships in the object-to-be-optimized category are re-assigned to more suitable categories, so that the sample imbalance problem is solved. Subsequently, the optimized classification result is used as part of the data set, combined with the original feature data, and used for training and testing by using the random forest model, so that the good generalization ability and noise resistance of the model on high-dimensional data are ensured. Not only the accuracy of ship classification is improved, but also the whole-process automation from data acquisition to classification prediction is realized, and good engineering practicability and expansibility are achieved. The application has wide application prospect and technical potential in the fields of ship management, fuel consumption optimization and maritime big data analysis, and provides strong technical support for related industries. BRIEF DESCRIPTION OF DRAWINGS
[0035] Figure 1 is a flowchart of the ship classification and prediction method based on cosine similarity optimization. DETAILED DESCRIPTION
[0036] The application will be described below in combination with the drawings.
[0037] The application relates to a ship classification and prediction method based on cosine similarity optimization. The cosine similarity between ships is calculated, and ships with a small number of individuals are redivided into ship categories with higher similarity, so that the sample imbalance problem in ship management and fuel consumption prediction is solved, and the accuracy and generalization ability of the classification model are improved. Subsequently, the ship classification mapping model (i.e. the ship classification prediction model) is constructed by using the random forest model for training based on the optimized ship categories, so that accurate ship classification and management are realized. The flowchart of the method is shown in Figure 1 and comprises the following steps in sequence:
[0038] The first step is to collect multi-dimensional original feature data of the ship, which includes physical characteristics of the ship, performance parameters reflecting the power level and running efficiency of the ship. The multi-dimensional original feature data is preprocessed, and the ship is preliminarily classified based on the preprocessed multi-dimensional original feature data and combined with empirical rules to obtain a preliminary classification result containing multiple ship categories. Then, the sample number of each ship category in the preliminary classification result is counted, and the ship category with a sample number less than a preset number threshold is recorded as an optimized category, and the other ship categories except the optimized category are recorded as reference categories.
[0039] Specifically, first, the multi-dimensional original feature data of the ship is collected, which includes the physical characteristics of the ship, the performance parameters reflecting the power level and running efficiency of the ship. The physical characteristics in the multi-dimensional original feature data include deadweight tonnage, total tonnage of the ship, net tonnage of the ship, ship length, ship width, designed draft, displacement, depth (the vertical distance from the deck beam upper edge to the ship bottom keel plate upper surface in the theoretical design line type (not including the thickness of the ship body steel plate)), main engine cylinder diameter, and main engine cylinder number, etc. The performance parameters reflecting the power level of the ship in the multi-dimensional original feature data include main engine stroke type, maximum power, braking power, operating power, main engine total horsepower, and main engine total power (main engine total kilowatt), etc. The performance parameters reflecting the running efficiency of the ship in the multi-dimensional original feature data include designed speed, maximum speed, and operating speed, etc.
[0040] Then, the collected multi-dimensional original feature data is processed for missing values and outliers to obtain processed multi-dimensional feature data. Then, the processed multi-dimensional feature data is standardized using a standardizer StandardScaler to obtain standardized feature data to eliminate the influence of different feature magnitude differences. Then, the ship is preliminarily classified based on the standardized feature data and combined with empirical rules (expert experience) to obtain a preliminary classification result containing multiple ship categories, and each ship category is labeled with a classification label, and the classification label is encoded using a label encoder LabelEncoder. The classification label is the ship's remark information. Finally, the sample number of each ship category in the preliminary classification result is counted, and the ship category with a sample number less than a preset number threshold is recorded as an optimized category, and the other ship categories except the optimized category are recorded as reference categories.
[0041] II. Ship class optimization step: convert each feature parameter in the pre-processed multi-dimensional original feature data into a feature vector, calculate the cosine similarity between each ship in the to-be-optimized class and all ships in the reference class based on the feature vector and using the cosine similarity method, and then generate a cosine similarity matrix; and find the target ship with the highest similarity to each ship in the to-be-optimized class in the cosine similarity matrix, and update the class label of each ship in the to-be-optimized class to the ship class to which the corresponding target ship belongs; after all ships in the to-be-optimized class are reclassified, delete the to-be-optimized class, and then obtain the optimized ship classification result.
[0042] In the preliminary classification, ships are divided into several groups, but some ship classes only contain a small number of samples, or even only one sample. This imbalance will adversely affect the training effect of the model and cannot provide enough training data for the classification model. In order to improve the problem of sample imbalance, the cosine similarity method is introduced to redivide the groups and merge these classes with insufficient samples into the most similar multi-sample class. Specifically, first, convert each feature parameter in the pre-processed multi-dimensional original feature data into a feature vector (the feature vector contains multi-dimensional information describing ship size, load, power, etc.), that is, represent the features of each ship sample in each ship class as a vector. Then, based on the feature vector and using the cosine similarity method, the cosine similarity between each ship in the to-be-optimized class and all ships in the reference class is calculated, and then a cosine similarity matrix is generated; the cosine similarity is an index that measures the similarity between two vectors, and measures the similarity between them by calculating the cosine of the angle between the two vectors. The value is between -1 and 1, and the closer the value is to 1, the more similar the two samples are. The cosine similarity calculation formula is as follows:
[0043]
[0044] In the above formula, A and B are the feature vectors of two ships, A i and B i represent the value of the i-th feature in the feature vector.
[0045] The cosine similarity between each ship in the to-be-optimized class and all ships in the reference class is calculated by the above formula, and then a cosine similarity matrix is generated, which is used to measure the similarity between each pair of ship samples. The cosine similarity matrix is a symmetric matrix, and each element sim(i,j) in the matrix represents the similarity between ship sample i and ship sample j.
[0046] Then, the target ship with the highest similarity in the cosine similarity matrix is found for each ship in the to-be-optimized category, that is, for each sample that needs to be re-allocated, the ship category to which the other sample with the highest cosine similarity belongs is found. And the category label of each ship in the to-be-optimized category is updated to the ship category to which the corresponding target ship belongs; when all ships in the to-be-optimized category are reclassified, the to-be-optimized category is deleted, and thus the optimized ship classification result is obtained. For example, take 664 ships as the experimental data set, of which container ship A is 188, liquid bulk carrier B is 120, dry bulk carrier C is 300, and special transport ship D is 56. Then, special transport ship D is recorded as the to-be-optimized category, and a ship E in special transport ship D is taken. The similarity of ship E with A, B and C is calculated: sim(E, A) = 0.94 (container ship), sim(E, B) = 0.76 (liquid bulk carrier), and sim(E, C) = 0.85. The allocation result is: ship E→container ship. In this way, all ships in special transport ship D are reclassified, and the to-be-optimized category is deleted. Through the above reclassification process, the number of samples in the originally insufficient sample category can be significantly increased, so that each ship category contains enough samples. This reclassification enables the classification model to fully learn the characteristics of each type of ship in subsequent training, avoiding model bias and overfitting problems caused by insufficient sample quantity.
[0047] III. Model construction and training steps: The pre-processed multi-dimensional original feature data and the optimized ship classification result are taken as a data set, the data set is divided into a training set and a test set according to a predetermined proportion, the training set is input into a random forest model for training, a trained random forest model is obtained and tested by the test set for analysis, the model performance is verified, and a final trained and tested ship classification prediction model is obtained.
[0048] After the re-partitioning is completed, the application utilizes a random forest model to construct a ship classification prediction model, realizing scientific classification of ship types; by introducing a machine learning method, useful information can be automatically extracted from multi-dimensional ship features, effectively improving the accuracy and generalization ability of ship classification. Specifically, first, the original feature data and the optimized ship classification results are taken as a data set, and the data set is divided into a training set and a test set according to a preset ratio, which can be 70%-80% for the training set and 20%-30% for the test set, for example, divided according to a ratio of 70% training set and 30% test set, the training set is used for model training, and the test set is used for verifying the performance of the model, so that the model can learn on sufficient data and can be verified on unknown data sets. Then the training set is input into the random forest model for training, and the performance is optimized by adjusting the model parameters, and the trained random forest model is obtained and tested by the test set, and the final trained and tested ship classification prediction model is obtained. Preferably, the performance of the trained random forest model is evaluated using the test set through evaluation indicators, including detection speed, detection accuracy, and complexity, detection accuracy includes accuracy, recall rate, and average accuracy, detection speed uses the number of samples processed per second, and complexity uses the number of parameters and floating point operations. The calculation steps of the accuracy rate specifically include: inputting each ship data in the test set into the trained random forest model to output the prediction result; and based on the prediction result and the true label manually labeled for each ship data in the test set, the true positive, true negative, false positive and false negative are counted, and the accuracy rate is calculated according to the true positive, true negative, false positive and false negative, and the accuracy rate is used as the main evaluation indicator. The accuracy rate formula is as follows:
[0049]
[0050] In the above formula, TP, TN, FP, and TN represent true positive, true negative, false positive, and false negative, respectively.
[0051] The random forest model is an ensemble learning model based on decision trees, which realizes classification by training multiple independent decision trees. Each decision tree uses a random subset of samples and a random selection of features for training to reduce the risk of overfitting. Each decision tree learns the relationship between the features and the class labels of the samples during training, and after all decision trees are trained, the final prediction result is determined by ensemble voting. By integrating multiple decision trees, the stability and accuracy of the model can be significantly improved. In addition, the random forest model can exhibit good generalization ability when processing high-dimensional data, and has strong robustness to noise, preferably, the model is composed of 100 independent decision trees, and the maximum depth of each tree is set to 20.
[0052] IV. Model category prediction step: using the ship classification prediction model to predict the category of the new ship to be detected, and predicting the ship category to which the new ship belongs. That is, the trained and tested ship classification prediction model can be used to predict the category of the new ship. By inputting the multi-dimensional feature data of the new ship into the ship classification prediction model, the ship classification prediction model uses the previously learned knowledge to classify it and predict the ship category to which the ship belongs. The prediction result is compared with the true label as shown in Table 1.
[0053] Table 1
[0054] New ship Ship classification prediction model Container ship 0.9123 Liquid bulk carrier 0.8056 Dry bulk carrier 0.8222 Specialized cargo ship 0.9412
[0055] It can be seen that the method of the present application can significantly improve the accuracy and generalization ability of ship classification, and provide strong technical support for ship management, fuel consumption optimization and maritime big data analysis.
[0056] The present application also relates to a ship classification and prediction system based on cosine similarity optimization, which corresponds to the above-mentioned ship classification and prediction method based on cosine similarity optimization, and can be understood as a system for implementing the above-mentioned method. The system comprises a ship feature acquisition and preliminary classification module, a ship category optimization module, a model construction and training module, and a model category prediction module connected in sequence, specifically,
[0057] The ship feature acquisition and preliminary classification module acquires multi-dimensional original feature data of the ship, which covers the physical characteristics of the ship and performance parameters reflecting the power level and running efficiency of the ship, pre-processes the acquired multi-dimensional original feature data, and preliminarily classifies the ship based on the pre-processed multi-dimensional original feature data and combined with experience rules to obtain a preliminary classification result containing multiple ship categories. Then, the sample number of each ship category in the preliminary classification result is counted, and the ship category with a sample number less than a preset number threshold is recorded as a to-be-optimized category, and other ship categories except the to-be-optimized category are recorded as reference categories;
[0058] The ship category optimization module converts each feature parameter in the pre-processed multi-dimensional original feature data into a feature vector, calculates the cosine similarity between each ship in the to-be-optimized category and all ships in the reference categories based on the feature vector and using the cosine similarity method, and further generates a cosine similarity matrix. In the cosine similarity matrix, the target ship with the highest similarity to each ship in the to-be-optimized category is found out, and the category label of each ship in the to-be-optimized category is updated to the ship category to which the corresponding target ship belongs. When all ships in the to-be-optimized category are reclassified, the to-be-optimized category is deleted, and an optimized ship classification result is obtained;
[0059] The model construction and training module takes the preprocessed multi-dimensional original feature data and the optimized ship classification result as a data set, divides the data set into a training set and a test set according to a preset ratio, inputs the training set into a random forest model for training, obtains a trained random forest model, and tests and analyzes the test set to verify the model performance, and obtains a final trained and tested ship classification prediction model.
[0060] The model class prediction module uses the ship classification prediction model to predict the class of a new ship to be detected.
[0061] Preferably, in the ship feature acquisition and preliminary classification module, the preprocessing includes missing value processing and outlier removal processing on the multi-dimensional original feature data to obtain processed multi-dimensional feature data, and a standardization processor is used to standardize the processed multi-dimensional feature data to obtain standardized feature data.
[0062] Preferably, in the model construction and training module, the test set is used for test analysis to verify the model performance, specifically including: using the test set to evaluate the performance of the trained random forest model through evaluation indexes, the evaluation indexes including detection speed, detection accuracy and complexity, the detection accuracy including accuracy, recall rate and average accuracy, the detection speed using the number of samples processed per second, and the complexity using the number of parameters and floating point operation times.
[0063] Preferably, in the model construction and training module, when the test set is used for test analysis to verify the model performance, the calculation steps of the accuracy rate specifically include: inputting each ship data in the test set into the trained random forest model to output a prediction result; and based on the prediction result and the true label manually labeled for each ship data in the test set, counting true positives, true negatives, false positives and false negatives, and calculating the accuracy rate according to the true positives, true negatives, false positives and false negatives.
[0064] The present application provides an objective and scientific ship classification and prediction method and system based on cosine similarity optimization, which redivides the ship classes by using the cosine similarity method, thereby reassigning the ships in the to-be-optimized classes to more suitable classes, solving the problem of sample imbalance; and uses a random forest model for training and testing to ensure good generalization ability and noise resistance of the model on high-dimensional data, not only improving the accuracy of ship classification, but also realizing full-process automation from data acquisition to classification prediction, having good engineering practicability and expansibility, and showing wide application prospect and technical potential in ship management, fuel consumption optimization and maritime big data analysis, etc., and providing strong technical support for related industries.
[0065] It should be noted that the above detailed description of the specific embodiments of the present application is not intended to limit the present application in any way. Thus, while the present application has been described in detail with reference to specific embodiments thereof, it will be apparent to those skilled in the art that various modifications and changes can be made thereto without departing from the spirit and scope of the present application.
Claims
1. A ship classification and prediction method based on cosine similarity optimization, characterized in that: The following steps are involved: Ship feature collection and preliminary classification steps: collecting multi-dimensional original feature data of the ship covering the ship's physical characteristics and performance parameters reflecting the ship's power level and operating efficiency, pre-processing the collected multi-dimensional original feature data, and preliminarily classifying the ships based on the pre-processed multi-dimensional original feature data and in combination with empirical rules to obtain preliminary classification results including multiple ship categories; then counting the number of samples of each ship category in the preliminary classification results, recording the ship categories with a sample number less than a preset number threshold as categories to be optimized, and recording the other ship categories except the categories to be optimized as reference categories; Ship category optimization step: converting each feature parameter in the preprocessed multidimensional original feature data into a feature vector, and calculating the cosine similarity between each ship in the category to be optimized and all ships in the reference category using a cosine similarity method based on the feature vector, thereby generating a cosine similarity matrix; and finding the target ship with the highest similarity to each ship in the category to be optimized in the cosine similarity matrix, and updating the category label of each ship in the category to be optimized to the ship category to which the corresponding target ship belongs; When all ships in the category to be optimized are reclassified, the category to be optimized is deleted, and the optimized ship classification result is obtained; Model construction and training steps: The preprocessed multi-dimensional original feature data and the optimized ship classification results are used as the data set. The data set is divided into a training set and a test set according to a preset ratio. The training set is input into the random forest model for training. The trained random forest model is tested and analyzed on the test set to verify the model performance, and the final trained and tested ship classification prediction model is obtained; Model category prediction step: Use the ship classification prediction model to predict the category of the new ship to be detected, and predict the ship category to which the new ship belongs.
2. The ship classification and prediction method based on cosine similarity optimization according to claim 1 is characterized in that: In the ship feature collection and preliminary classification steps, the preprocessing includes performing missing value processing and outlier elimination processing on the multidimensional original feature data to obtain processed multidimensional feature data; and using a standardization processor to standardize the processed multidimensional feature data to obtain standardized feature data.
3. The ship classification and prediction method based on cosine similarity optimization according to claim 1 is characterized in that: In the model construction and training steps, the preset ratio is 70%-80% for the training set and 20%-30% for the test set; the random forest model is an ensemble learning model based on decision trees, consisting of 100 independent decision trees. When each decision tree is trained, a sample subset is generated by random sampling with replacement from the training set, and a feature subset is randomly selected during each node splitting process; at the same time, the maximum depth of each decision tree is set to 20; after all decision trees are trained, the final prediction result is determined by ensemble voting.
4. The ship classification and prediction method based on cosine similarity optimization according to claim 3 is characterized in that: During the model building and training steps, the test set is used for testing and analysis to verify the model performance, specifically including: The performance of the trained random forest model was evaluated using the test set using evaluation indicators. The evaluation indicators included detection speed, detection accuracy, and complexity. The detection accuracy included accuracy, recall, and average precision. The detection speed was measured by the number of samples processed per second, and the complexity was measured by the number of parameters and floating-point operations.
5. The ship classification and prediction method based on cosine similarity optimization according to claim 4 is characterized in that: The accuracy calculation step specifically includes: inputting each ship data in the test set into a trained random forest model and outputting a prediction result; and based on the prediction result and the real label manually annotated for each ship data in the test set, counting the true positive examples, true negative examples, false positive examples and false negative examples, and calculating the accuracy rate based on the true positive examples, true negative examples, false positive examples and false negative examples.
6. The ship classification and prediction method based on cosine similarity optimization according to claim 1 or 2, characterized in that: In the ship feature collection and preliminary classification step, the physical characteristics in the multidimensional original feature data include deadweight tonnage, gross tonnage of the ship, net tonnage of the ship, length, width, design draft, displacement, draft, main engine cylinder diameter and number of main engine cylinders; the performance parameters reflecting the ship's power level in the multidimensional original feature data include main engine stroke type, maximum power, braking power, operating power, main engine total horsepower and main engine total power; the performance parameters reflecting the ship's operating efficiency in the multidimensional original feature data include design speed, maximum speed and operating speed.
7. A ship classification and prediction system based on cosine similarity optimization, characterized in that: It includes the ship feature collection and preliminary classification module, the ship category optimization module, the model construction and training module and the model category prediction module, which are connected in sequence. The ship feature collection and preliminary classification module collects multi-dimensional original feature data of ships, including physical characteristics of the ships and performance parameters reflecting the ship's power level and operating efficiency, pre-processes the collected multi-dimensional original feature data, and preliminarily classifies the ships based on the pre-processed multi-dimensional original feature data and in combination with empirical rules to obtain preliminary classification results including multiple ship categories; then, the number of samples of each ship category in the preliminary classification results is counted, and the ship categories with a sample number less than a preset number threshold are recorded as categories to be optimized, and the other ship categories except the categories to be optimized are recorded as reference categories; The ship category optimization module converts each feature parameter in the preprocessed multidimensional original feature data into a feature vector, and calculates the cosine similarity between each ship in the category to be optimized and all ships in the reference category using a cosine similarity method based on the feature vector, thereby generating a cosine similarity matrix; and finding the target ship with the highest similarity to each ship in the category to be optimized in the cosine similarity matrix, and updating the category label of each ship in the category to be optimized to the ship category to which the corresponding target ship belongs; When all ships in the category to be optimized are reclassified, the category to be optimized is deleted, and the optimized ship classification result is obtained; The model building and training module uses the preprocessed multi-dimensional original feature data and the optimized ship classification results as a data set, divides the data set into a training set and a test set according to a preset ratio, inputs the training set into the random forest model for training, obtains the trained random forest model, and performs test analysis on the test set to verify the model performance, thereby obtaining the final trained and tested ship classification prediction model; The model category prediction module uses the ship classification prediction model to perform category prediction on a new ship to be detected, and predicts the ship category to which the new ship belongs.
8. The ship classification and prediction system based on cosine similarity optimization according to claim 7, characterized in that: In the ship feature collection and preliminary classification module, the preprocessing includes performing missing value processing and outlier elimination processing on the multidimensional original feature data to obtain processed multidimensional feature data; and using a standardization processor to standardize the processed multidimensional feature data to obtain standardized feature data.
9. The ship classification and prediction system based on cosine similarity optimization according to claim 7, characterized in that: In the model building and training module, the test set is used for testing and analysis to verify the model performance, specifically including: The performance of the trained random forest model was evaluated using the test set using evaluation indicators. The evaluation indicators included detection speed, detection accuracy, and complexity. The detection accuracy included accuracy, recall, and average precision. The detection speed was measured by the number of samples processed per second, and the complexity was measured by the number of parameters and floating-point operations.
10. The ship classification and prediction system based on cosine similarity optimization according to claim 9, characterized in that: In the model building and training module, when testing and analyzing the test set to verify the model performance, the accuracy calculation steps specifically include: inputting each ship data in the test set into the trained random forest model and outputting the prediction results; and based on the prediction results and the real labels manually annotated for each ship data in the test set, counting the true positive examples, true negative examples, false positive examples and false negative examples, and calculating the accuracy based on the true positive examples, true negative examples, false positive examples and false negative examples.
Citation Information
Patent Citations
Polarimetric SAR image classification method based on eigenvector measurement spectral clustering
CN104463219A
Text classification method and device
CN111858917A
Disease auxiliary diagnosis system and equipment based on oral acid, and storage medium
CN112185571A
Imbalanced marine ship target detection method and system based on virtual and real data mixed learning, medium and product
CN118823450A
Few-shot image recognition method and apparatus, device, and storage medium
US20240029397A1