Data processing method, data processing system, electronic equipment and readable storage medium
By collecting, cleaning, collecting and feature extraction, and building a database based on ontology models, the problem that the diversity and correlation of data source platforms in ship path inference in the existing technology is not effectively utilized, and the accuracy and efficiency of inference results are improved.
Patent Information
- Application Number
- CN202411915722.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing technology fails to effectively consider the diversity of data source platforms and the correlation in a single data source platform in ship path reasoning, resulting in low accuracy of inference results.
Provide a data processing method, by collecting multiple open source data, performing data cleaning and collection, extracting features, and constructing a database based on ontology models to perform ship path inference. This method fully combines business needs, data source platform and the correlation between the acquired data.
It improves the operation efficiency of data processing and the accuracy of results of ship path inference, ensuring the completeness, efficiency and accuracy of data.
Smart Images

Figure CN120045838A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ship path inference, and more specifically, to a data processing method, a data processing system, an electronic device, and a readable storage medium. Background Art
[0002] The data collected for ship path prediction involves a wide variety of data types and numerous data source platforms. In order to serve the inference of ship paths, it is also necessary to extract, correlate, and fuse the above diverse data, while ensuring the integrity, efficiency, and accuracy of the obtained data.
[0003] The data processing methods adopted in the related art do not consider the diversity of data source platforms and the correlation within a single data source platform, resulting in a relatively low accuracy of the inference results for ship paths. Summary of the Invention
[0004] In order to solve or improve at least one of the above technical problems, an object of the present invention is to provide a data processing method.
[0005] Another object of the present invention is to provide a data processing system.
[0006] Another object of the present invention is to provide an electronic device.
[0007] Another object of the present invention is to provide a readable storage medium.
[0008] To achieve the above object, a first aspect of the present invention provides a data processing method for ship path inference. The steps of the data processing method include:
[0009] First, collect multiple open-source data according to data requirements, and record the data source platform corresponding to each open-source data; wherein, the data requirements include at least one or a combination of the following: historical ship route data, hot event data, supply port data, and weather data.
[0010] Second, perform data cleaning on the open-source data to obtain cleaned result data.
[0011] Third, perform data aggregation on the cleaned result data according to the data source platform to obtain multiple aggregated data sets corresponding one-to-one to multiple data source platforms.
[0012] Fourth, perform feature extraction on the aggregated data sets to obtain a feature representation set.
[0013] In the fifth step, an ontology model is established based on the analysis requirements, and a database is constructed according to the ontology model and multiple feature representation sets for ship path reasoning. Among them, the analysis requirements include at least one of the following or a combination thereof: surface route analysis, underwater route analysis, and waypoint heat analysis.
[0014] The present invention aims to provide a data processing method that fully combines the relevance among business requirements (including data requirements and analysis requirements), data source platforms, and the acquired data. Whether in the process of feature extraction or in the process of database establishment, it can improve the operation efficiency of data processing and also improve the accuracy of the results of ship path reasoning.
[0015] In addition, the above technical solution provided by the present invention may also have the following additional technical features:
[0016] In some technical solutions, optionally, when the data requirements include historical ship route data, the data source platforms include a ship monitoring system and an underwater acoustic monitoring system; when the data requirements include hot event data, the data source platform includes a news platform; when the data requirements include supply port data, the data source platform includes a map service platform; when the data requirements include weather data, the data source platform includes a meteorological service platform.
[0017] In this technical solution, the historical ship route data is collected through the ship monitoring system and the underwater acoustic monitoring system. The hot event data is collected through the news platform. The supply port data is collected through the map service platform. The weather data is collected through the meteorological service platform.
[0018] By comprehensively analyzing the historical ship route data, considering the external impact brought by hot event data, combining the supply port data to plan reasonable docking points, and responding to meteorological changes based on the weather data, it is possible to more comprehensively and accurately reason and plan the future navigation path of the ship, improving the safety, efficiency, and economy of ship navigation, etc.
[0019] In some technical solutions, optionally, data cleaning is performed on the open-source data to obtain cleaned result data. The steps include: setting a configuration policy corresponding to the file of the open-source data and importing the open-source data through the configuration policy; determining the metadata of the data attributes corresponding to the open-source data, comparing the imported open-source data with the metadata to obtain a filtered open-source data set; performing abnormal data processing based on the filtered open-source data set to obtain the cleaned result data.
[0020] In this technical solution, the imported open-source data is compared with the metadata for the purpose of checking the integrity, accuracy, and consistency of the imported open-source data. After the comparison, a filtered open-source data set will be obtained, which includes overall data (data that meets the requirements), data that can be corrected (data with some attributes not meeting the requirements but can be corrected through certain methods), and data that cannot be processed (data that does not meet the requirements and cannot be corrected). The processing of abnormal data includes correcting the data that can be corrected and removing the data that cannot be processed.
[0021] The cleaned result data better meets the requirements of subsequent data processing steps, providing high-quality data support for ship path inference.
[0022] In some technical solutions, optionally, feature extraction is performed on the aggregated data set to obtain a set of feature representations, including: performing entity recognition on multiple aggregated data sets based on a large language model to obtain multiple entities; establishing a knowledge graph, calculating the similarity between the identified multiple entities and the entity categories in the knowledge graph and matching them to obtain a set of feature representations.
[0023] In this technical solution, by adopting a large language model, compared with traditional data processing models, the large language model can better understand and generate natural text, and can also exhibit certain logical thinking and reasoning abilities.
[0024] By accurately matching entities to appropriate categories, when performing relevant analyses such as ship path inference later, it is more convenient to utilize various relationships, attributes, etc. of the entities of this category already existing in the knowledge graph, thereby more efficiently mining the information in the data and providing more accurate and comprehensive support for ship path inference.
[0025] In some technical solutions, optionally, the set of feature representations includes at least one of the following or a combination thereof: semantic features, statistical features, temporal features, and geographic information features.
[0026] In this technical solution, when the number of the set of feature representations is one, the set of feature representations includes only any one of semantic features, statistical features, temporal features, and geographic information features. When the number of the set of feature representations is multiple, the set of feature representations includes any combination of semantic features, statistical features, temporal features, and geographic information features.
[0027] Optionally, semantic features include but are not limited to event type semantic features, severity semantic features, and impact scope semantic features.
[0028] Optionally, statistical features include but are not limited to mean, median, standard deviation, variance, extreme values, and quantiles.
[0029] Optionally, the time series features include trend features, seasonal features, and noise features.
[0030] Optionally, the geographical information features include location coordinates, distance calculation, and regional clustering.
[0031] In some technical solutions, optionally, an ontology model is established based on the analysis requirements, and a database is constructed according to the ontology model and multiple feature representation sets for ship path reasoning. The steps include: establishing an ontology model based on the analysis requirements; taking the ontology in the ontology model as the target ontology and the ontology in the feature representation set as the source ontology to determine the association relationship between the target ontology and the source ontology; constructing a database based on the target ontology, the source ontology, and the association relationship for ship path reasoning.
[0032] In this technical solution, after establishing the ontology model, a database is constructed according to the ontology model and multiple feature representation sets obtained through data collection and feature extraction. The database constructed in this way can effectively support the ship path reasoning work and provide data support for accurately and efficiently planning the ship navigation path.
[0033] In some technical solutions, optionally, the ontology model includes: ontology name, ontology type, and ontology description; and / or the association relationship includes: the relationship type between the target ontology and the source ontology, and the relationship direction between the target ontology and the source ontology.
[0034] In this technical solution, the ontology name can specifically be "ship", "port", "waterway", etc., which is convenient for quickly identifying and referring to specific ontologies.
[0035] The ontology type includes attribute code, attribute name, and whether it can be used as an extraction primary key.
[0036] The ontology description is used to explain and describe the ontology in detail, including the definition, connotation, extension, related rules and constraints, etc. of the ontology.
[0037] The relationship type is used to determine the nature of the connection between the target ontology and the source ontology. The relationship type includes but is not limited to inclusion relationship, causal relationship, and attribute relationship.
[0038] The relationship direction is used to indicate the direction of the association, that is, from the target ontology to the source ontology or from the source ontology to the target ontology. The relationship direction is crucial for constructing the database, which determines the sequence of reasoning and analysis when using the association relationship.
[0039] The second aspect of the present invention provides a data processing system, including a data acquisition module, a data cleaning module, a data collection module, a feature extraction module, and a data processing module.
[0040] The data acquisition module is used to acquire multiple open-source data according to data requirements and record the data source platforms corresponding to each open-source data. Among them, the data requirements include at least one of the following or a combination thereof: historical ship route data, hot event data, supply port data, and weather data.
[0041] The data cleaning module is used to clean the open-source data to obtain cleaned result data.
[0042] The data collection module is used to collect the cleaned result data according to the data source platforms to obtain multiple collection data sets corresponding one by one to the multiple data source platforms.
[0043] The feature extraction module is used to extract features from the collection data sets to obtain a feature representation set.
[0044] The data processing module is used to establish an ontology model based on analysis requirements, construct a database according to the ontology model and multiple feature representation sets for ship route reasoning. Among them, the analysis requirements include at least one of the following or a combination thereof: surface route analysis, underwater route analysis, and route point heat analysis.
[0045] The present invention aims to provide a data processing system that fully combines the relevance among business requirements (including data requirements and analysis requirements), data source platforms, and the acquired data. Whether in the process of feature extraction or in the process of database establishment, it can improve the operation efficiency of data processing and also improve the accuracy of the results of ship route reasoning.
[0046] A third aspect of the present invention provides an electronic device, including a memory and a processor. Among them, a program or instruction that can run on the processor is stored on the memory, and when the processor executes the program or instruction, the steps of the data processing method in any of the above technical solutions are implemented.
[0047] A fourth aspect of the present invention provides a readable storage medium, which stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the data processing method in any of the above technical solutions are implemented.
[0048] The additional aspects and advantages of the technical solution of the present invention will become obvious in the following description part or be understood through the practice of the present invention. Description of the Drawings
[0049] Figure 1 Shows a flowchart of a data processing method according to an embodiment of the present invention;
[0050] Figure 2 Shows a flowchart of a data processing method according to another embodiment of the present invention;
[0051] Figure 3 The flowchart of a data processing method according to another embodiment of the present invention is shown;
[0052] Figure 4 The flowchart of a data processing method according to another embodiment of the present invention is shown;
[0053] Figure 5 The block diagram of the structure of a data processing system according to an embodiment of the present invention is shown;
[0054] Figure 6 The block diagram of the structure of an electronic device according to an embodiment of the present invention is shown;
[0055] Figure 7 The flowchart of a data processing method according to another embodiment of the present invention is shown;
[0056] Figure 8 The flowchart of a data processing method according to another embodiment of the present invention is shown;
[0057] Figure 9 The flowchart of a data processing method according to another embodiment of the present invention is shown.
[0058] Wherein, Figure 5 and Figure 6 The corresponding relationship between the reference numerals and the component names in is:
[0059] 200: data processing system; 210: data acquisition module; 220: data cleaning module; 230: data collection module; 240: feature extraction module; 250: data processing module; 300: electronic device; 310: memory; 320: processor. Detailed implementation manners
[0060] In order to be able to more clearly understand the above objects, features and advantages of the embodiments of the present invention, the embodiments of the present invention will be further described in detail below with reference to the drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.
[0061] Many specific details are set forth in the following description in order to fully understand the present application. However, the embodiments of the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present application is not limited to the limitations of the specific embodiments disclosed below.
[0062] The following refers to Figures 1 to 9 Describe a data processing method, a data processing system, an electronic device and a readable storage medium provided according to some embodiments of the present invention.
[0063] In an embodiment according to the present invention, the data processing method is used for ship route inference.
[0064] As Figure 1 shown, the steps of the data processing method include:
[0065] S102, collect multiple open-source data according to data requirements, and record the data source platform corresponding to each open-source data; wherein, the data requirements include at least one of the following or a combination thereof: ship route historical data, hot event data, supply port data, and weather data.
[0066] Open-source data refers to data that can be publicly obtained, used, and shared. In the data processing method for ship route inference, these open-source data come from multiple data source platforms and are data resources that can be utilized. Open-source data is the data basis for ship route inference.
[0067] The data requirements include at least one of the following or a combination thereof: ship route historical data, hot event data, supply port data, and weather data. In other words, when the number of data requirements is one, the data requirement is any one of ship route historical data, hot event data, supply port data, and weather data. When the number of data requirements is multiple, the data requirements are any combination of ship route historical data, hot event data, supply port data, and weather data.
[0068] Regarding the ship route historical data, this part of the data is crucial for understanding the ship's past travel routes, habitual routes, etc. By analyzing the ship route historical data, the navigation trajectory rules of the ship at different time periods and under different conditions can be summarized, providing a basic reference for subsequent route inference. Its sources are mainly ship monitoring systems and underwater acoustic monitoring systems, which can record relevant information such as the ship's navigation position in real time or regularly, thus forming historical data for analysis.
[0069] Regarding the hot event data, these data mainly come from news platforms. Ship navigation may be affected by various hot events occurring around the waterway, such as sudden natural disasters in the waterway, waterway control caused by maritime accidents, major events at the port, etc. Understanding these hot events helps to consider the impact of external sudden factors on the ship's route selection during ship route inference, so as to more reasonably plan future routes.
[0070] For the replenishment port data, this data is collected by the map service platform. The replenishment port is an important node for ships to conduct activities such as material replenishment and personnel rest during navigation. Mastering relevant data such as the location, facility configuration, and operation status of the replenishment port can reasonably plan the port for docking replenishment and the subsequent navigation route based on the current state of the ship (such as the remaining fuel, material reserves, etc.) during ship route inference, ensuring the continuity and safety of ship navigation.
[0071] For weather data, this data is collected by the meteorological service platform. Weather conditions have a direct and significant impact on ship navigation safety and route selection. Severe weather such as typhoons, heavy rains, and heavy fog may force ships to change their routes to avoid danger, while under good weather conditions, ships may also choose a more optimal direct route, etc. Therefore, weather data is one of the key factors that must be considered in ship route inference, which can help accurately predict the feasible routes of ships under different weather conditions.
[0072] In a specific embodiment, optionally, the data requirements include ship route historical data, hot event data, replenishment port data, and weather data. These different types of data requirements together constitute the data basis required for ship route inference. By comprehensively analyzing ship route historical data, considering the external impacts brought by hot event data, combining replenishment port data to plan reasonable docking points, and responding to meteorological changes based on weather data, it is possible to more comprehensively and accurately infer and plan the future navigation route of the ship, improving the safety, efficiency, and economy of ship navigation, etc.
[0073] In the process of collecting multiple open-source data in the present invention, the multi-source nature of collecting open-source data and the correlation between open-source data are reflected.
[0074] Multi-source nature: All types of open-source data involved in the data requirements come from different platforms (data source platforms), such as the ship's own monitoring and acoustic monitoring systems (ship monitoring system and underwater acoustic monitoring system), news platforms, map service platforms, meteorological service platforms, etc., fully reflecting the diversity of data sources to obtain information related to ship navigation from multiple perspectives.
[0075] Correlation: Although each type of open-source data seems independent, they are interrelated during the ship route inference process. For example, hot events may affect the route selection of ships to replenishment ports, and weather data will also affect the route adjustment of ships based on route historical data, etc., working together to achieve accurate route inference.
[0076] S104, perform data cleaning on the open-source data to obtain the cleaned result data.
[0077] Data cleaning is an important step in the data processing process, and its purpose is to improve data quality. For open source data, due to its wide and complex sources, there may be various problems, such as incomplete data, inconsistent data formats, data errors, duplicate data, etc. These problems can be solved through data cleaning, thereby obtaining cleaner and more usable cleaning result data, providing a reliable foundation for subsequent analysis and processing.
[0078] The cleaned result data is more in line with the requirements of subsequent data processing steps and provides high-quality data support for ship path reasoning.
[0079] S106, performing data aggregation on the cleaning result data according to the data source platform, and obtaining a plurality of aggregated data sets corresponding to the plurality of data source platforms one by one.
[0080] Data collection is to classify and organize the cleaned data (cleaning result data) according to its data source platform. The purpose is to maintain the original relevance and traceability of the data, facilitate subsequent targeted analysis and processing according to the characteristics of data from different platforms, and also facilitate the separate evaluation and management of the data quality of each source platform.
[0081] Specifically, the cleaning result data from the ship monitoring system are aggregated into a set to form a first aggregated data set; the cleaning result data from the underwater acoustic monitoring system are aggregated into a set to form a second aggregated data set; the cleaning result data from the news platform are aggregated into a set to form a third aggregated data set; the cleaning result data from the map service platform are aggregated into a set to form a fourth aggregated data set; the cleaning result data from the meteorological service platform are aggregated into a set to form a fifth aggregated data set.
[0082] There is a one-to-one correspondence between these collected data sets and the data source platforms. This correspondence makes the management and use of data more orderly.
[0083] S108, extracting features from the collected data set to obtain a feature representation set.
[0084] There are multiple feature representation sets, and the multiple feature representation sets correspond one-to-one to the multiple imputed data sets.
[0085] Feature extraction is a key step in mining valuable information from large amounts of data. For the collection of data sets, its purpose is to transform the original data into a more representative and analytically valuable form, namely, a feature representation set. These features can better describe the essential attributes of the data, provide more concise and effective information for subsequent analysis work such as ship path reasoning, and help improve the accuracy and efficiency of the model.
[0086] Each aggregated data set is processed by a specific method to extract the key features therein. These features may cover multiple dimensions, such as statistical features, semantic features, temporal features, and geographic information features, etc. After extraction, the obtained feature representation set will reflect the important information in the original data in a new and more compact data structure.
[0087] S110, establish an ontology model based on the analysis requirements, and construct a database according to the ontology model and multiple feature representation sets for ship path reasoning; wherein, the analysis requirements include at least one of the following or a combination thereof: surface route analysis, underwater route analysis, and waypoint heat analysis.
[0088] In the case where the number of analysis requirements is one, the analysis requirements only include any one of surface route analysis, underwater route analysis, and waypoint heat analysis. In the case where the number of analysis requirements is multiple, the analysis requirements include any combination of surface route analysis, underwater route analysis, and waypoint heat analysis.
[0089] The content of surface route analysis includes but is not limited to the width range of the waterway, water depth conditions, water flow speed and direction, and surface navigation rules.
[0090] Specifically, for different types of surface waterways, it is necessary to analyze their physical characteristics in detail, including the width range of the waterway. From narrow inland waterways to wide sea lanes, different widths will limit the passage of ships of different sizes; water depth conditions are equally crucial. Shallow water areas may only allow ships with a shallow draft to pass, while deep water areas have less restrictions on the draft of ships. These water depth data will affect the safe navigation path of ships; water flow speed and direction are also factors that cannot be ignored. Sailing downstream and upstream will greatly affect the ship's speed and fuel consumption, and thus affect path selection. In addition, surface navigation rules are also one of the contents of surface route analysis, such as traffic control rules in different waters, the meaning of waterway markings, etc. These rules will restrict the navigation behavior of ships on the water surface.
[0091] Underwater route analysis mainly focuses on the impact of the underwater environment on the ship's navigation path. The content of underwater route analysis includes but is not limited to the distribution of reefs and the depth and location information of trenches.
[0092] The complexity of the underwater terrain is an important factor. For example, the distribution of reefs, and ships need to avoid these potential dangerous areas; the depth and location information of trenches are crucial for the path planning of some special operation ships (such as diving operation ships). At the same time, underwater obstacles, whether naturally formed or man-made remnants (such as sunken ships), need to be considered in path planning to ensure the safety of ships' underwater navigation.
[0093] Waypoint thermal analysis mainly analyzes the attributes of each waypoint, such as the busyness and importance. By collecting and analyzing a large amount of ship navigation data, it is possible to determine which waypoints are hot spots where ships frequently pass through, which may be important port entrances and exits, waterway intersections, etc. Understanding the thermal conditions of waypoints helps to reasonably select waypoints in ship path reasoning, improve navigation efficiency, and avoid unnecessary congestion.
[0094] In a specific embodiment, optionally, the analysis requirements include surface route analysis, underwater route analysis, and waypoint thermal analysis. Based on these analysis requirements, an ontology model is established. The ontology model is an abstract representation of the knowledge in the field of ship path reasoning. After the ontology model is established, a database is constructed based on the ontology model and multiple feature representation sets obtained through data collection and feature extraction. The database constructed in this way can effectively support the ship path reasoning work and provide data support for accurate and efficient planning of ship navigation paths.
[0095] The present invention aims to provide a data processing method that fully combines the business needs (including data needs and analysis needs), the data source platform and the correlation between the acquired data. Whether in the process of feature extraction or in the process of database establishment, it will improve the operating efficiency of data processing and can also improve the accuracy of the results of reasoning about the ship path.
[0096] Specifically, the data processing method provided by the present invention combines the business needs, data source platforms and acquired data, and the correlation between multiple parties. In the process of data feature extraction and database establishment, it can improve efficiency and make data storage faster. It should be emphasized that in the process of database establishment, an ontology model is established in combination with the needs of ship path reasoning (an ontology model is established based on multiple analysis needs), so that the subsequent data retrieval is more timely and accurate. This data processing method can ensure that when a part of the collected open source data fluctuates, it will not have a correlation effect on the open source data of other data source platforms, which is conducive to improving the fault tolerance of the system.
[0097] In some embodiments, optionally, when the data demand includes historical ship route data, the data source platform includes a ship monitoring system and an underwater acoustic monitoring system.
[0098] The ship's route historical data is collected through the ship monitoring system and underwater acoustic monitoring system.
[0099] By analyzing the historical data of ship routes, we can summarize the navigation trajectory patterns of ships in different time periods and under different conditions, providing a basic reference for subsequent path reasoning.
[0100] The ship monitoring system is installed on the hull of the ship or at monitoring stations along the shore, and is equipped with various types of sensors. Among them, the Global Positioning System (GPS) sensor can accurately obtain the real-time geographical location coordinates of the ship. As the ship sails, the position information is recorded at regular time intervals. The continuous arrangement of these position data points in the time dimension constitutes the basic framework of the ship's navigation track. At the same time, the speed sensor and the heading sensor in the ship monitoring system work together. The speed sensor measures the moving speed of the ship relative to the surrounding water body, and the heading sensor determines the direction of the ship's travel using the geomagnetic or gyro principle. These data are correlated with the position data, which can not only accurately depict the ship's navigation state at different times, but also reflect the changes in the ship's traveling speed in the waterway and navigation actions such as turning and turning around.
[0101] The underwater acoustic monitoring system utilizes the characteristics of underwater acoustic propagation. Acoustic sensors are deployed around the ship or at specific positions on the seabed, which can transmit and receive underwater sound waves. When the sound waves generated by the ship's underwater navigation or the rotation of its propeller are received by these sensors, after signal processing and analysis, various information about the ship underwater can be obtained. In addition, the underwater acoustic monitoring system can also detect the distances between the ship and the surrounding underwater terrain (such as underwater mountains, canyons, etc.) or potential obstacles (such as sunken ships, reefs, etc.) through the principle of acoustic wave reflection. These distance information are crucial for understanding the environmental conditions of the ship's underwater navigation, and combined with the ship's navigation data on the water surface, comprehensively enrich the connotation of the ship's route historical data.
[0102] In some embodiments, optionally, when the data requirement includes hot event data, the data source platform includes a news platform.
[0103] The hot event data is collected through the news platform.
[0104] Ship navigation may be affected by various hot events occurring around the waterway, such as sudden natural disasters in the waterway, waterway control caused by maritime accidents, major events at the port, etc. Understanding these hot events helps to consider the impact of external sudden factors on the ship's route selection during ship path reasoning, so as to more reasonably plan the future path.
[0105] The news platform includes but is not limited to news media websites, shipping news platforms, and news channels of social media platforms.
[0106] In some embodiments, optionally, when the data requirement includes replenishment port data, the data source platform includes a map service platform.
[0107] The replenishment port data is collected through the map service platform.
[0108] Mastering relevant data such as the location, facility configuration, and operation status of supply ports (supply port data) enables reasonable planning of the ports for refueling and subsequent sailing routes based on the current status of the ship (such as remaining fuel, material reserves, etc.) during ship route reasoning, ensuring the continuity and safety of ship navigation.
[0109] The map service platform includes, but is not limited to, a general map platform and a professional shipping map platform.
[0110] In some embodiments, optionally, when the data requirements include weather data, the data source platform includes a meteorological service platform.
[0111] Weather data is collected through the meteorological service platform.
[0112] Weather conditions have a direct and significant impact on ship navigation safety and route selection. Severe weather such as typhoons, heavy rains, and fog may force ships to change their routes to avoid danger, while under good weather conditions, ships may also choose a more optimal direct route, etc. Therefore, weather data is one of the key factors that must be considered in ship route reasoning, which can help accurately predict the feasible routes of ships under different weather conditions.
[0113] The meteorological service platform is a system built by institutions that professionally collect, analyze, and release meteorological information, and has a high degree of professionalism and authority.
[0114] By comprehensively analyzing the historical data of ship routes, considering the external impacts brought by hotspot event data, planning reasonable docking points in combination with supply port data, and responding to meteorological changes based on weather data, it is possible to reason and plan the future sailing routes of ships more comprehensively and accurately, improving the safety, efficiency, and economy of ship navigation, etc.
[0115] In the process of collecting multiple open-source data in the present invention, the multi-source nature of collecting open-source data and the relevance between open-source data are reflected.
[0116] In some embodiments, optionally, as Figure 2 shown, for S104 (performing data cleaning on the open-source data to obtain cleaned result data), the steps include:
[0117] S1042, setting a configuration policy corresponding to the file of the open-source data and importing the open-source data through the configuration policy.
[0118] Open-source data from different sources often differ in file format, data structure, storage method, etc. Therefore, it is necessary to formulate corresponding configuration policies according to the file characteristics of these specific open-source data.
[0119] The configuration policy can be understood as the import rules and parameter settings for specific open-source data files. The configuration policy includes, but is not limited to, the delimiter of the data, the encoding method, and the mapping of data types.
[0120] In the case where the open-source data is in text format and has a delimiter, it is necessary to determine the delimiter of the data. For example, a CSV file (a file format) commonly uses a comma as the delimiter.
[0121] Optionally, the UTF-8 encoding (an encoding method) is adopted to ensure the correct interpretation of characters.
[0122] Regarding the mapping of data types, numeric fields in text form can be correctly recognized as numeric types for subsequent calculation and processing.
[0123] After setting the configuration policy for a specific open-source data file, the data import operation is performed according to these policies. This means that the system will read the open-source data from its original storage location and convert it into a format that can be used for subsequent operations in the current data processing environment based on the previously set rules such as delimiter, encoding method, and data type mapping. For example, if the configuration policy stipulates that the encoding method of a certain open-source data file is UTF-8 and the delimiter is a comma, then during import, the system will decode and split the file according to these requirements to correctly import the data into the corresponding data structures (such as data tables, data matrices, etc.), thus preparing for subsequent data cleaning, analysis, and other steps.
[0124] S1044, determine the metadata of the data attributes corresponding to the open-source data, compare the imported open-source data with the metadata, and obtain a filtered set of open-source data.
[0125] Metadata is closely related to open-source data. It is an abstract description of the characteristics, sources, quality, etc. of open-source data. Each set of open-source data has corresponding metadata to characterize its own properties.
[0126] Comparing the imported open-source data with the metadata aims to check the integrity, accuracy, and consistency of the imported open-source data. By comparing with the metadata, it can be found whether the data meets the expected attribute requirements, such as whether the data range is correct and whether the data format is consistent with the specified one.
[0127] In a specific embodiment, optionally, the filtered set of open-source data includes overall data, data that can be corrected, and data that cannot be processed.
[0128] Among them, the overall data are data that meet the requirements; the data that can be corrected are data whose partial attributes do not meet the requirements but can be corrected by certain methods; the data that cannot be processed are data that do not meet the requirements and cannot be corrected. This provides a basis for subsequent abnormal data processing to ensure the quality of the data ultimately used for ship route inference.
[0129] S1046, perform abnormal data processing on the screened open-source data set to obtain the cleaned result data.
[0130] Abnormal data processing includes correcting the data that can be corrected and removing the data that cannot be processed.
[0131] The cleaned result data better meet the requirements of subsequent data processing steps and provide high-quality data support for ship route inference.
[0132] In some embodiments, optionally, as Figure 3 shown, S108 (perform feature extraction on the aggregated data set to obtain the feature representation set), the steps include:
[0133] S1081, receive multiple aggregated data sets.
[0134] There is a one-to-one correspondence between the multiple aggregated data sets and the multiple data source platforms, and this correspondence makes the management and use of data more orderly.
[0135] S1082, perform entity recognition on the multiple aggregated data sets based on the large language model to obtain multiple entities.
[0136] The large language model (LLM model, Large Language Model) is a model based on machine learning and natural language processing technologies.
[0137] The large language model is trained on a large amount of text data, thereby learning and mastering the ability to serve human language understanding and generation. Its core idea is to learn the patterns and language structures of natural language through large-scale unsupervised training, and then to a certain extent simulate the human language cognition and generation process.
[0138] By adopting the large language model, compared with traditional data processing models, the large language model can better understand and generate natural text, and can also show certain logical thinking and reasoning abilities.
[0139] An entity refers to a specific thing, concept, person, place, etc. that has an independent meaning of existence. For example, in the scenario of ship route inference, the ship itself, ports, waterways, etc. can all be regarded as entities.
[0140] The purpose of this step is to identify specific relevant things or concepts (entities) from the aggregated data set through a large language model.
[0141] S1084, build a knowledge graph, calculate the similarity between the identified multiple entities and the entity categories in the knowledge graph and match them to obtain a set of feature representations.
[0142] A knowledge graph is a structured graph composed of nodes and edges. Nodes usually represent entities, and edges represent the relationships between entities. For example, in a knowledge graph about shipping, there may be an edge with a "docking" relationship between the "ship" node and the "port" node. The entity categories in the knowledge graph are obtained by classifying numerous entities according to certain attributes, features, functions, etc., such as different categories like ship category, port category, meteorological category, etc.
[0143] There are various methods for similarity calculation. Common ones such as the method based on the vector space model will convert both entities and entity categories into vector representations, and then measure the similarity between them by calculating metrics such as the cosine value of the angle between vectors; there is also the feature-based method, that is, comparing the various features possessed by entities and entity categories, and counting the number or proportion of the same features to determine the similarity.
[0144] By accurately matching entities to appropriate categories, in subsequent related analyses such as ship route reasoning, it is more convenient to utilize various relationships, attributes, etc. of entities of this category already existing in the knowledge graph, thereby more efficiently mining information in the data and providing more accurate and comprehensive support for ship route reasoning.
[0145] In some embodiments, optionally, the set of feature representations includes at least one of the following or a combination thereof: semantic features, statistical features, temporal features, and geographical information features.
[0146] When the number of the set of feature representations is one, the set of feature representations includes only any one of semantic features, statistical features, temporal features, and geographical information features. When the number of the set of feature representations is multiple, the set of feature representations includes any combination of semantic features, statistical features, temporal features, and geographical information features.
[0147] Semantic features refer to various characteristics that can reflect the meaning of text content. It goes beyond the surface form of the text (such as words, characters themselves) and delves into the semantic-level information such as concepts, relationships, and intentions conveyed by the text.
[0148] Taking natural disaster events as an example, if the semantic features in news texts indicate that a typhoon has occurred on the scheduled route of a ship, by analyzing semantic information such as the intensity (one of the semantic features) and moving path of the typhoon, the risk level faced by the ship can be judged. If the text also contains semantic features of the affected area, such as a certain port being affected by the typhoon and suspending operations, the ship can adjust its route in advance to avoid the affected port and the sea area affected by the typhoon.
[0149] Optionally, the semantic features include but are not limited to event type semantic features, severity semantic features, and impact range semantic features. For example, a news report mentions pirate activities in a certain sea area (event type semantic feature), and through text analysis, semantic information such as the frequency of the activities (severity semantic feature) and the sea area involved (impact range semantic feature) is obtained. Ship operators can decide whether to change the route and choose a safer waterway based on these semantic features.
[0150] Optionally, the statistical features include but are not limited to mean, median, standard deviation, variance, extreme values, and quantiles.
[0151] The mean is the arithmetic average of the data and can reflect the average level of the data. In the data related to ship route inference, for example, the mean of the ship speed can reflect the approximate traveling speed of the ship within a certain period of time and is a basic description of the overall speed situation.
[0152] The median is the value at the middle position after sorting the data from smallest to largest (if the number of data is odd) or the average of the two middle numbers (if the number of data is even). It can better reflect the central position of the data set with outliers. For example, in the data of the cargo throughput of a supply port, if there are individual extremely large values (possibly due to the transportation of special goods, etc.), the median can more stably represent the general throughput level.
[0153] The standard deviation and variance are used to measure the dispersion degree of the data. The standard deviation is the square root of the variance. A larger standard deviation or variance indicates that the data is more dispersed. In ship-related data, for example, in meteorological data, a large standard deviation of temperature indicates that the temperature fluctuates violently, which may have different impacts on ship equipment and navigation conditions.
[0154] The extreme values include the maximum and minimum values, which define the range of the data. The maximum and minimum values of the ship's cargo capacity can show the boundaries of the ship's cargo capacity and are helpful for understanding the limits of the ship's transportation capacity.
[0155] Quantiles divide the data according to a certain proportion, which can observe the data distribution more carefully.
[0156] Optionally, the time series features include trend features, seasonal features, and noise features.
[0157] Trend features are used to reflect the overall direction of data changes over time and represent a long-term and relatively stable change trend. In data related to ship path inference, trend features are of great significance for prediction and planning. For example, the trend of a ship's sailing speed over a long period can help determine changes in ship performance. If the speed shows a gradually decreasing trend, it may imply problems with the ship's power system or an increase in hull resistance, etc.
[0158] Seasonal features are used to reflect the regular changes presented by data within a specific time cycle. For ship-related data, such seasonal features are very common. For example, the wind and wave conditions in certain sea areas may be seasonal, with larger waves in specific seasons and relatively calm in other seasons, which has a direct impact on ship navigation safety and route selection.
[0159] Noise features represent the irregular and random fluctuation parts in the data. In ship data, noise may be caused by various factors. For example, sudden meteorological changes (such as short-term severe convective weather), temporary waterway control, and some minor malfunctions of the ship itself may all cause noise in ship navigation data (such as speed, heading, etc.). Identifying noise features helps to eliminate abnormal interference, analyze the essential laws of data more accurately, and also remain vigilant about sudden situations that may affect ship navigation.
[0160] Optionally, geographical information features include position coordinates, distance calculation, and regional clustering.
[0161] Position coordinates are the key information for determining the position of geographical entities on the earth's surface and usually adopt coordinate systems such as longitude and latitude. In ship path inference, the position coordinates of entities such as ships, ports, and waterways are basic data. For example, knowing the current position coordinates of a ship and the position coordinates of the target port is a basic prerequisite for planning the ship's navigation path.
[0162] Distance calculation is used to measure the spatial interval between geographical entities. In the shipping field, distance calculation has various uses. For example, calculating the distance between a ship and a port can help evaluate the time and fuel consumption required for the ship to reach the port; calculating the distance between a ship and obstacles or other ships on the waterway is crucial for avoiding collisions and ensuring navigation safety.
[0163] Regional clustering is a method of classifying regions with similar geographical features or attributes. In ship route analysis, regional clustering can be used to classify sea areas, waterways, etc. For example, sea areas can be clustered according to factors such as marine environment (such as water temperature, salinity, etc.), shipping density, and geographical features, and similar sea areas are divided into the same type of region. This can help ship operators understand the characteristics of different regions, select appropriate regions to pass through according to the needs and capabilities of the ship when planning routes, and also contribute to analyzing the navigation rules and risks of ships in different types of regions.
[0164] In some embodiments, optionally, as Figure 4 shown, S110 (establish an ontology model based on analysis requirements, construct a database according to the ontology model and multiple feature representation sets for ship route reasoning), the steps include:
[0165] S1102, establish an ontology model based on analysis requirements.
[0166] Analysis requirements include surface route analysis, underwater route analysis, and waypoint heat analysis. Based on these analysis requirements, an ontology model is established. The ontology model is a tool for structuring the representation of knowledge in a specific domain.
[0167] Optionally, the ontology model includes multiple ontologies. In the ontology model for ship route reasoning, ontologies may include various concepts related to the ship's navigation route, such as the ship itself, ports, waterways, and meteorological conditions.
[0168] For example, the "ship" ontology may contain attributes such as the type, size, speed, and load capacity of the ship; the "port" ontology may cover information such as the location, throughput capacity, and service facilities of the port; the "waterway" ontology may involve features such as the width, depth, and water flow speed of the waterway; the "meteorological conditions" ontology may contain meteorological elements such as wind speed, wind direction, and visibility. These ontologies and their attributes are connected to each other through certain relationships to form a complete ontology model, so as to better meet the analysis requirements of ship route reasoning.
[0169] S1104, take the ontologies in the ontology model as target ontologies, take the ontologies in the feature representation set as source ontologies, and determine the association relationships between the target ontologies and the source ontologies.
[0170] Determining the association relationship between the target ontology and the source ontology is the key to realizing knowledge fusion and data-driven decision-making. This association relationship can complement the database designed based on the ontology model with the information obtained through data feature extraction. For example, if a certain ship feature in the source ontology is associated with the ship type in the target ontology, then when performing ship route reasoning, this association can be used to determine the ship type according to the ship's features, and then use the knowledge such as the navigation rules of this type of ship in the ontology model for route planning.
[0171] S1106, construct a database based on the target ontology, the source ontology, and the association relationship for ship route reasoning.
[0172] As the core part of the ontology model, the target ontology provides a high-level conceptual framework and semantic structure for the database. In the context of ship route reasoning, the target ontology includes but is not limited to ship types, channel categories, and navigation rules. These target ontologies enable the database to be designed according to the logic in the field of ship route reasoning, storing and managing different types of ship information, channel information, navigation rule information, etc. separately, facilitating subsequent queries and reasoning.
[0173] The source ontology reflects the information contained in the features extracted from actual data, such as the ontology corresponding to the features extracted from ship route historical data, hot event data, supply port data, and weather data.
[0174] The association relationship is the bridge connecting the target ontology and the source ontology. It clarifies the correspondence and connection between ontologies from different sources and levels, making the data in the database coherent and consistent. For example, the association relationship can determine that a certain ship type (target ontology) is associated with specific feature such as speed range and equipment configuration (source ontology), so that these related information can be accurately stored and retrieved in the database, providing a basis for obtaining relevant feature data according to the ship type during ship route reasoning.
[0175] In some embodiments, optionally, the ontology model includes: ontology name, ontology type, and ontology description.
[0176] The ontology name can specifically be "ship", "port", "channel", etc., which facilitates the quick identification and reference to specific ontologies.
[0177] The ontology type includes attribute code, attribute name, and whether it can be used as an extraction primary key.
[0178] The attribute code is an encoded representation of the ontology attribute, with uniqueness and standardization. Through the attribute code, relevant attribute data can be quickly located and operated on.
[0179] The attribute name is an intuitive description of the ontology attribute, such as "ship speed", "channel water depth", etc.
[0180] If an attribute is marked as extractable primary key, it means that when extracting ontology-related information from a large amount of data, this attribute can be used as the main basis for identification and differentiation. For example, in the data processing of ship path reasoning, if the "International Maritime Organization number (IMO number)" of a ship is determined to be an extractable primary key, then when collecting ship information from different data sources, the relevant data of each ship can be accurately identified and integrated through this unique number, avoiding data confusion and incorrect association.
[0181] The ontology description is used to provide a detailed explanation and illustration of the ontology, including the definition, connotation, extension, relevant rules and constraints of the ontology, etc.
[0182] In some embodiments, optionally, the association relationship includes: the relationship type between the target ontology and the source ontology, and the relationship direction between the target ontology and the source ontology.
[0183] The relationship type is used to determine the nature of the connection between the target ontology and the source ontology. The relationship type includes but is not limited to inclusion relationship, causal relationship and attribute relationship.
[0184] Regarding the inclusion relationship, for example: the target ontology "ship" includes the source ontology "ship power system".
[0185] Regarding the causal relationship, for example: there is a causal connection between the target ontology "severe weather" and the source ontology "change in ship sailing speed".
[0186] Regarding the attribute relationship, for example: the target ontology "port" and the source ontology "port throughput capacity".
[0187] The relationship direction is used to indicate the direction of the association, that is, from the target ontology to the source ontology, or from the source ontology to the target ontology. The relationship direction is crucial for building a database, and it determines the sequence when using the association relationship for reasoning and analysis.
[0188] In some embodiments, optionally, the association relationship further includes: the names of the target ontology and the source ontology.
[0189] In some embodiments, optionally, as Figure 7 shown, the steps of the data processing method include:
[0190] S1, collect open-source data based on multiple open-source data platforms according to data requirements, and save the source platform corresponding to each open-source data.
[0191] S2, perform data cleaning on the open-source data to obtain the cleaned result data.
[0192] S3. Aggregate the cleaned result data according to the source platforms to obtain multiple aggregated data sets corresponding to multiple source platforms respectively.
[0193] S4. Extract features from multiple aggregated data sets respectively to obtain multiple feature representation sets corresponding to multiple aggregated data sets.
[0194] S5. Establish analysis requirements including surface route analysis, underwater route analysis and route point heat analysis, establish an ontology model based on the analysis requirements, and construct a database according to the ontology model and multiple feature representation sets to complete data processing.
[0195] In some embodiments, optionally, as Figure 8 shown, S2 (perform data cleaning on the open-source data to obtain the cleaned result data), the steps include:
[0196] S21. Set the configuration policy according to the file correspondence of the open-source data, and import the open-source data through the configuration policy.
[0197] S22. Compare the imported open-source data with the metadata of the data attributes corresponding to the open-source data to obtain a filtered open-source data set, and the filtered open-source data set includes overall data, correctable data and unprocessable data.
[0198] S23. Perform abnormal data processing based on the filtered open-source data to obtain the cleaned result data.
[0199] In some embodiments, optionally, as Figure 9 shown, S4 (extract features from multiple aggregated data sets respectively to obtain multiple feature representation sets corresponding to multiple aggregated data sets), the steps include:
[0200] S41. Receive multiple aggregated data sets.
[0201] S42. Perform entity recognition on multiple aggregated data sets through the LLM model to obtain multiple entities.
[0202] Among them, the LLM model (Large Language Model) is a large language model, which is a model based on machine learning and natural language processing technologies.
[0203] S43. Establish a knowledge graph, calculate the similarity between the identified multiple entities and the entity categories in the knowledge graph and match them to obtain multiple feature representation sets.
[0204] In an embodiment according to the present invention, as Figure 5As shown, the data processing system 200 includes a data acquisition module 210, a data cleaning module 220, a data aggregation module 230, a feature extraction module 240, and a data processing module 250.
[0205] The data acquisition module 210 is used to collect multiple open-source data according to data requirements and record the data source platforms corresponding to each open-source data. Among them, the data requirements include at least one or a combination of the following: ship route historical data, hot event data, supply port data, and weather data.
[0206] Open-source data refers to data that can be publicly obtained, used, and shared. In the data processing method for ship route inference, these open-source data come from multiple data source platforms and are data resources that can be utilized. Open-source data is the data basis for ship route inference.
[0207] The data requirements include at least one or a combination of the following: ship route historical data, hot event data, supply port data, and weather data. In other words, when the number of data requirements is one, the data requirement is any one of ship route historical data, hot event data, supply port data, and weather data. When the number of data requirements is multiple, the data requirement is any combination of ship route historical data, hot event data, supply port data, and weather data.
[0208] Regarding the ship route historical data, this part of the data is crucial for understanding the ship's previous driving routes, habitual routes, etc. By analyzing the ship route historical data, the navigation trajectory rules of the ship at different time periods and under different conditions can be summarized, providing a basic reference for subsequent route inference. Its sources are mainly the ship monitoring system and the underwater acoustic monitoring system, which can record relevant information such as the ship's navigation position in real time or regularly, thus forming historical data available for analysis.
[0209] Regarding the hot event data, these data mainly come from news platforms. Ship navigation may be affected by various hot events occurring around the waterway, such as sudden natural disasters in the waterway, waterway control caused by maritime accidents, major events at the port, etc. Understanding these hot events helps to consider the impact of external sudden factors on the ship's navigation route selection during ship route inference, so as to more reasonably plan the future route.
[0210] For the replenishment port data, this data is collected by the map service platform. The replenishment port is an important node for ships to conduct activities such as material replenishment and crew rest during navigation. Mastering relevant data such as the location, facility configuration, and operation status of the replenishment port can, in the ship route inference, reasonably plan the port for docking and replenishment and the subsequent navigation route based on the current state of the ship (such as the remaining fuel quantity, material reserve situation, etc.), ensuring the continuity and safety of ship navigation.
[0211] For the weather data, this data is collected by the meteorological service platform. Weather conditions have a direct and significant impact on ship navigation safety and route selection. Severe weather such as typhoons, heavy rains, and heavy fog may force ships to change their routes to avoid danger, while under good weather conditions, ships may also choose a more optimal direct route, etc. Therefore, weather data is one of the key factors that must be considered in ship route inference, which can help accurately predict the feasible routes of ships under different weather conditions.
[0212] In a specific embodiment, optionally, the data requirements include ship route historical data, hotspot event data, replenishment port data, and weather data. These different types of data requirements together constitute the data basis required for ship route inference. By comprehensively analyzing ship route historical data, considering the external impacts brought by hotspot event data, combining replenishment port data to plan reasonable docking points, and responding to meteorological changes based on weather data, it is possible to more comprehensively and accurately infer and plan the future navigation routes of ships, improving the safety, efficiency, and economy of ship navigation, etc.
[0213] In the process of collecting multiple open-source data in the present invention, the multi-source nature of collecting open-source data and the correlation between open-source data are reflected.
[0214] Multi-source nature: All kinds of open-source data involved in the data requirements come from different platforms (data source platforms), such as the ship's own monitoring and acoustic monitoring systems (ship monitoring system and underwater acoustic monitoring system), news platforms, map service platforms, meteorological service platforms, etc., fully reflecting the diversity of data sources to obtain information related to ship navigation from multiple perspectives.
[0215] Correlation: Although various types of open-source data seem independent, they are interrelated in the process of ship route inference. For example, hotspot events may affect the route selection of ships to replenishment ports, and weather data will also affect the route adjustment of ships based on route historical data, etc., working together to achieve accurate route inference.
[0216] The data cleaning module 220 is used to clean the open-source data to obtain the cleaned result data.
[0217] Data cleaning is an important step in the data processing process, and its purpose is to improve data quality. For open-source data, due to its wide and complex sources, various problems may exist, such as incomplete data, inconsistent data formats, data errors, duplicate data, etc. Through data cleaning, these problems can be solved, and thus cleaner and more usable cleaned result data can be obtained, providing a reliable basis for subsequent analysis and processing.
[0218] The cleaned result data better meets the requirements of subsequent data processing steps and provides high-quality data support for ship path inference.
[0219] The data collection module 230 is used to collect the cleaned result data according to the data source platform, and obtain multiple collected data sets corresponding one by one to multiple data source platforms.
[0220] Data collection is to classify and organize the data (cleaned result data) that has undergone data cleaning according to its data source platform. The purpose is to maintain the original relevance and traceability of the data, facilitate subsequent targeted analysis and processing according to the characteristics of data from different platforms, and at the same time is conducive to separately evaluating and managing the data quality of each source platform.
[0221] Specifically, for the cleaned result data from the ship monitoring system, it is collected into a set to form the first collected data set; for the cleaned result data from the underwater acoustic monitoring system, it is collected into a set to form the second collected data set; for the cleaned result data from the news platform, it is collected into a set to form the third collected data set; for the cleaned result data from the map service platform, it is collected into a set to form the fourth collected data set; for the cleaned result data from the meteorological service platform, it is collected into a set to form the fifth collected data set.
[0222] These collected data sets and the data source platforms are in a one-to-one correspondence relationship, and this correspondence relationship makes the management and use of data more orderly.
[0223] The feature extraction module 240 is used to extract features from the collected data sets to obtain a set of feature representations.
[0224] The number of sets of feature representations is multiple, and multiple sets of feature representations correspond one by one to multiple collected data sets.
[0225] Feature extraction is a crucial step in mining valuable information from a large amount of data. For the aggregated data set, the aim is to transform the original data into a more representative and analytically valuable form, namely the feature representation set. These features can better describe the essential attributes of the data, providing more concise and effective information for subsequent analysis tasks such as ship route inference, and helping to improve the accuracy and efficiency of the model.
[0226] Each aggregated data set is processed through specific methods to extract the key features. These features may cover multiple dimensions, such as statistical features, semantic features, temporal features, and geographical information features, etc. After extraction, the resulting feature representation set will embody the important information in the original data in a new and more compact data structure.
[0227] The data processing module 250 is used to establish an ontology model based on the analysis requirements, and construct a database according to the ontology model and multiple feature representation sets for ship route inference. Among them, the analysis requirements include at least one or a combination of the following: surface route analysis, underwater route analysis, and waypoint heat analysis.
[0228] In the case where the number of analysis requirements is one, the analysis requirement only includes any one of surface route analysis, underwater route analysis, and waypoint heat analysis. In the case where the number of analysis requirements is multiple, the analysis requirements include any combination of surface route analysis, underwater route analysis, and waypoint heat analysis.
[0229] The content of surface route analysis includes but is not limited to the width range of the waterway, water depth conditions, water flow speed and direction, and surface navigation rules.
[0230] Specifically, for different types of surface waterways, it is necessary to analyze their physical characteristics in detail. This includes the width range of the waterway. From narrow inland waterways to wide sea lanes, different widths will restrict the passage of ships of different sizes. Water depth conditions are also crucial. Shallow water areas may only allow ships with a shallow draft to pass, while deep water areas have less restrictions on the draft of ships. These water depth data will affect the safe navigation route of ships. The water flow speed and direction are also factors that cannot be ignored. Sailing with the current and against the current will greatly affect the ship's speed and fuel consumption, and thus affect the route selection. In addition, surface navigation rules are also part of the surface route analysis content, such as traffic control rules in different waters, the meaning of waterway markings, etc. These rules will restrict the navigation behavior of ships on the water surface.
[0231] Underwater route analysis mainly focuses on the impact of the underwater environment on the ship's navigation route. The content of underwater route analysis includes but is not limited to the distribution of underwater reefs and the depth and location information of trenches.
[0232] The complexity of the underwater terrain is an important factor. For example, the distribution of reefs, and ships need to avoid these potential dangerous areas; the depth and location information of trenches are crucial for the path planning of some special operation ships (such as diving operation ships). At the same time, underwater obstacles, whether naturally formed or man-made remnants (such as sunken ships), need to be considered in path planning to ensure the safety of underwater navigation of ships.
[0233] The heat analysis of waypoints mainly analyzes attributes such as the busyness and importance of each waypoint. By collecting and analyzing a large amount of ship navigation data, it is possible to determine which waypoints are hot spots where ships frequently pass, and these areas may be important port entrances and exits, channel intersections, etc. Understanding the heat situation of waypoints helps to reasonably select waypoints in ship path reasoning, improve navigation efficiency, and avoid unnecessary congestion.
[0234] In a specific embodiment, optionally, the analysis requirements include surface route analysis, underwater route analysis, and heat analysis of waypoints. Based on these analysis requirements, an ontology model is established. The ontology model is an abstract representation of the knowledge in the field of ship path reasoning. After establishing the ontology model, a database is constructed according to the ontology model and multiple feature representation sets obtained through data collection and feature extraction. The database constructed in this way can effectively support the ship path reasoning work and provide data support for accurately and efficiently planning the ship navigation path.
[0235] The present invention aims to provide a data processing system 200, which fully combines the relevance among the business requirements (including data requirements and analysis requirements), the data source platform, and the acquired data. Whether in the process of feature extraction or in the process of database establishment, it can improve the operation efficiency of data processing and also improve the accuracy of the result of reasoning about the ship path.
[0236] Specifically, the data processing system 200 provided by the present invention combines the business requirements, the data source platform, and the acquired data, and the relevance among multiple parties. During the process of data feature extraction and database establishment, it can improve the efficiency and make the data storage speed fast. It should be emphasized that during the process of database establishment, an ontology model is established in combination with the requirements of ship path reasoning (an ontology model is established based on multiple analysis requirements), making the subsequent data retrieval more timely and accurate. This data processing method can ensure that when there are fluctuations in a part of the open-source data collected, it has no associated impact on the open-source data of other data source platforms, which is beneficial to improving the fault tolerance rate of the system.
[0237] In a specific embodiment, optionally, both the map service platform and the meteorological service platform are connected to the data acquisition module 210 of the data processing system 200 through an API (Application Programming Interface) interface. Among them, the API interface is a technical means that allows different software systems to communicate with each other.
[0238] The data acquisition module 210 is connected to the map service platform and the meteorological service platform through the API interface, which reflects the convenience and efficiency of data integration. This connection method enables the ship route inference system (data processing system 200) to obtain data from these two important external platforms in a standardized and programmed manner, without complex underlying interaction mechanisms, so as to better integrate geographical information and meteorological information into the overall data processing process of ship route inference and provide a more comprehensive basis for ship route planning.
[0239] In an embodiment according to the present invention, as Figure 6 shown, the electronic device 300 includes a memory 310 and a processor 320. Among them, a program or instruction that can run on the processor 320 is stored on the memory 310, and when the processor 320 executes the program or instruction, the steps of the data processing method in any of the above embodiments are implemented.
[0240] The electronic device 300 has the beneficial effects of any of the above embodiments, which will not be elaborated here.
[0241] In an embodiment according to the present invention, a readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the data processing method in any of the above embodiments are implemented.
[0242] The readable storage medium has the beneficial effects of any of the above embodiments, which will not be elaborated here.
[0243] In the present invention, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance; the term "plural" means two or more, unless otherwise clearly defined. Terms such as "installation", "connection", "connection", and "fixation" should all be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; "connection" can be a direct connection or an indirect connection through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0244] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "left", "right", "front", "rear", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or unit referred to must have a specific direction, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.
[0245] In the description of this specification, the descriptions of terms such as "one embodiment", "some embodiments", "specific embodiments", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or instance. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0246] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data processing method, characterized in that: For ship path reasoning, the data processing method includes: Collect multiple open source data according to data requirements, and record the data source platform corresponding to each open source data; wherein the data requirements include at least one of the following or a combination thereof: ship route history data, hot event data, supply port data and weather data; Performing data cleaning on the open source data to obtain cleaning result data; Collecting the cleaning result data according to the data source platform to obtain multiple collected data sets corresponding to multiple data source platforms; Performing feature extraction on the collected data set to obtain a feature representation set; An ontology model is established based on analysis requirements, and a database is constructed according to the ontology model and a plurality of feature representation sets to perform ship path reasoning; wherein the analysis requirements include at least one of the following or a combination thereof: surface route analysis, underwater route analysis, and waypoint thermal analysis.
2. The data processing method according to claim 1, characterized in that: In the case where the data demand includes the ship route history data, the data source platform includes a ship monitoring system and an underwater acoustic monitoring system; In the case where the data demand includes the hot event data, the data source platform includes a news platform; In the case where the data demand includes the supply port data, the data source platform includes a map service platform; In the case where the data demand includes the weather data, the data source platform includes a meteorological service platform.
3. The data processing method according to claim 1, characterized in that: The step of performing data cleaning on the open source data to obtain cleaning result data includes: Setting a configuration policy corresponding to a file of the open source data, and importing the open source data through the configuration policy; Determine metadata of data attributes corresponding to the open source data, compare the imported open source data with the metadata, and obtain a screened open source data set; Abnormal data processing is performed based on the screened open source data set to obtain the cleaning result data.
4. The data processing method according to claim 1, characterized in that: The step of extracting features from the collected data set to obtain a feature representation set includes: Performing entity recognition on the plurality of collected data sets based on a large language model to obtain a plurality of entities; A knowledge graph is established, and similarities are calculated and matched between the multiple entities identified and the entity categories in the knowledge graph to obtain the feature representation set.
5. The data processing method according to claim 1, characterized in that: The feature representation set includes at least one of the following or a combination thereof: semantic features, statistical features, temporal features and geographic information features.
6. The data processing method according to any one of claims 1 to 5, characterized in that: The ontology model is established based on the analysis requirements, and a database is constructed according to the ontology model and a plurality of feature representation sets to perform ship path reasoning, including: Establishing the ontology model based on the analysis requirements; Taking the ontology in the ontology model as the target ontology and the ontology in the feature representation set as the source ontology, and determining the association relationship between the target ontology and the source ontology; Based on the target ontology, the source ontology and the association relationship, the database is constructed to perform ship path reasoning.
7. The data processing method according to claim 6, characterized in that: The ontology model includes: ontology name, ontology type and ontology description; and / or The association relationship includes: the relationship type between the target ontology and the source ontology, and the relationship direction between the target ontology and the source ontology.
8. A data processing system, characterized in that: include: A data collection module (210) is used to collect a plurality of open source data according to data requirements, and record the data source platform corresponding to each of the open source data; wherein the data requirements include at least one of the following or a combination thereof: ship route history data, hot event data, supply port data and weather data; A data cleaning module (220), used for cleaning the open source data to obtain cleaning result data; A data collection module (230) is used to collect the cleaning result data according to the data source platform to obtain multiple collection data sets corresponding to multiple data source platforms; A feature extraction module (240) is used to extract features from the collected data set to obtain a feature representation set; A data processing module (250) is used to establish an ontology model based on analysis requirements, and to construct a database according to the ontology model and a plurality of feature representation sets to perform ship path reasoning; wherein the analysis requirements include at least one of the following or a combination thereof: surface route analysis, underwater route analysis, and waypoint thermal analysis.
9. An electronic device, characterized in that: include: A memory (310) and a processor (320), wherein the memory (310) stores a program or instruction that can be run on the processor (320), and when the processor (320) executes the program or the instruction, the steps of the data processing method according to any one of claims 1 to 7 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or an instruction, and when the program or the instruction is executed by a processor, the steps of the data processing method according to any one of claims 1 to 7 are implemented.