Navigation data comprehensive treatment platform based on trusted data space
By adopting a comprehensive governance platform based on trusted data space in the shipping industry, the problems of data silos, low quality and difficulty in sharing in the shipping industry have been solved, and comprehensive, efficient and safe data governance has been achieved, and the operational efficiency and competitiveness of enterprises have been improved.
Patent Information
- Application Number
- CN202510227113.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-20
AI Technical Summary
The shipping industry faces problems such as data silos, low data quality, and difficulty in data sharing, which seriously restricts its intelligent and efficient development.
The comprehensive shipping data governance platform based on trusted data space is adopted, and comprehensive, efficient and safe governance of shipping data is achieved through modules such as data extraction, specification, processing, quality management, resource management, sharing and security services.
It effectively solves the problems of data silos, low data quality and difficult data sharing in the shipping industry, realizes comprehensive, efficient and safe data governance, and improves the operational efficiency and competitiveness of shipping companies.
Smart Images

Figure CN120181783A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a comprehensive shipping data governance platform based on a trusted data space. Background Art
[0002] At present, as an important pillar of the global economy, the shipping industry has a huge and complex amount of data. However, traditional data governance methods often have problems such as data islands, low data quality, and difficulty in data sharing, which seriously restrict the intelligent and efficient development of the shipping industry. To solve these problems, the shipping industry urgently needs a comprehensive, efficient, and secure data governance solution. In this context, the trusted data space technology has emerged. The trusted data space is a data sharing and circulation platform built based on advanced technologies such as blockchain and privacy computing. It can achieve "usable but invisible, controllable and measurable" for data, ensuring the security and privacy of data during the sharing process. The emergence of this technology provides new ideas and solutions for the data governance of the shipping industry. Summary of the Invention
[0003] The purpose of the present invention is to provide a comprehensive shipping data governance platform based on a trusted data space to solve the above problems.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions:
[0005] A comprehensive shipping data governance platform based on a trusted data space includes a data extraction module, a data specification module, a data processing module, a data quality module, a data resource module, a data sharing module, and a data security module. The data extraction module is used to comprehensively and flexibly extract relevant shipping data from various data sources and provide various extraction methods and data storage supports. The data specification module is used to conduct a comprehensive and systematic standard specification design for shipping data to ensure data quality and interface compatibility. The data processing module is used to perform secondary processing on the extracted shipping data to form standardized and normalized data and provide various processing components to meet the data processing requirements in different scenarios. The data quality module is used to conduct quality audits on shipping data from multiple dimensions and generate corresponding quality reports. The data resource module is used to conduct comprehensive management of resources for the standardized shipping data. The data sharing module is used to generate APIs for shipping data and externally open data resources in the form of APIs. The data security module is used to provide security services during the process of shipping data processing.
[0006] Preferably, the types of the shipping data include ship static data, dynamic data, and marine environment data.
[0007] Preferably, the data extraction module supports real-time extraction of structured data, semi-structured data, unstructured data, and network data from various AIS devices and data source types. Its extraction methods include full extraction, timed incremental extraction, and real-time incremental extraction, and its storage methods include offline data storage and real-time data storage.
[0008] Preferably, the data standardization module includes a theme design unit, a standard design unit, a model design unit, and an index design unit; the theme design unit standardizes data by performing theme management and data management on shipping data; the standard design unit is expressed in the form of a class UML diagram, provides clear service definitions and service technical designs, emphasizes the creation process of service instances, service data models, and physical data models, and provides functions for adding service definitions, adding service technical designs, adding service instances, adding service data models, adding physical data models, adding relationships, and adding annotations; the model design unit designs data models in a drag-and-drop manner and draws them in the form of a class UML diagram to intuitively create and edit data models; the index design unit configures dataset indicators through SQL statements and a visual interface, defines and analyzes various indicators according to business requirements, and visually displays the execution results of the indicators.
[0009] Preferably, the specific process of the secondary processing of shipping data by the data processing module is as follows:
[0010] A1. Create a custom job type and configure the job;
[0011] A2. Add the configured job to the workflow in a drag-and-drop manner to combine and optimize the data processing flow;
[0012] A3. Configure the periodic scheduling or custom time scheduling method as needed to execute tasks regularly;
[0013] A4. After the task is executed, monitor whether the task status is successfully executed or whether an exception occurs. If an exception occurs, predict the risk and handle the problem. At the same time, the task supplement function is supported during the task execution process.
[0014] Preferably, the data quality module supports database-level, table-level, field-level, and custom-type auditing tasks, outputs a table-level quality report on quality issues regarding data integrity, uniqueness, authority, legality, and consistency, and provides a quality improvement plan at the task instance level in the quality report.
[0015] Preferably, the specific method for the source comprehensive management of shipping data by the data resource module is as follows: Set up a data resource library containing data directories, data maps, and data items of various nautical data, and retrieve data through the data directory for online preview.
[0016] Preferably, the data resource module records data resource applications and access logs, supports the application approval and subscription usage processes of data resources, and analyzes the sources, destinations, and dependencies of data by drawing a relational network diagram of the data.
[0017] Preferably, the data sharing module includes an API online development unit, an API online debugging unit, and an API secure access unit; the API online development unit directly creates objects through an API development platform and defines data items. The data objects adopt a multi-layer tree structure. If a data object already exists, multiple objects are combined into a composite object structure; the API online debugging unit conducts online debugging of the interfaces within the API development platform; the API secure access unit realizes the security authentication of API calls through an API gateway and can statistically analyze the API call situations; the security services provided by the data security module are specifically as follows: based on the importance of the data, the shipping data is classified through a data classification evaluation formula, and at the same time, sensitive data is desensitized and encrypted, access rights to the data are controlled, and the entire process of data usage is recorded to generate a data usage log.
[0018] Preferably, the data classification evaluation formula comprehensively scores based on three factors: the value of the data, the release scope, and the leakage impact. According to the scoring results, the shipping data is divided into four different confidentiality levels: top secret level, confidential level, secret level, and general level, and corresponding security services are matched according to the confidentiality levels. The data classification evaluation formula is: K = 20% * data value score + 30% * data release scope score + 50% * data leakage impact score, where
[0019] C4: Top secret level: 3.8 ≤ K ≤ 4.0 C3:
[0020] C3: Confidential level: 2.8 ≤ K < 3.8
[0021] C2: Secret level: 1.8 ≤ K < 2.8
[0022] C1: General level: K < 1.8.
[0023] After adopting the above technical solutions, compared with the background technology, the present invention has the following advantages:
[0024] The present invention provides a comprehensive governance platform for shipping data based on a trusted data space. By comprehensively governing all aspects and elements of the shipping data engineering process, it ensures the complete functionality and complete process of the data engineering package, and realizes the comprehensive, efficient, and secure governance of shipping data. By constructing a trusted data space for shipping data, it effectively solves problems such as data islands, low data quality, and difficult data sharing in the shipping industry. It realizes the comprehensive, efficient, and secure governance of data, provides more convenient and reliable data support for shipping enterprises, and helps to improve the operation efficiency and competitiveness of enterprises. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a system structure diagram of the method of the present invention;
[0026] Figure 2 It is a business flow diagram of data extraction of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0028] Embodiment
[0029] Please refer to Figure 1 and Figure 2 As shown, the present invention discloses a comprehensive governance platform for shipping data based on a trusted data space, including a data extraction module, a data specification module, a data processing module, a data quality module, a data resource module, a data sharing module, and a data security module. The data extraction module is used to comprehensively and flexibly extract relevant shipping data from various data sources and provide various extraction methods and data storage supports. The data specification module is used to conduct a comprehensive and systematic standard specification design for shipping data to ensure data quality and interface compatibility. The data processing module is used to perform secondary processing on the extracted shipping data to form standardized and normalized data and provide various processing components to meet the data processing requirements in different scenarios. The data quality module is used to conduct quality audits on shipping data from multiple dimensions and generate corresponding quality reports. The data resource module is used to comprehensively manage the resources of the standardized shipping data. The data sharing module is used to generate APIs for shipping data and open up data resources in the form of APIs. The data security module is used to provide security services during the processing of shipping data.
[0030] The types of shipping data include ship static data, dynamic data, and marine environment data.
[0031] The data extraction module supports real-time extraction of structured data, semi-structured data, unstructured data, and network data from various AIS devices and data source types. Its extraction methods include full extraction, timed incremental extraction, and real-time incremental extraction, and its storage methods include offline data storage and real-time data storage.
[0032] When extracting data, it is necessary to ingest and transform all (or part) of this dataset because it contains relevant machine / data knowledge for our use case. These datasets come in various structured or semi-structured formats. To scale this data, we have developed programming methods as an alternative to ETL excellence. Collections of Python or Javaocos are bundled in packages for data engineers. The latter can generate large-scale rdf triples from large structured data files such as database extractions, and we link these datasets to their records and etl pipeline systems. Whenever available, the data models (entity-relationship models or class diagrams) that describe the structure and content of the datasets are transformed or modeled in RDF format, loaded into the KG, and linked to the dataset uris.
[0033] The semantic etl pipeline uses the form of normalized structured and semi-structured datasets. We apply matching algorithms (similar to OpenRefine(92)(93)(103)) that perform content reconciliation (based on ontology entities and reference data) and structural validation according to predefined SHACL shapes. This step is crucial for obtaining a logical representation (a model) of the dataset that seamlessly integrates with the KG and is linked to existing DG concepts. All details related to semantic transformation (such as code, language, execution environment) are recorded in the KG and associated with the original input da task and output semantic model.
[0034] During the data extraction process, the extraction tasks are configured using visual drag-and-drop, providing an intuitive interface that allows users to easily define and configure the data extraction process. Through visual drag-and-drop, users can easily select the data sources to be extracted, specify the fields to be extracted, define the rules for data transformation and processing, and set the storage destinations for the data. This configuration method can quickly build data extraction tasks, reducing the workload of writing complex code and improving the efficiency and accuracy of data extraction. The visual drag-and-drop method also supports the dynamic generation of collection functions. Users can, through simple configuration, achieve the dynamic generation of data collection functions without writing additional code. This method can greatly simplify the configuration process of data collection and improve the collection efficiency.
[0035] Full extraction means extracting all the content of the source data at once and loading all the data into the target system. This extraction method is applicable to situations where the data volume is small or the data update frequency is low. In the extraction of navigation-related data, if the data volume is not large or the data update frequency is low, the full extraction method can be adopted.
[0036] Regular incremental extraction means regularly extracting the updated data part from the source system at a certain time interval. This extraction method is applicable to situations where the data volume is large or the data update frequency is high. In the extraction of navigation-related data, if the data volume is large or the data update frequency is high, the regular incremental extraction method can be adopted. By setting an appropriate time interval, the real-time nature and accuracy of the data can be ensured.
[0037] Real-time incremental extraction means extracting the updated data part from the source system in real time and immediately loading it into the target system. This extraction method is applicable to situations where high real-time requirements for data are needed. In the extraction of navigation-related data, if it is necessary to obtain and process the latest ship positions, navigation status, etc. in real time, the real-time incremental extraction method can be adopted. By updating the data in real time, the real-time nature and accuracy of the data can be ensured.
[0038] Offline data storage means the data that has already occurred, which can be stored and accessed at any time without real-time updates. This kind of data is usually used in scenarios such as historical data analysis and data mining, and can be stored and processed in the data storage system for a long time. In the extraction of navigation-related data, offline data may include global ship data, crew profile data, electronic nautical charts, global port data, etc.
[0039] Real-time data storage means the latest data that occurs, which needs to be stored and accessed in real time to support real-time business decision-making and monitoring. This kind of data is usually used in scenarios such as real-time monitoring, early warning systems, and real-time analysis, and needs to be updated and accessed in real time in the data storage system. In the extraction of navigation-related data, real-time data may include global ship AIS data, global container tracking data, bulk shipping bulk cargo tracking data, etc.
[0040] Data source types include MySQL, Oracle, SQL Server, PostgreSQL, DB2, Kingbase, Gbase, DM, Greenplum, Hdfs, Hive, Hbase, Elasticsearch, Clickhouse, Kudu, Kafka, Redis, MongoDb, Ftp, and Http, etc. The diversity of data source types enables the data engineering process package to extract and process data from various different data sources. By supporting multiple data source types, the data engineering process package can meet the needs of different users and adapt to various complex data environments and business scenarios. For each data source type, the data engineering process package provides corresponding connectors and adapters to enable communication and interaction with various data sources. By using these connectors and adapters, users can easily establish connections with different data sources and perform operations such as data extraction, transformation, and storage.
[0041] The data specification module includes a theme design unit, a standard design unit, a model design unit, and an indicator design unit; the theme design unit standardizes data by performing theme management and data management on shipping data; the standard design unit is expressed in the form of a class UML diagram, provides clear service definitions and service technology designs, emphasizes the creation process of service instances, service data models, and physical data models, and provides functions such as adding service definitions, adding service technology designs, adding service instances, adding service data models, adding physical data models, adding relationships, and adding annotations; the model design unit designs data models in a drag-and-drop manner and draws them in the form of a class UML diagram to intuitively create and edit data models; the indicator design unit configures dataset indicators through SQL statements and a visual interface, defines and analyzes various indicators according to business requirements, and visualizes the execution results of the indicators.
[0042] In model design, users can add elements such as classes, interfaces, methods, and relationships in a drag-and-drop manner and can use annotations to describe and record design ideas and details. By using a way similar to UML diagrams for drawing, users can more clearly express the structure and relationships of data models, making data processing and analysis more accurate and reliable. The underlying data is stored in the database in text form during the model design process. This storage method has flexibility and scalability, enabling users to easily manage and maintain data models. At the same time, model design also provides data import and export functions, enabling users to conveniently interact and integrate data models with other systems or tools.
[0043] The specific process of the secondary processing of shipping data by the data processing module is as follows:
[0044] A1. Create custom job types and configure jobs. The custom job configuration allows users to create custom job types, which can be shell scripts, SQL statements, and also support various other data processing methods such as Hive, HiveSQL, Spark, SparkSQL, Sqoop, Streaming SQL, PrestoSql, ImpalaSql, etc. This custom job configuration function provides greater flexibility and scalability, enabling users to create and run various different types of jobs according to their own needs and skill levels.
[0045] A2. Add the configured jobs to the workflow in a drag-and-drop manner to combine and optimize the data processing flow to meet different data processing requirements.
[0046] A3. Configure periodic scheduling or custom time scheduling methods as needed to execute tasks regularly.
[0047] Among them, periodic scheduling allows users to schedule and execute jobs at specified time intervals. Users can set the time intervals for scheduling, such as daily, weekly, or monthly, etc. Through periodic scheduling, users can achieve automated execution and regular processing of jobs. Custom scheduling allows users to customize the scheduling time and method of jobs according to their needs. Users can specify the execution time, execution order, and execution conditions of jobs, etc., to meet the data processing requirements in different scenarios. Custom scheduling provides greater flexibility and scalability, enabling users to adjust and optimize the scheduling strategy at any time according to business changes and demand changes.
[0048] A4. After the task is executed, monitor whether the task status is successfully executed or whether an exception occurs. If an exception occurs, predict the risk and handle the problem. At the same time, the task supplement function is supported during the task execution process.
[0049] After the task scheduling is executed, it is necessary to monitor the task status to understand whether it is successfully executed or whether an exception occurs. The task warning function can help operation and maintenance personnel predict risks in advance and handle problems, solving problems faster and more accurately. It supports warnings for task execution failure or success, server node exceptions, server startup timeouts, and task startup timeouts to monitor the task running status.
[0050] Task monitoring is one of the important functions in the data processing process. It monitors the processing tasks in real time and provides a series of operation options, such as rerunning, closing, and log viewing. When an exception or failure warning occurs in the task. Task monitoring supports rerunning the processing task. If the task is executed successfully but the user needs to close the task, the close operation can be used. In addition, if you want to monitor the status of the processing task in real time, you can view the running log of the task. The task monitoring function monitors the processing tasks in real time and provides a series of operation options, enabling users to control the task execution process more flexibly.
[0051] The task supplement function includes a retry mechanism, a timeout mechanism, and a complement mechanism, aiming to improve the reliability and stability of data processing and analysis.
[0052] Retry mechanism: When the task execution fails, the platform can automatically or manually trigger the retry mechanism to re-execute the task. The retry mechanism can configure the number of retries and the interval time to ensure that the task can be successfully executed. Through the retry mechanism, users can solve temporary failures or accidental problems and improve the quality and efficiency of data processing.
[0053] Timeout mechanism: During the task execution process, if the task execution time exceeds the preset time limit, the platform can trigger the timeout mechanism to stop the task execution and return an error message. The timeout mechanism can configure the timeout time to prevent the task from occupying system resources for a long time or causing system crashes due to abnormal task execution. Through the timeout mechanism, users can promptly discover and handle tasks that have not responded for a long time, improving the stability and availability of the system.
[0054] Complement mechanism: During the data processing process, if data is found to be missing or abnormal, the platform can trigger the complement mechanism to complete or correct the data. The complement mechanism can configure the complement strategy and method to supplement data according to the actual situation. Through the complement mechanism, users can solve data quality problems and improve the integrity and accuracy of data processing and analysis.
[0055] The data quality module supports database-level, table-level, field-level, and custom-type auditing tasks, outputs table-level quality reports on quality issues regarding data integrity, uniqueness, authority, legality, and consistency, and provides quality improvement solutions at the task instance level in the quality report. The quality improvement solutions include completing missing data, removing duplicate data, selecting authoritative data sources, establishing data integrity judgment rules, and consistency preprocessing.
[0056] The specific method for the data resource module to comprehensively manage the sources of shipping data is as follows: Set up a data resource library containing various nautical data, data maps, and data items, and retrieve data through the data directory for online preview.
[0057] The data resource module records data resource applications and access logs, supports the application approval and subscription usage processes of data resources, facilitates school teachers and students to quickly obtain the required data information, and apply for the use of data resources as needed. By drawing a network diagram of the relationships between data, it analyzes the sources, destinations, and dependencies of data.
[0058] A data catalog is a tool for managing and organizing data. It helps users understand the structure, classification, and metadata information of data. The data catalog can include a two-layer structure of data groups and data sets, and each layer of the structure includes functions for adding, deleting, and modifying.
[0059] The system automatically reads the pre-data and derived data fields of the data resource set to form a relationship network diagram with data flow. It helps users understand the upstream and downstream flow rules of data in different systems, as well as the impact relationships between data. The relationship network diagram manages various data source and usage scenario information, and analyzes the dependencies between data tasks through data lineage tracking, thus assisting users to efficiently understand and utilize data resources.
[0060] The data resource process adopts a permission-based access control mechanism. Users need to apply to the user group administrator for access to specific data. After being approved by the user group administrator, further approval by the data administrator is required. Users can view the introductions of data groups, data sets, and data items, but can only access the data items for which they are granted permissions in the user permission table. This process ensures the secure access and legal utilization of data.
[0061] The data sharing module includes an API online development unit, an API online debugging unit, and an API secure access unit; the API online development unit directly creates objects through the API development platform and defines data items. The data object adopts a multi-layer tree structure, and one object can correspond to multiple tables in the database. If the data object already exists, multiple objects are combined into a composite object structure; the API online debugging unit conducts online debugging of the interface within the API development platform; the API secure access unit realizes the secure authentication of API calls through the API gateway and can statistically analyze the API call situations.
[0062] The system automatically generates an API documentation interface according to the interfaces defined by developers. Users can select the interfaces to be debugged in the documentation and select the corresponding HTTP request methods, such as GET, POST, PUT, etc. The interface will display information such as request parameters and return data formats in real time, and at the same time supports inputting request parameters and immediately sending requests to view and check the request response results, effectively verifying whether the interface works as expected. Compared with directly debugging interfaces in the development environment, such an online debugging function is more intuitive and simple, and the interface debugging work can be completed without deploying a test environment, greatly improving the development efficiency.
[0063] API security access adopts the gateway authentication method. All API calls need to go through the gateway, which is responsible for security authentication. Only users with an AK (access key) can access API resources. Users need to create an application in the API application management interface and apply for access rights to the corresponding data. After the application is successfully created, the platform will assign an independent AK. The platform realizes fine-grained control of data access by restricting the access scope of each AK.
[0064] API call situation analysis is of great significance for optimizing the usage efficiency of APIs and monitoring call security. The platform provides a detailed API call statistics function. It can record the number of calls and traffic usage of each AK (access key) within a certain period of time, regularly count the total number of API calls of each application, and also count the call frequency of each IP address. These data will be fed back to the background management interface in real time. System administrators can log in to the interface to query the detailed call records of specified users or IP addresses within any time period, including indicators such as the number of successful and failed calls and the average response time of each API interface, so as to monitor abnormal calls and adjust API strategies in a timely manner.
[0065] The security services provided by the data security module are specifically as follows: Based on the importance of the data, the shipping data is classified through a data classification evaluation formula. At the same time, sensitive data is desensitized and encrypted, access rights to the data are controlled, and the entire process of data usage is recorded to generate a data usage log.
[0066] The data classification evaluation formula comprehensively scores based on three factors: data value, release scope, and leakage impact. According to the scoring results, the shipping data is divided into four different confidentiality levels: top secret, confidential, secret, and general, and corresponding security services are matched according to the confidentiality level. The data classification evaluation formula is: K = 20% * data value score + 30% * data release scope score + 50% * data leakage impact score, where
[0067] C4: Top secret level: 3.8 ≤ K ≤ 4.0 C3:
[0068] C3: Confidential level: 2.8 ≤ K < 3.8
[0069] C2: Secret level: 1.8 ≤ K < 2.8
[0070] C1: General level: K < 1.8.
[0071] The calculation method of the data value score is shown in Table 1 below:
[0072]
[0073] The data release scope is shown in Table 2 below:
[0074]
[0075] The impact of data leakage is shown in Table 3 below:
[0076]
[0077] To better protect users' personal information, users are allowed to set different desensitization methods for different fields. For example, the name can be desensitized to display only the first letter, the last few digits of the mobile phone number can be desensitized, and the last few digits of the ID number can be hidden. These strategies can effectively hide personal private information while retaining the value of data analysis, maximizing the protection of user privacy. The system also records the desensitization history of each field, facilitating administrators to optimize strategies and comprehensively ensuring the security of the data usage process, providing an important guarantee for protecting user privacy and enhancing user experience.
[0078] Set different access policies according to the data level, only open data access that meets the minimum permissions, and achieve "data statistics and export on demand". When users access data, the platform will conduct security checks such as identity authentication, permission verification, black and white list verification, and access limit. The permission control has functions of identity check, permission grading check, black and white list check, and access limit check. Each user obtains general permissions when creating an account, and the additional permissions applied for later are realized through the priority override mechanism to achieve flexible user permission definition. The platform will first match the user's precise permissions for the data set. If the match fails, the general permissions will be used, and at the same time, white list and black list controls are supported to ensure secure data access.
[0079] Each detail of the entire process of accessing user data is carefully recorded and logged. When a user views or downloads a certain data set, the system background will automatically generate a detailed operation event record, recording key information such as the user's identity, operation time, data set name, operation type (view or download), etc. These event records will be pushed to the administrator's account in real time, so that the administrator can timely understand the usage situation of each data set access in the system, forming a complete and detailed operation history record chain. At the same time, the system also gives the administrator the function of querying all event records generated within a certain time range in chronological order.
[0080] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A shipping data integrated management platform based on a trusted data space, characterized by: It includes a data extraction module, a data specification module, a data processing module, a data quality module, a data resource module, a data sharing module and a data security module. The data extraction module is used to comprehensively and flexibly extract relevant shipping data from various data sources, and provide a variety of extraction methods and data storage support. The data specification module is used to carry out comprehensive and systematic standard specification design for shipping data to ensure data quality and interface compatibility. The data processing module is used to perform secondary processing on the extracted shipping data to form standardized and normalized data, and provide a variety of processing components to meet the data processing requirements in different scenarios. The data quality module is used to perform quality audits on shipping data in multiple dimensions and generate corresponding quality reports. The data resource module is used to perform comprehensive resource management on shipping data after standardization. The data sharing module is used to generate APIs for shipping data and open data resources to the outside in the form of APIs. The data security module is used to provide security services during the processing of shipping data.
2. The shipping data integrated management platform based on a trusted data space as claimed in claim 1, characterized in that: The types of shipping data include ship static data, dynamic data and marine environment data.
3. The shipping data integrated management platform based on a trusted data space as claimed in claim 1, characterized in that: The data extraction module supports real-time extraction of structured data, semi-structured data, unstructured data and network data from various AIS devices and data source types, and its extraction methods include full extraction, timed incremental extraction and timely incremental extraction, and its storage methods include offline data storage and real-time data storage.
4. The shipping data integrated management platform based on a trusted data space as claimed in claim 1, characterized in that: The data specification module includes a theme design unit, a standard design unit, a model design unit and an indicator design unit; the theme design unit normalizes the data by performing theme management and data management on the shipping data; the standard design unit is expressed in a UML-like manner, provides a clear service definition and service technical design, emphasizes the creation process of service instances, service data models and physical data models, and provides the functions of adding service definitions, adding service technical designs, adding service instances, adding service data models, adding physical data models, adding relationships, and adding comments; the model design unit uses a drag-and-drop method to design data models, and draws them in a UML-like manner to intuitively create and edit data models; the indicator design unit configures data set indicators through SQL statements and a visual interface, defines and analyzes various indicators according to business needs, and displays indicator execution results through visualization.
5. The shipping data integrated management platform based on a trusted data space as claimed in claim 1, characterized in that: The secondary processing of shipping data by the data processing module is specifically as follows: A1. Create a custom job type and configure the job; A2. Add configured jobs to the workflow by dragging and dropping to combine and optimize the data processing process; A3. Configure periodic scheduling or custom time scheduling as needed to execute tasks regularly; A4. After the task is executed, monitor the task status to see whether it is successfully executed or whether any exception occurs. If an exception occurs, predict the risk and handle the problem. At the same time, the task supplement function is supported during the task execution process.
6. The shipping data integrated management platform based on a trusted data space as claimed in claim 1, characterized in that: The data quality module supports database audit tasks at the library level, table level, field level and custom types, outputs table-level quality reports on quality issues such as data integrity, uniqueness, authority, legality and consistency, and provides task instance-level quality improvement solutions in the quality reports.
7. The shipping data integrated management platform based on a trusted data space as claimed in claim 1, characterized in that: The data resource module's method for comprehensive management of shipping data sources is specifically as follows: setting up a data directory, data map and data resource library containing various types of navigation data, retrieving data through the data directory and performing online preview.
8. The shipping data integrated management platform based on a trusted data space as claimed in claim 7, characterized in that: The data resource module records data resource applications and access logs, supports the application approval and subscription use process of data resources, and analyzes the source, destination and dependency of data by drawing a relationship network diagram between data.
9. The shipping data integrated management platform based on a trusted data space as claimed in claim 1, characterized in that: The data sharing module includes an API online development unit, an API online debugging unit and an API security access unit; The API online development unit directly creates objects through the API development platform and defines data items. The data objects adopt a multi-layer tree structure. If the data objects already exist, multiple objects are combined into a composite object structure. The API online debugging unit performs online debugging of the interface through the API development platform. The API security access unit implements security authentication of API calls through the API gateway, and can perform statistical analysis on API calls; the security services provided by the data security module are as follows: based on the importance of the data, the shipping data is classified into grades through a data classification evaluation formula, and at the same time, sensitive data is desensitized and encrypted, access rights to the data are controlled, and the entire process of data use is recorded to generate a data usage log.
10. The shipping data integrated management platform based on a trusted data space as claimed in claim 9, characterized in that: The data classification evaluation formula is based on the three factors of data value, release scope and leakage impact. According to the scoring results, the shipping data is divided into four different confidentiality levels: top secret, confidential, secret and ordinary. The corresponding security services are matched according to the confidentiality level. The data classification evaluation formula is: K = 20% * data value score + 30% * data release scope score + 50% * data leakage impact score, where C4: Top Secret: 3.8≤K≤4.0C3: C3: Confidential level: 2.8≤K<3.8 C2: Secret level: 1.8≤K<2.8 C1: Ordinary level: K<1.8.