Maritime document knowledge graph construction method and system

Through the maritime document knowledge graph construction method, a large language model is used to extract and structure key information from complex maritime data to build a high-quality knowledge graph, solving the problem of difficult information extraction and real-time update requirements in maritime data, and achieving efficient retrieval and intelligent analysis support.

CN119938916AActive Publication Date: 2025-05-06COSCO SHIPPING TECH CO LTD

Patent Information

Application Number
CN202510025104.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately extract information from maritime data with complex sources and build high-quality knowledge maps, especially in terms of inconsistent data quality, difficulty in extracting information and real-time update requirements.

Method used

A method for building a maritime document knowledge graph is proposed. Through steps such as data collection and preprocessing, text segmentation, key information extraction, information integration and summary, community classification, community report generation, vectorization processing, etc., a large language model is used to automatically identify and structure key entities and their relationships to build a high-quality knowledge graph.

Benefits of technology

It realizes accurate extraction of information from complex maritime data and builds a high-quality knowledge graph, significantly improves the accuracy and refinement of knowledge, supports efficient retrieval and intelligent analysis, and adapts to the real-time update requirements of maritime information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938916A_ABST
    Figure CN119938916A_ABST
Patent Text Reader

Abstract

The invention provides a maritime document knowledge graph construction method and system, and the method comprises the steps: carrying out the collection, preprocessing and text segmentation of a maritime document, carrying out the extraction, information integration and description summarization of key information through a large language model, and obtaining the node data, edges and the attributes of nodes and edges of a graph; through community classification, community abstract generation and vectorization processing, the data are grouped and optimized, and finally a structured knowledge graph is constructed through integration. A large language model is adopted to automatically identify the processed data and extract key entities and relationships thereof, so that the accuracy and the refining degree of data information are improved; a potential relation network between entities is extracted through community classification, and a basis is provided for efficient retrieval; through the steps, key information can be intelligently extracted from maritime data updated in real time, the knowledge graph construction accuracy and comprehensiveness are improved, the efficiently retrieved knowledge graph is constructed, and powerful support is provided for information retrieval, intelligent analysis and strategic decision of a shipping company.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to a method and system for constructing a maritime document knowledge graph. Background Art

[0002] The shipping industry is an important pillar of the global economy, responsible for about 80% of the world's cargo transportation needs, covering a wide range from energy supply to commodity trade. With the rapid development of data and information technology, the maritime industry has accumulated a large amount of unstructured data resources, including ship operation records, technical documents, laws and regulations, and news reports. These data resources contain rich information and can be used to support decisions in many aspects such as ship management, fuel consumption optimization, and carbon emission monitoring. In order to better respond to the policy requirements of carbon emission reduction and environmental protection, shipping companies need more comprehensive and timely information support to make scientific decisions in ship management, fuel optimization, and risk control. The emergence of large language models (LLMs) provides a new solution for the construction of knowledge graphs for unstructured maritime data. These technologies can automatically identify and structure key entities and their relationships based on the processing of a large number of documents, and realize efficient retrieval and association analysis of massive information. However, there are still many challenges in building high-quality knowledge graphs for documents, mainly including:

[0003] 1. Data quality: Maritime data usually comes from complex sources, and there are problems such as inconsistent data quality, duplication and missing data, which affect the accuracy of knowledge extraction;

[0004] 2. Difficulty of information extraction: Unstructured documents have diverse data types and involve complex professional terms. Extracting accurate entities and relationships requires a high level of natural language processing capabilities;

[0005] 3. Real-time update requirements: Maritime information is updated frequently, and the knowledge graph needs to have the ability to update and dynamically maintain in real time to ensure the timeliness and accuracy of the information.

[0006] Based on the above background, how to accurately extract information based on maritime data from complex sources and construct high-quality knowledge graphs has become an urgent problem that needs to be solved. Summary of the invention

[0007] In order to solve the technical problem of how to accurately extract information based on maritime data with complex sources and construct a high-quality knowledge graph, the present invention proposes a method and system for constructing a maritime document knowledge graph, which realizes the accurate extraction of information from complex maritime information and constructs a set of high-quality knowledge graphs.

[0008] The present invention provides the following specific solutions:

[0009] In a first aspect, a method for constructing a maritime document knowledge graph includes:

[0010] S1. Data collection and preprocessing: Collect maritime data, including ship operation records, technical documents, laws and regulations, and news reports;

[0011] S2. Text segmentation: set the segmentation length, segment the maritime field data into new text units according to the segmentation length, and obtain a text unit list;

[0012] S3. Extraction of key information: Use a large language model to extract entity information, relationship information and event information from each text unit in the text unit list described in S2, and store the extracted information in a structured manner to obtain entity data, relationship data and event data; the entity information includes name, type and entity description, the relationship information includes source entity, target entity, relationship weight and relationship description, and the event information includes event initiator, event type, status, start date, end date and cause description; wherein the name is used as node data of the graph, the entity description is used as a node attribute, the directed or undirected connection from the source entity to the target entity is used as an edge of the graph, the relationship weight is used as the weight of the edge, the relationship description is used as an edge attribute, and the event information is used to enrich the attributes of the node or as an additional edge;

[0013] S4. Generate key information summary based on large language model: Based on the entity data and the relationship data, store different entity descriptions of the same entity information in different text units in a first description list, and store different relationship descriptions of the same relationship information in a second description list; and use a large language model to summarize the descriptions of the same entity in the first description list and generate an entity description summary, and summarize the descriptions of the same relationship in the second description list and generate a relationship description summary; use the entity description summary and the relationship description summary to optimize the attributes of the node data and the edge respectively;

[0014] S5. Community classification: clustering the entity data, relationship data and event data, attributing the entity data to a specific community, completing the grouping of node data, and obtaining a community classification result;

[0015] S6. Generate community report: Based on the community classification result, generate a community report for each community, generate community description information of the community by a large language model, and use the community description information to enrich the node attributes; the community report includes entity information, relationship information and event information of the community;

[0016] S7. Vectorization: Vectorize the entity data, relationship data and community description information in the same community to generate a vector representation that can represent its semantic content, that is, form semantic embedding;

[0017] S8. Construct a knowledge graph: construct a knowledge graph based on the node data, node attributes, edges, edge weights, edge attributes and the vector representation described in S7.

[0018] Preferably, the construction method S8 also includes: using a preset large language model to convert each of the text units into a vector representation, wherein the vector representation is used to capture the semantic information of the text unit, generate semantic embedding, and provide basic vectorization support for graph retrieval.

[0019] Preferably, if the maritime field data in S1 is unstructured data, it is pre-processed and then input into S2; the unstructured data is pre-processed, including: removing irrelevant content, formatting and cleaning data.

[0020] Preferably, the information integration step in S4 further includes: using a large language model to uniformly process the names of the same entity information in different text units based on a synonym matching method.

[0021] Preferably, the communities in the community classification result in S5 are arranged in different levels.

[0022] In the second aspect, a maritime document knowledge graph construction system includes: a data collection and processing module, a text segmentation module, an information extraction module, an information integration and summary module, a community classification module, a community summary generation module, a vectorization processing module and a graph generation module;

[0023] The data acquisition and processing module includes: a data acquisition submodule for collecting maritime data, and a data preprocessing submodule for preprocessing the maritime data, removing irrelevant content, formatting and cleaning the data, and obtaining the data to be processed; the maritime data includes ship operation records, technical documents, laws and regulations, and news reports;

[0024] The text segmentation module is used to set a segmentation length, and to segment the data to be processed into new text units according to the segmentation length to obtain a text unit list;

[0025] The information extraction module comprises: an information extraction module for extracting entity information, relationship information and event information from each of the text units based on a large language model, and for storing the information of the information extraction module in a structured manner to obtain a structured storage module for entity data, relationship data and event data; the entity information comprises a name, a type and an entity description, the relationship information comprises a source entity, a target entity and a relationship description, and the event information comprises an event initiator, an event type, a status, a start date, an end date and a cause description; wherein the name is used as the node data of the graph, the entity description is used as the node attribute, the directed or undirected connection from the source entity to the target entity is used as the edge of the graph, the relationship weight is used as the weight of the edge, the relationship description is used as the attribute of the edge, and the event information is used to enrich the attributes of the node or as an additional edge;

[0026] The information integration and summarization module includes: based on the entity data and the relationship data, an information integration submodule stores different entity descriptions of the same entity information in a first description list and different relationship descriptions of the same relationship information in a second description list in different text units; a description summary submodule that uses a large language model to summarize all descriptions of the same entity in the first description list and generate an entity description summary, summarizes all descriptions of the same relationship in the second description list and generates a relationship description summary, and uses the entity description summary and the relationship description summary to optimize the attributes of node data and edges respectively;

[0027] The community classification module is used to cluster the entity data, relationship data and event data, attribute the entity data to a specific community, complete the grouping of node data, and generate a community classification result;

[0028] The community summary generation module generates a community summary for each community based on the community classification result, generates community description information of the community by a large language model, and enriches the node attributes by using the community description information; the community report includes entity information, relationship information and event information of the community;

[0029] The vectorization processing module vectorizes entity data, relationship data and community description information in the same community to generate a vector representation that can represent its semantic content, that is, to form a semantic embedding;

[0030] The graph generation module constructs a knowledge graph based on the node data, node attributes, edges, edge weights, edge attributes and the vector representation of S7.

[0031] Beneficial effects of the present invention:

[0032] The present invention provides a method and system for constructing a maritime document knowledge graph. By collecting and preprocessing maritime documents, text segmentation is performed, and key information is extracted, information is integrated and descriptions are summarized using a large language model to obtain the node data of the graph, the edges of the graph, and the attributes of the nodes and edges; through community classification, generation of community summaries and vectorization processing, the above data are grouped and optimized, and finally integrated to construct a structured knowledge graph. Maritime data from complex sources are collected and preprocessed into a standardized format, providing a high-quality data foundation for subsequent data processing; a large language model is used to automatically identify and extract key entities and their relationships, reducing redundant information and improving the accuracy and refinement of knowledge; community classification helps to reveal the potential structure and relationship network between entities, providing a basis for efficient retrieval; vectorization processing generates vector representations that can represent semantic content, providing strong support for subsequent query, reasoning, analysis and visualization; after the implementation of the above steps, the problem of intelligent information extraction from real-time updated maritime data is solved, which significantly improves the accuracy and comprehensiveness of knowledge graph construction, effectively overcomes the limitations of traditional manual processing methods, and constructs an efficient retrieval knowledge graph, providing strong support for shipping companies' information retrieval, intelligent analysis and strategic decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A flowchart of a method for constructing a maritime document knowledge graph provided by the present invention.

[0034] Figure 2 A schematic diagram of a graph structure in a method for constructing a maritime document knowledge graph provided by the present invention.

[0035] Figure 3 A framework diagram of a maritime document knowledge graph construction system provided by the present invention. DETAILED DESCRIPTION

[0036] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0037] Figure 1 The present invention provides a flowchart of a method for constructing a maritime document knowledge graph. Figure 1 As shown, a method for constructing a maritime document knowledge graph includes:

[0038] S1. Data collection and preprocessing: Collect maritime data, including ship operation records, technical documents, laws and regulations, and news reports.

[0039] Specifically, if the maritime data in S1 is unstructured data, it is pre-processed and then input into S2; the unstructured data is pre-processed, including: removing irrelevant content, formatting and cleaning data.

[0040] Exemplarily, first, a collection of resolutions and guidelines on ship carbon intensity of the International Maritime Organization (IMO) is collected. The file is in PDF format and converted into a CII-related file in txt format named "CII.txt". During the conversion process, the document content is preprocessed to remove irrelevant content, delete duplicate records, correct erroneous data and unify the format, and then stored in the directory of the input file for future use.

[0041] S2. Text segmentation: set segmentation length, segment the maritime field data into new text units according to the segmentation length, and obtain a text unit list.

[0042] For example, the segmentation length is set to 300 tokens, and the content of the file "CII.txt" is segmented into several text units according to the length of 300 tokens, and a text unit list is obtained for subsequent use. In this embodiment, the Python-based GraphRAG toolkit is used to store the text in the form of Parquet files after segmentation. The specific structure is:

[0043] Text unit list create_base_text_units

[0044] id: unique identifier;

[0045] chunk: text block content;

[0046] chunk_id: unique identifier of the text chunk;

[0047] document_ids: a list of document IDs associated with the text block;

[0048] n_tokens: The number of tokens in the text block.

[0049] Documents are associated with text units create_base_documents

[0050] id: A unique identifier for the document.

[0051] text_unit_ids: List of text unit IDs.

[0052] raw_content: Original document content.

[0053] title: Document title.

[0054] S3. Extraction of key information: Use a large language model to extract entity information, relationship information and event information from each text unit in the text unit list described in S2, and store the extracted information in a structured manner to obtain entity data, relationship data and event data; the entity information includes name, type and entity description, the relationship information includes source entity, target entity, relationship weight and relationship description, and the event information includes event initiator, event type, status, start date, end date and cause description. Among them, the name is used as the node data of the graph, the entity description is used as the node attribute, the directed or undirected connection from the source entity to the target entity is used as the edge of the graph, the relationship weight is used as the weight of the edge, the relationship description is used as the attribute of the edge, and the event information is used to enrich the attributes of the node or as an additional edge.

[0055] Exemplarily, a large language model is used to extract entity information (such as CII guidelines, regulatory names), relationship information (such as "includes", "definition") and event information (such as regulatory releases, voyage adjustments) for each text unit. The extracted information is structured and stored to obtain entity data, relationship data and event data; wherein entity information includes name, type and entity description; relationship information includes source entity, target entity and relationship description; event information includes event initiator, event type, status, start date, end date and cause description. And establish an association relationship between entity data and relationship data and text units. Similarly, the above data is stored in the form of Parquet files, and the specific structure is:

[0056] Entity data create_final_entities

[0057] id: A unique identifier for the entity.

[0058] name: entity name;

[0059] type: entity type;

[0060] description: entity description;

[0061] human_readable_id: human readable entity ID;

[0062] graph_embedding: graph embedding information;

[0063] text_unit_ids: list of text unit IDs associated with the entity;

[0064] description_embedding: Description of embedding information.

[0065] Entity and text unit association join_text_units_to_entity_ids

[0066] text_unit_ids: text unit ID;

[0067] entity_ids: Entity IDs.

[0068] Relationship data create_final_relationships

[0069] source: the source entity of the relationship;

[0070] target: the target entity of the relationship;

[0071] weight: relationship weight;

[0072] description: Relationship description;

[0073] text_unit_ids: list of text unit IDs associated with the relationship;

[0074] id: unique identifier of the relationship;

[0075] human_readable_id: human readable relationship ID;

[0076] source_degree: the degree of the source entity;

[0077] target_degree: the degree of the target entity;

[0078] rank: Relationship ranking.

[0079] join_text_units_to_relationship_ids

[0080] id: text unit ID;

[0081] relationship_ids: relationship IDs.

[0082] create_final_text_units

[0083] id: A unique identifier for the text unit.

[0084] text: The content of the text unit.

[0085] n_tokens: The number of tokens in a text unit.

[0086] document_ids: A list of document IDs associated with the text unit.

[0087] entity_ids: A list of entity IDs associated with the text unit.

[0088] relationship_ids: A list of relationship IDs associated with the text unit.

[0089] In the above data, the entity name is used as the node data of the graph, the entity description is used as the node attribute, the directed or undirected connection from the source entity to the target entity is used as the edge of the graph, the relationship weight is used as the edge weight, and the relationship description is used as the edge attribute.

[0090] S4. Information integration and description summary: Based on the entity data and the relationship data, different entity descriptions of the same entity information in different text units are stored in a first description list, and different relationship descriptions of the same relationship information are stored in a second description list; and a large language model is used to summarize the descriptions of the same entity in the first description list and generate an entity description summary, and to summarize the descriptions of the same relationship in the second description list and generate a relationship description summary.

[0091] Specifically, a large language model is adopted to uniformly process the names of the same entity information in different text units based on a synonym matching method.

[0092] Exemplarily, for the above-mentioned entity data and relationship data, different text units provide different descriptions for the same entity, and these descriptions are stored as an entity description list, namely, a first description list, to ensure the integrity of the description information. Different relationship descriptions of different text units are all saved in the second description table of the description column to support a multi-dimensional understanding of the relationship. In the process of information integration, a large language model is also used to unify the names of the same entity information in different text units based on a synonym matching method. For example, "International Maritime Organization" and "IMO" are unified as "International Maritime Organization" to simplify the unification and maintenance of subsequent relationships.

[0093] Then, a large language model is used to summarize the descriptions of the same entity in the generated first description list and generate an entity description summary, and to summarize the descriptions of the same relationship in the generated second description list to generate a relationship description summary, and the entity description summary and relationship description summary are used to optimize the attributes of the node data and edges respectively.

[0094] S5. Community classification: cluster the entity data, relationship data and event data, attribute the entity data to a specific community, complete the grouping of node data, and obtain the community classification result.

[0095] Specifically, the communities in the community classification result described in S5 are set at different levels.

[0096] Exemplarily, clustering algorithms are used to cluster the entity data, relationship data, and event data obtained above, and the corresponding entities are assigned to the same community, and finally the community classification result is obtained; for example, entities related to CII calculation are assigned to a community, which can reveal the internal structure and connection of CII document content and store it in the form of Parquet files, specifically including:

[0097] Community data create_final_communities

[0098] id: community ID;

[0099] title: community title;

[0100] level: community level;

[0101] raw_community: original community ID;

[0102] relationship_ids: a list of relationship IDs associated with the community;

[0103] text_unit_ids: List of text unit IDs associated with the community.

[0104] Community classification data create_final_nodes

[0105] level: node level;

[0106] title: node title;

[0107] type: node type;

[0108] description: node description;

[0109] source_id: node source ID;

[0110] community: the community to which the node belongs;

[0111] degree: node degree;

[0112] human_readable_id: human readable node ID;

[0113] id: unique identifier of the node;

[0114] size: node size;

[0115] graph_embedding: graph embedding information;

[0116] top_level_node_id: top-level node ID;

[0117] x: the x coordinate of the node in the graph;

[0118] y: The y coordinate of the node in the graph.

[0119] S6. Generate a community report: Based on the community classification results, generate a community summary for each community, generate community description information of the community by a large language model, and use the community description information to enrich the node attributes; the community report includes entity information, relationship information and event information of the community.

[0120] Exemplarily, a large language model is used to generate a community summary containing entity information, relationship information, and event information for each community in the above community classification results. At the same time, community description information is generated for the community; and stored in the form of Parquet files, specifically including:

[0121] Community report data create_final_community_reports

[0122] community: community number;

[0123] full_content: the full content of the community;

[0124] level: community level;

[0125] rank: community ranking;

[0126] title: community title;

[0127] rank_explanation: community ranking explanation;

[0128] summary: Community summary;

[0129] findings: community discovery;

[0130] full_content_json: full content JSON representation of the community;

[0131] id: Unique identifier of the community.

[0132] S7. Vectorization processing: Vectorize the entity data, relationship data and community description information within the same community to generate a vector representation that can represent its semantic content, that is, form semantic embedding.

[0133] For example, the entity data, relationship data and community description information generated above are converted into word vectors using a pre-trained large language model for vectorization processing; the relationship data between entities are converted into relationship vectors using graph embedding technology (such as Graph Embedding), which can represent the social interaction and connection strength of the entities; different weights can be assigned according to different types of relationships to more finely capture the relationship between entities. And stored in the form of Parquet files, specifically including:

[0134] create_base_extracted_entities

[0135] entity_graph: XML representation of the entity graph.

[0136] create_summarized_entities

[0137] entity_graph: XML representation of the abstract entity graph.

[0138] create_base_entity_graph

[0139] level: clustering level.

[0140] clustered_graph: XML representation of the clustered graph.

[0141] S8. Construct a knowledge graph: construct a knowledge graph based on the node data, node attributes, edges, edge weights, edge attributes and the vector representation described in S7.

[0142] Specifically, the above S8 also includes: using a preset large language model to convert each of the text units into a vector representation, wherein the vector representation is used to capture the semantic information of the text unit, generate semantic embedding, and provide basic vectorization support for graph retrieval.

[0143] For example, Figure 2As shown, all Parquet files in the specified directory, that is, the various Parquet file data stored above, are traversed, read and merged into a DataFrame, which contains all the information extracted from the CII document, that is, all nodes and edge information generated by entity data, relationship data, event data, first description list, first description result, entity description summary, relationship description summary, community classification result, and community description information. Then, clean the DataFrame, remove the null values, and convert the `source` and `target` columns to string types; use `networkx` to create a directed graph, convert each row of data in the DataFrame into an edge in the graph, which represents the entities and relationships in the CII document; use `networkx`'s layout algorithm to generate 3D coordinates, generate 3D trajectories of nodes and edges through Plotly, and add edge labels, which describe the entities and relationships in the CII document; save the final generated graph visualization result as an HTML file and display it in the browser. This result allows the content and relationships of the CII document to be intuitively understood and analyzed.

[0144] Figure 3 This is a framework diagram of a maritime document knowledge graph construction system provided by the present invention. Figure 3 As shown, a maritime document knowledge graph construction system includes: a data collection and processing module, a text segmentation module, an information extraction module, an information integration and summary module, a community classification module, a community summary generation module, a vectorization processing module and a graph generation module.

[0145] The data acquisition and processing module includes: a data acquisition submodule for collecting maritime data, and a data preprocessing submodule for preprocessing the maritime data, removing irrelevant content, formatting and cleaning data, and obtaining data to be processed; the maritime data includes ship operation records, technical documents, laws and regulations, and news reports.

[0146] The text segmentation module is used to set the segmentation length, segment the data to be processed into new text units according to the segmentation length, and obtain a text unit list.

[0147] The information extraction module comprises: an information extraction module for extracting entity information, relationship information and event information from each of the text units based on a large language model, and for storing the information of the information extraction module in a structured manner to obtain a structured storage module for entity data, relationship data and event data; the entity information comprises a name, a type and an entity description, the relationship information comprises a source entity, a target entity and a relationship description, and the event information comprises an event initiator, an event type, a status, a start date, an end date and a cause description; wherein the name is used as node data of a graph, the entity description is used as a node attribute, the directed or undirected connection from the source entity to the target entity is used as an edge of the graph, the relationship weight is used as an edge weight, the relationship description is used as an edge attribute, and the event information is used to enrich the attributes of the node or as an additional edge.

[0148] The information integration and summarization module includes: based on the entity data and the relationship data, the information integration submodule stores different entity descriptions of the same entity information in the first description list and different relationship descriptions of the same relationship information in the second description list in different text units; uses a large language model to summarize all descriptions of the same entity in the first description list into an entity description summary, summarizes all descriptions of the same relationship in the second description list to generate a relationship description summary, and uses the entity description summary and the relationship description summary to optimize the description summary submodule of the node data and the attributes of the edge respectively.

[0149] The community classification module is used to cluster the entity data, relationship data and event data, attribute the entity data to a specific community, complete the grouping of node data, and generate a community classification result.

[0150] A community summary generation module includes a community summary generation submodule for generating a community summary for each community based on the community classification result, and a community description information generation submodule for generating community description information for each community by a large language model and enriching the node attributes using the community description information; the community summary includes entity information, relationship information and event information of the community.

[0151] The vectorization processing module vectorizes the entity data, relationship data and community description information within the same community to generate a vector representation that can represent its semantic content, that is, to form a semantic embedding.

[0152] The graph generation module constructs a knowledge graph based on the node data, node attributes, edges, edge weights, edge attributes and the above-mentioned vector representations.

[0153] It should be noted that the above-described specific implementations can enable those skilled in the art to more fully understand the invention, but do not limit the invention in any way. Therefore, although this specification has described the invention in detail with reference to the drawings and embodiments, those skilled in the art should understand that the invention can still be modified or replaced by equivalents. In short, all technical solutions and improvements that do not deviate from the spirit and scope of the invention should be included in the protection scope of the patent for the invention.

Claims

1. A method for constructing a maritime document knowledge graph, characterized in that: include: S1. Data collection and preprocessing: Collect maritime data, including ship operation records, technical documents, laws and regulations, and news reports; S2. Text segmentation: set the segmentation length, segment the maritime field data into new text units according to the segmentation length, and obtain a text unit list; S3. Extraction of key information: Use a large language model to extract entity information, relationship information and event information from each text unit in the text unit list described in S2 as key information, and store the extracted key information in a structured manner to obtain entity data, relationship data and event data; use the name in the entity data as the node data of the graph, the entity description as the node attribute, and the directed or undirected connection from the source entity to the target entity as the edge of the graph; use the relationship weight in the relationship data as the edge weight, the relationship description as the edge attribute, and the event information is used to enrich the node attributes or as an additional edge; S4. Generate key information summary based on large language model: based on the entity data and the relationship data, store the entity data and relationship data in different text units in different lists and mark them; And use the large language model to generate key information summaries and optimize the attributes of node data and edges, the key information summaries include: entity description summaries and relationship description summaries; S5. Community classification: clustering the entity data, relationship data and event data, attributing the entity data to a specific community, completing the grouping of node data, and obtaining a community classification result; S6. Generate community summary: Based on the community classification result, generate a community summary for each community, the community summary includes entity information, relationship information and event information of each community; S7. Vectorization: vectorize the entity information, relationship information and community description information in the same community to generate a vector representation of each community; S8. Construct a knowledge graph: Call the node data, node attributes, edges, edge weights, edge attributes and vector representations generated by steps S3-S7 to construct a knowledge graph.

2. The method for constructing a maritime document knowledge graph according to claim 1, characterized in that: The construction method S8 also includes: using a preset large language model to convert each of the text units into a vector representation, wherein the vector representation is used to capture the semantic information of the text unit, generate semantic embedding, and provide basic vectorization support for graph retrieval.

3. The method for constructing a maritime document knowledge graph according to claim 1, characterized in that: If the maritime domain data described in S1 is unstructured data, it is pre-processed and then input into S2; The unstructured data is preprocessed, including: removing irrelevant content, formatting and cleaning data.

4. A method for constructing a maritime document knowledge graph according to claim 1 or 3, characterized in that: The entity information includes name, type and entity description, the relationship information includes source entity, target entity, relationship weight and relationship description, and the event information includes event initiator, event type, status, start date, end date and cause description.

5. A method for constructing a maritime document knowledge graph according to claim 1 or 3, characterized in that: The method for generating entity description summary and relationship description summary in S4 is: in different text units, different entity descriptions of the same entity information are stored in a first description list, and different relationship descriptions of the same relationship information are stored in a second description list; the descriptions of the same entity in the first description list are summarized and an entity description summary is generated, and the descriptions of the same relationship in the second description list are summarized and a relationship description summary is generated; the entity description summary and relationship description summary are used to optimize the attributes of the node data and edges respectively.

6. A method for constructing a maritime document knowledge graph according to claim 1 or 3, characterized in that: The information integration step in S4 also includes: using a large language model to unify the names of the same entity information in different text units based on a synonym matching method.

7. The method for constructing a maritime document knowledge graph according to claim 1, characterized in that: The community settings in the community classification result in S5 are of different levels; the next level of the community summary also includes community description information of the community generated by the large language model; The node attributes are enriched using the community description information.

8. A maritime document knowledge graph construction system using the method described in any one of claims 1 to 7, comprising: Data collection and processing module, text segmentation module, information extraction module, information integration and summary module, community classification module, community summary generation module, vectorization processing module and graph generation module; The data acquisition and processing module includes: a data acquisition submodule for collecting maritime data, and a data preprocessing submodule for preprocessing the maritime data, removing irrelevant content, formatting and cleaning the data, and obtaining the data to be processed; the maritime data includes ship operation records, technical documents, laws and regulations, and news reports; The text segmentation module is used to set a segmentation length, and to segment the data to be processed into new text units according to the segmentation length to obtain a text unit list; The information extraction module comprises: an information extraction module for extracting entity information, relationship information and event information from each of the text units based on a large language model, and for storing the information of the information extraction module in a structured manner to obtain a structured storage module for entity data, relationship data and event data; the entity information comprises a name, a type and an entity description, the relationship information comprises a source entity, a target entity and a relationship description, and the event information comprises an event initiator, an event type, a status, a start date, an end date and a cause description; wherein the name is used as the node data of the graph, the entity description is used as the node attribute, the directed or undirected connection from the source entity to the target entity is used as the edge of the graph, the relationship weight is used as the weight of the edge, the relationship description is used as the attribute of the edge, and the event information is used to enrich the attributes of the node or as an additional edge; The information integration and summarization module includes: based on the entity data and the relationship data, an information integration submodule stores different entity descriptions of the same entity information in a first description list and different relationship descriptions of the same relationship information in a second description list in different text units; a description summary submodule that uses a large language model to summarize the descriptions of the same entity in the first description list and generate an entity description summary, summarizes the descriptions of the same relationship in the second description list and generates a relationship description summary, and uses the entity description summary and the relationship description summary to optimize the description summary of the node data and the attributes of the edge respectively; The community classification module is used to cluster the entity data, relationship data and event data, attribute the entity data to a specific community, complete the grouping of node data, and generate a community classification result; The community summary generation module includes a community summary generation submodule for generating a community summary for each community based on the community classification result, and a community description information generation submodule for generating community description information for each community by a large language model and enriching the node attributes with the community description information; the community summary includes entity information, relationship information and event information of the community; The vectorization processing module vectorizes entity data, relationship data and community description information in the same community to generate a vector representation that can represent its semantic content, that is, to form a semantic embedding; The graph generation module constructs a knowledge graph based on the node data, node attributes, edges, edge weights, edge attributes and the vector representation of S7.

Citation Information

Patent Citations

  • Security tool knowledge graph construction method and device for open source community

    CN113901466A

  • Information extraction and knowledge graph construction system and method

    CN117313850A

  • Prompt text generation method and system, computer equipment and storage medium

    CN118012990A

  • Method and system for generating questions and answers based on knowledge graph

    KR102697127B1

  • Methods and systems for automated generation of personalized messages

    US20200065857A1

Cited By

  • Fault library modeling method, device and product based on decision chain atlas

    CN120494070A

  • Method, system and equipment for constructing integrated digital model of full-motion simulator and medium

    CN120563744A

  • A method, system, device and medium for constructing an integrated digital model of a full-motion simulator

    CN120563744B

  • Ultra-large file deep analysis method and system based on dynamic segmentation and knowledge graph

    CN120850991A

  • Matching processing method and device for maritime information

    CN121144491A