Knowledge graph-driven water transportation scientific research knowledge service method and system
Through the knowledge graph-driven method, stuttering, TF-IDF, K-means clustering algorithm and Bayesian classifier are used to extract and classify water transport scientific research literature word segmentation and keywords, and use Neo4j to build a knowledge graph, solving the fragmentation and simplicity of the knowledge service platform in the water operation industry, and achieving efficient knowledge retrieval and service.
Patent Information
- Application Number
- CN202311628073.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-03
AI Technical Summary
The existing knowledge service platform of the water operation industry has problems such as fragmentation of knowledge, lack of systematicity, and simple search services, which cannot meet users' needs for rapid and efficient acquisition and learning related knowledge resources.
Using a knowledge graph-driven method, the word segmentation and keyword extraction and classification of water transport research literature is realized through stuttering, TF-IDF, K-means clustering algorithm and Bayesian classifier, and the knowledge graph is constructed using the Neo4j graph database to build a knowledge service system based on the knowledge graph, and provide efficient knowledge retrieval and visualization services.
It has realized the efficient classification of water transport scientific research literature and the construction of knowledge graphs, improved the accuracy and efficiency of knowledge services, and met users' needs for rapid knowledge acquisition and learning.
Smart Images

Figure CN120086380A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining, and particularly relates to a knowledge graph-driven water transportation scientific research knowledge service method and system. Background Art
[0002] The water transportation industry is an important part of the modern integrated transportation system. Under the background of the digital age, the water transportation industry still faces many challenges in the knowledge processing and application of the increasingly large and continuously updated massive domain information data. The current resource organization in the water transportation industry still remains at the digital stage, and traditional databases have problems of being too complicated and lacking systematicness when the public retrieves knowledge. The main reason behind this is that current systems such as search engines can only provide users with fragmented knowledge, for example, individual URLs, texts, pictures or videos, and it is difficult to provide users with the associations between different pieces of knowledge and to construct a complete knowledge system for users. Therefore, how to use a knowledge graph to build a knowledge service model and achieve a deep match between knowledge resources and service requirements is the focus of future knowledge service research.
[0003] Knowledge services based on knowledge graphs are gradually being applied in various fields and have seen rapid development in resources in the fields of library and information science, agronomy, medicine, finance, etc. This fully demonstrates the great feasibility of knowledge graph technology in resource organization. In the early days, the construction of knowledge graphs was mainly based on structured data on encyclopedia websites and was built using the Resource Description Framework standard (RDF). For example, the knowledge base YAGO containing spatio-temporal knowledge, the open-source shared knowledge base Freebase, and the crowdsourced open-source DBpedia, etc. The Knowledge Vault knowledge base published by Google in 2014 used natural language processing (NLP) tools to perform named entity recognition, part-of-speech tagging, and entity linking on natural language, transforming unstructured text information into structured information, which promoted the application of knowledge graphs in various fields. Mohamed S.K. et al. in the "Biological applications of knowledge graph embedding models" published in 2021 constructed a biological knowledge graph with a high-precision and highly scalable mining method and conducted a comparative study with previous knowledge graph construction methods. Ko H. et al. in the "Machine learning and knowledge graph based design rule construction for additive manufacturing" published in 2021 proposed a new method for 3D printing rules based on machine learning and knowledge graphs. Koho M. et al. in the "WarSampo knowledge graph: Finland in the second world war as linked open data" published in 2021 constructed the WarSampo knowledge graph, realizing the application of knowledge graphs in the field of military history. Liu Jing et al. in the "Device fault diagnosis method driven by knowledge and data fusion" published in 2021 proposed a device fault diagnosis method driven by knowledge and data fusion, forming a fault graph diagnosis system by fusing mechanism knowledge or operation data, and demonstrating detailed fault information and similar faults. Jiang Zhihao et al. in the "Construction and application of combat target knowledge graph" published in 2020 proposed the basic architecture for constructing a combat target knowledge graph and realized the application of knowledge graphs in combat through main links such as entity extraction, relation extraction, attribute extraction, entity disambiguation, and anaphora resolution.
[0004] In the field of waterway transportation, literature such as "Knowledge Graph Analysis of Domestic Port Logistics Research Based on CiteSpace", "Research Status of Port Logistics in China", and "Visual Analysis of Research Hotspots and Evolution of Port Logistics at Home and Abroad" uses the CNKI database, bibliometric methods, and CiteSpace visualization tools to analyze and summarize relevant literature in the field of port logistics. Through visual analysis of spatio-temporal distribution characteristics, keyword co-occurrence, keyword clustering, and time zone knowledge graphs, the evolution and emerging trends of port logistics are explored. Literature such as "Review of Domestic Port Construction Research Based on Literature Visualization", "Knowledge Graph Analysis of Blockchain Technology Empowering Port Supply Chain Finance", and "Construction and Analysis of the Knowledge Graph of Shipping Economy Based on CiteSpace" uses the same method to mine and construct knowledge graphs for port construction, port supply chain finance, and shipping economy literature information, and analyzes the distribution of literature publication years, research institution networks, research scholar networks, and research trend distributions, providing a reference for scientific research and engineering applications in inland shipping economy. The literature "Research on the Storage and Query System of Shipping Information in the Semantic Web" and the literature "Research on the Construction and Application of Shipping Linked Data under the Background of Open Government Data" use semantic web technology, introduce the RDF data model to define shipping information triples, and realize the semantic storage and query of shipping information, meeting the demand for enhancing the potential value of exploring shipping data and promoting the effective utilization of shipping data. The literature "Construction of the Knowledge Graph of Waterway Transportation of Dangerous Goods" uses a top-down approach to build the framework construction layer of the knowledge graph and fills the data layer with dangerous goods knowledge to construct the knowledge graph of waterway transportation of dangerous goods. The literature "Application of Knowledge Graph Technology in the Storage of Multiple Dangerous Goods in Ports" uses ontology representation to represent cargo knowledge and establish associations between knowledge. On this basis, implicit knowledge is further mined through custom knowledge inference rules to construct the knowledge graph of the storage of multiple dangerous goods in ports. The literature "Method for Constructing the Knowledge Graph of Ship On-site Supervision Operations" proposes a knowledge graph construction method that combines top-down and bottom-up approaches, uses a sequence annotation model and web crawler technology for entity recognition and knowledge extraction, fuses domain knowledge through a binary classification model, and uses the graph database Neo4j for knowledge storage and visualization.
[0005] The above studies construct knowledge graphs of port logistics, shipping economy, dangerous goods, and ships based on methods such as CiteSpace visualization tools, semantic web technology, RDF data model, and ontology representation, realizing the research on the development trend of the water transportation industry. However, these studies all use existing visualization tools to construct knowledge graphs, resulting in deficiencies in aspects such as the breadth, depth, adaptability, and flexibility of knowledge services. In addition, the information service mode of existing knowledge service platforms only provides simple retrieval services, which can no longer meet the needs and experiences of users to quickly and efficiently obtain and learn relevant knowledge resources and reduce the cost of knowledge search and discovery. Summary of the Invention
[0006] The object of the present invention is to overcome the problems and deficiencies existing in the above-mentioned prior art, and provide a knowledge graph-driven water transportation scientific research knowledge service method and system. The key is to use Jieba word segmentation, TF-IDF (term frequency–inverse document frequency) statistical algorithm, K-means clustering algorithm, and Bayesian classifier to realize word segmentation, keyword extraction and classification of water transportation scientific research literature knowledge, and then use Neo4j graph database to store water transportation scientific research literature knowledge, and construct a knowledge service system based on knowledge graph to meet the needs of user knowledge services.
[0007] To solve the above technical problems, a knowledge graph-driven water transportation scientific research knowledge service method provided by the technical solution of the present invention includes the following steps:
[0008] Step 1: Collect water transportation scientific research literature related to the water transportation industry and perform preprocessing;
[0009] Step 2: Perform word segmentation processing on the preprocessed water transportation scientific research literature; extract the subject words whose frequencies of occurrence after word segmentation are greater than the set value, and fuse them with the keywords of the water transportation scientific research literature; use clustering methods to cluster the fused keywords to obtain the literature categories of the water transportation scientific research literature;
[0010] Step 3: Use Neo4j graph database to construct a knowledge graph including literature information, author information and literature categories of water transportation scientific research literature;
[0011] Step 4: Build a service framework based on literature information, author information and literature categories, use Echarts data visualization library for visualization, and realize knowledge retrieval by users in the form of human-computer dialogue.
[0012] As an improvement of the above method, the step 1 includes:
[0013] Step 1.1: Collect water transportation scientific research literature through search keywords: Through search keywords, initially collect water transportation scientific research literature related to the water transportation industry;
[0014] Step 1.2: Judge whether the initially collected water transportation scientific research literature is unstructured data. If it is, the collection is completed. If not, execute step 1.3;
[0015] Step 1.3: Preprocess the water transport research literature: Among the initially collected water transport research literature, filter out the water transport research literature whose relevance to the water transport industry is lower than the preset relevance; among the initially collected water transport research literature, filter out the water transport research literature whose author address, keywords and abstract fields are missing, so as to collect water transport industry-related literature including title, author, institution, journal, keyword, abstract, citation volume and download volume information.
[0016] As an improvement of the above method, the search keywords include: water transport, water transport industry, waterway, port, shipping, lock, ship industry, maritime affairs and shipping.
[0017] As an improvement of the above method, step 2 comprises:
[0018] Step 2.1: Use the Jieba word segmentation to segment the abstract information of the preprocessed water transport scientific research literature; use the TF-IDF technology to extract the subject words with a frequency greater than the set value, and merge them with the keywords of the water transport scientific research literature to form a keyword document;
[0019] Step 2.2: Use K-means clustering method to cluster the fused keywords to obtain the document category of water transport scientific research literature, and save the document category to the category information document;
[0020] Step 2.3: Match the keywords in the category information document with the keywords in the abstract fields from different water transport scientific research documents to achieve classification of different water transport scientific research documents.
[0021] As an improvement of the above method, step 3 comprises:
[0022] Step 3.1: Extract the author information of water transport scientific research documents and save it to the author information file; extract the document information of water transport scientific research documents and save it to the document information file;
[0023] Step 3.2: Import the category information documents, author information documents and document information documents of water transport scientific research literature into the Neo4j graph database to realize the construction, retrieval, query and modification of the knowledge graph of water transport scientific research literature.
[0024] As an improvement of the above method, the document information includes: document name, keywords, abstract and publication time information.
[0025] As an improvement of the above method, step 4 includes:
[0026] Step 4.1: Build a question library based on document information, author information, and document categories, and use a Bayesian classifier to classify input questions, thereby realizing knowledge services;
[0027] Step 4.2: Build a knowledge Q&A system using the Flask framework, use the Neo4j graph database as the database, perform visualization using the Echarts data visualization library, and realize the retrieval of knowledge by users in the form of human-computer dialogue.
[0028] To achieve another object of the present invention, the present invention also provides a knowledge graph-driven water transportation scientific research knowledge service system, including:
[0029] A collection module, which is used to collect water transportation scientific research literature related to the water transportation industry and perform preprocessing;
[0030] A classification module, which uses the Jieba word segmentation technology to perform word segmentation on the preprocessed water transportation scientific research literature; uses the TF-IDF technology to extract the subject words with a frequency greater than a set value after word segmentation, and fuse them with the keywords of the water transportation scientific research literature; uses a clustering method to cluster the fused keywords to obtain the literature categories of the water transportation scientific research literature;
[0031] A knowledge graph construction module, which uses the Neo4j graph database to construct a knowledge graph including the literature information, author information, and literature categories of water transportation scientific research literature; and
[0032] A service module, which builds a service framework based on literature information, author information, and literature categories, performs visualization using the Echarts data visualization library, and realizes the retrieval of knowledge by users in the form of human-computer dialogue.
[0033] The present invention has the following advantages:
[0034] 1. Use Jieba word segmentation to segment the literature abstract, then extract high-frequency words through TF-IDF, and expand the water transportation industry category information through fusion with the literature keywords.
[0035] 2. Use the K-means clustering method to cluster the keywords, realize the construction of the literature category document, use the Bayesian classifier to classify the input question, and then improve the accuracy of question matching.
[0036] 3. Construct a knowledge graph based on author, category, and literature information, realize the visualization of the knowledge graph through Neo4j, transform the knowledge graph into a Q&A form, and realize the retrieval of knowledge through human-computer interaction, meet the user's knowledge needs, and improve the accuracy and efficiency of knowledge services.
[0037] Combining the above three points, a knowledge graph-driven water transportation scientific research knowledge service method and system adopted by the present invention can realize knowledge services by constructing a knowledge graph, thereby improving the performance of knowledge services. Description of the Drawings
[0038] Figure 1 It is a schematic flowchart of the knowledge graph-driven water transportation scientific research knowledge service method provided by the present invention. Specific implementation manners
[0039] The following further illustrates the technical solution provided by the present invention in conjunction with embodiments.
[0040] Embodiment 1
[0041] Figure 1 It is a schematic flowchart of a knowledge graph-driven water transportation scientific research knowledge service method of the present invention.
[0042] The water transportation scientific research knowledge service method based on a knowledge graph in this embodiment uses a crawler tool to crawl the literature related to waterway transportation on CNKI. The items crawled mainly include information such as the literature name, author, publication time, author's unit, keywords, abstract, and publication journal. The jieba word segmentation technology is used to segment the abstract information, and then the TF-IDF technology is used to extract the frequently occurring topic words, and the topic words and keywords are fused to obtain category keywords; the obtained category keywords are analyzed by the K-means clustering method to obtain different categories of keywords for classifying the literature; information extraction is performed on the crawled data to construct category, author, and literature title documents, which are imported into Neo4j to construct knowledge graphs such as category-literature and author-literature; a question-answering knowledge base is constructed, the question category is determined, the question is classified based on a Bayesian classifier to achieve question-answer matching; finally, a knowledge service question-answering system is constructed based on the Flask framework, with Neo4j as the database, and the Echarts data visualization library is used for visualization.
[0043] The method provided in this embodiment includes the following steps:
[0044] Step 1: Filter the obtained water transportation scientific research literature information and preprocess the filtered water transportation scientific research literature information;
[0045] Step 2: Use the jieba word segmentation technology to segment the abstract information of the water transportation scientific research literature, use the TF-IDF technology to extract the keywords with more frequent occurrences after word segmentation, fuse the keywords in combination with the keywords of the literature, and then use the clustering method for classification to obtain the type information of the keywords, and then use these category information to classify the literature to obtain literature information belonging to different categories;
[0046] Step 3: Process the data, construct data information of category, author, and literature title information, store the data information in a.csv file, import the constructed data file into Neo4j, and use python programming to implement the construction, retrieval, query, and modification of the knowledge graphs of category, author, and literature title information.
[0047] Step 4: Build a Q&A knowledge base related to authors, categories, literature, etc., use a Bayesian classifier to classify the input knowledge, and then realize the retrieval of knowledge to meet the user's needs for knowledge services.
[0048] Preferably, the detailed steps of Step 1 are as follows:
[0049] Step 1.1: Filter the information of water transportation research literature. Mainly search for literature related to the water transportation industry using keywords such as waterway transportation, water transportation industry, waterway, port, shipping, ship lock, shipbuilding industry, maritime affairs, and sea transportation.
[0050] Filter out the literature with low relevance to the water transportation industry and missing fields such as author address, keywords, and abstracts. Only retain journal literature from sources such as SCI, EI, Peking University Core, CSSCI, and CSCD to obtain water transportation industry literature including main data information such as title, author, institution, journal, keywords, abstract, citation volume, and download volume.
[0051] Preferably, the detailed steps of Step 2 are as follows:
[0052] Step 2.1: For the filtered data, use Jieba segmentation to segment the abstract information, and then use TF-IDF to extract the words with higher occurrence frequencies as supplementary keywords. Integrate the extracted keywords with the original keywords in the literature to form a keyword document.
[0053] Step 2.2: Use the K-means algorithm to cluster the keywords to obtain the categories of water transportation industry research information and save them to the category information document;
[0054] Step 2.3: Match the keywords in the category information with the abstract keywords from different literatures to realize the classification of different literatures.
[0055] Preferably, the detailed steps of Step 3 are as follows:
[0056] Step 3.1: Extract the data, extract the author information and literature information respectively. The literature information includes information such as literature name, keywords, abstract, and publication time. Save the extracted author and literature information to.csv files respectively.
[0057] Step 3.2: Import the category, author, and literature documents constructed in Step 3.1 into Neo4j, start Neo4j, write code to realize the construction, retrieval, query, and modification of the knowledge graph of category, author, and literature title information.
[0058] Preferably, the detailed steps of Step 4 are as follows:
[0059] Step 4.1: Construct a question bank based on authors, categories, literature, etc., and use a Bayesian classifier to classify the input questions, so as to realize knowledge services.
[0060] Step 4.2: Build a knowledge Q&A system using the Flask framework, with Neo4j as the database, visualize the data using Echarts, and realize the retrieval of knowledge by users in the form of human-computer dialogue.
[0061] Embodiment 2
[0062] A knowledge graph-driven water transportation scientific research knowledge service system, comprising:
[0063] A collection module, used to collect water transportation scientific research literature related to the water transportation industry and perform preprocessing;
[0064] A classification module, which uses the Jieba word segmentation technology to segment the preprocessed water transportation scientific research literature; uses the TF-IDF technology to extract the subject words with a frequency greater than a set value after word segmentation, and fuse them with the keywords of the water transportation scientific research literature; uses a clustering method to cluster the fused keywords to obtain the literature categories of the water transportation scientific research literature;
[0065] A knowledge graph construction module, which uses the Neo4j graph database to construct a knowledge graph including the literature information, author information, and literature categories of water transportation scientific research literature; and
[0066] A service module, which builds a service framework based on literature information, author information, and literature categories, visualizes it using the Echarts data visualization library, and realizes the retrieval of knowledge by users in the form of human-computer dialogue.
[0067] In summary, a knowledge graph-driven water transportation scientific research knowledge service method and system provided by the present invention mainly consists of two parts: one is to construct a water transportation scientific research knowledge graph; the other is to use the water transportation scientific research knowledge graph to realize knowledge services. Construct a water transportation scientific research knowledge graph, obtain water transportation scientific research literature knowledge, use technologies such as Jieba word segmentation technology, TF-IDF, and K-means algorithm to extract subject words, classify the literature, and construct knowledge graphs of literature titles, authors, journals, etc. For the part of realizing knowledge services with the water transportation scientific research knowledge graph, build a Q&A knowledge base related to authors, categories, literature, etc., use a Bayesian classifier to classify the input knowledge, and then realize the retrieval of knowledge to meet the needs of users for knowledge services.
[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present invention does not depart from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A knowledge graph-driven water transportation scientific research knowledge service method, comprising the following steps: Step 1: Collect water transportation scientific research literature related to the water transportation industry and perform preprocessing; Step 2: Perform word segmentation on the preprocessed water transportation scientific research literature; Extract the subject words with a frequency greater than the set value after word segmentation, and fuse them with the keywords of the water transportation scientific research literature; Use the clustering method to cluster the fused keywords to obtain the literature categories of the water transportation scientific research literature; Step 3: Use the Neo4j graph database to construct a knowledge graph including the literature information, author information, and literature categories of the water transportation scientific research literature; Step 4: Build a service framework based on the literature information, author information, and literature categories, use the Echarts data visualization library for visualization, and realize the retrieval of knowledge by users in the form of human-computer dialogue.
2. The knowledge graph-driven water transportation scientific research knowledge service method according to claim 1, characterized in that, the said Step 1 includes: Step 1.1: Collect water transportation scientific research literature through search keywords: Through search keywords, preliminarily collect water transportation scientific research literature related to the water transportation industry; Step 1.2: Judge whether the preliminarily collected water transportation scientific research literature is unstructured data. If it is, the collection is completed. If not, execute Step 1.3; Step 1.3: Perform preprocessing on the water transportation scientific research literature: In the preliminarily collected water transportation scientific research literature, filter out the water transportation scientific research literature with a relevance lower than the preset relevance to the water transportation industry. In the preliminarily collected water transportation scientific research literature, filter out the water transportation scientific research literature with missing author addresses, keywords, and abstract fields, so as to collect and obtain water transportation industry-related water transportation industry literature including information such as title, author, institution, journal, keyword, abstract, citation volume, and download volume.
3. The knowledge graph-driven water transportation scientific research knowledge service method according to claim 2, characterized in that, the said search keywords include: waterway transportation, water transportation industry, waterway, port, shipping, ship lock, shipbuilding industry, maritime affairs, and sea transportation.
4. The knowledge graph-driven water transportation scientific research knowledge service method according to claim 1, characterized in that, the said Step 2 includes: Step 2.1: Use Jieba to perform word segmentation on the abstract information of the preprocessed water transportation scientific research literature; Use the TF-IDF technology to extract the subject words with a frequency greater than the set value, and fuse them with the keywords of the water transportation scientific research literature to form a keyword document; Step 2.2: Use the K-means clustering method to cluster the fused keywords to obtain the literature categories of the water transportation scientific research literature, and save the literature categories to the category information document; Step 2.3: Match the keywords in the category information document with the keywords in the abstract fields of different water transportation scientific research literatures to realize the classification of different water transportation scientific research literatures.
5. The knowledge graph-driven water transportation scientific research knowledge service method according to claim 4, characterized in that, the said Step 3 includes: Step 3.1: Extract the author information of the water transportation scientific research literature and save it to the author information document, and extract the literature information of the water transportation scientific research literature and save it to the literature information document; Step 3.2: Import the category information document, author information document, and literature information document of water transportation scientific research literature into the Neo4j graph database to realize the construction, retrieval, query, and modification of the knowledge graph of water transportation scientific research literature.
6. The knowledge graph-driven water transportation scientific research knowledge service method according to claim 5, wherein, the literature information includes: literature name, keywords, abstract, and publication time information.
7. The knowledge graph-driven water transportation scientific research knowledge service method according to claim 5, wherein, the step 4 includes: Step 4.1: Construct a question bank based on literature information, author information, and literature categories, and use a Bayesian classifier to classify the input questions, thereby realizing knowledge services; Step 4.2: Build a knowledge Q&A system using the Flask framework, use the Neo4j graph database as the database, use the Echarts data visualization library for visualization, and realize the retrieval of knowledge by users in the form of human-computer dialogue.
8. A knowledge graph-driven water transportation scientific research knowledge service system, wherein, it includes: a collection module, which is used to collect water transportation scientific research literature related to the water transportation industry and perform preprocessing; a classification module, which uses the Jieba word segmentation technology to perform word segmentation on the preprocessed water transportation scientific research literature; uses the TF-IDF technology to extract the subject words whose occurrence frequency after word segmentation is greater than the set value, and integrates them with the keywords of the water transportation scientific research literature; uses the clustering method to cluster the integrated keywords to obtain the literature categories of the water transportation scientific research literature; a knowledge graph construction module, which uses the Neo4j graph database to construct a knowledge graph including the literature information, author information, and literature categories of water transportation scientific research literature; and a service module, which builds a service framework based on literature information, author information, and literature categories, uses the Echarts data visualization library for visualization, and realizes the retrieval of knowledge by users in the form of human-computer dialogue.