Construction Method and System of Query Software System Based on Academic Knowledge Graph

By designing the schema of RDF and document database, combining Virtuoso and ElasticSearch, an academic knowledge graph query system was built, which solved the problem of the lack of a unified construction method of the academic knowledge graph system, and achieved low-cost and efficient development of the academic knowledge graph query system.

CN115757824BActive Publication Date: 2025-07-08SHANGHAI JIAOTONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211411474.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-11
Publication Date
2025-07-08
Estimated Expiration
2042-11-11

AI Technical Summary

Technical Problem

The existing academic knowledge graph system lacks a unified construction methodology, and the functions of each system are different, so it is impossible to effectively explore the knowledge in the knowledge graph. The existing patents only involve the construction of knowledge graphs in the medical field and do not involve the academic field.

Method used

Design the schema of RDF database and document database, use D2RQ tools to generate ttl files, store and query data through Virtuoso and ElasticSearch databases, build a back-end query module, and visualize it on the front-end, providing paper aggregation and content query functions.

Benefits of technology

It significantly reduces the cost and cycle of software development, has a simple and clear system architecture, and has a wide range of applications. It provides a general method for academic knowledge graph query system, and quickly iterates out displayable query systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115757824B_ABST
    Figure CN115757824B_ABST
Patent Text Reader

Abstract

The present invention provides a construction method and system for a query software system based on an academic knowledge graph, including the following steps: designing the schema of the RDF database; according to the designed schema, exporting corresponding data from the database and storing it in the RDF database Virtuoso; designing the schema of the document database; according to the designed schema, exporting the relevant document data and partial meta-information of the papers from the database and storing them in the document database ElasticSearch; constructing a backend query module according to the query capabilities provided by the cooperation of the above two databases; and completing the visual display of relevant functions at the front end according to the interfaces provided by the backend. The present invention significantly reduces the development cost and development cycle of the software, the system architecture is simple and clear, the use process is fast and convenient, the applicable range is relatively wide, and it can provide an effective way for the upper-layer application development of the academic knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer software, and in particular, to a method and system for constructing a query software system based on an academic knowledge graph. Background Art

[0002] With the continuous progress of science and technology, the number of papers as a carrier of knowledge has been increasing rapidly. In this process, many highly influential academic papers have emerged. Each paper cites other papers and is also cited by other papers. Regarding papers as nodes in a network and the citation relationship between papers as edges in the network, an academic network can be obtained. More generally, entities such as authors, institutions, scenarios (journals and conferences), topics, etc. can be regarded as nodes in the network. Through relationships such as (author, writes, paper) and (paper, published in, journal), with "writes" and "published in" as edges, a more extensive heterogeneous academic network can be formed. Such a network is called an academic knowledge graph. Currently, there are many systems based on academic knowledge graphs, such as Aminer, Acemap, etc. However, the construction of these systems is closed, with different underlying storage methods, and they do not exploit more knowledge in the academic knowledge graph.

[0003] Literature 1, the paper "GAKG: A Multimodal Geoscience Academic Knowledge Graph" conducts a deeper exploration of the academic knowledge graph of geoscience papers, extracting some features of geoscience, such as the geographical location of geoscience paper research and the ages of the research contents of rocks and water collected, etc., to construct more attributes and entities, enriching the academic knowledge graph of geoscience. However, at the application level, there is currently no methodology summarizing how a system based on an academic knowledge graph should be designed. The functions of the above-mentioned various systems are different, but there are obvious contents that can be learned from each other, which also indicates the lack of such a unified methodology.

[0004] The patent document with the publication number CN112820400A discloses a disease diagnosis method, device, and equipment based on knowledge reasoning of a medical knowledge graph. The method includes: obtaining the interaction content between a user and a medical Q&A system, and obtaining the user's symptom set based on the interaction content; according to the user's symptom set, calculating the first probability of the user's existence event; calculating the second probability of the co-occurrence of user S and disease di based on the first probability; calculating the third probability of the user suffering from disease di based on the first probability and the second probability; and outputting the user's final disease diagnosis result based on a dynamic threshold and the third probability. However, this patent document only relates to the construction of a method based on a knowledge graph in the medical field. Summary of the Invention

[0005] Aiming at the defects in the prior art, the purpose of the present invention is to provide a construction method and system of a query software system based on an academic knowledge graph.

[0006] A construction method of a query software system based on an academic knowledge graph provided by the present invention includes the following steps:

[0007] Step 1: Design the schema of the RDF database;

[0008] Step 2: According to the designed schema, export the corresponding data from the database and store it in the RDF database Virtuoso;

[0009] Step 3: Design the schema of the document database;

[0010] Step 4: According to the designed schema, export the relevant document data and part of the meta-information of the papers from the database and store it in the document database ElasticSearch;

[0011] Step 5: Construct a backend query module according to the query capabilities provided by the cooperation of the above two databases;

[0012] Step 6: Complete the visual display of relevant functions on the front end according to the interfaces provided by the backend.

[0013] Preferably, in the above Step 4, exporting the relevant document data and part of the meta-information of the papers from the database includes: the abstract of the paper, the title of the paper;

[0014] In the above Step 6, completing the visual display of relevant functions on the front end includes: paper aggregation function, content query function.

[0015] Preferably, in the above Step 1, according to the existing paper attributes in the database, construct the association relationship of triples and design in combination with the content to be queried.

[0016] Preferably, the above Step 2 includes the following steps:

[0017] Step 2.1: Use the D2RQ tool to generate and modify the corresponding ttl file according to the designed schema. The ttl file describes the attributes of the nodes in the schema and the relationships between the nodes;

[0018] Step 2.2: According to the ttl file, convert the data in the database into an RDF triple file in.nt format;

[0019] Step 2.3: Use docker to install the Virtuoso image;

[0020] Step 2.4: Start the Virtuoso image. When starting, mount the above-mentioned.nt file in the Virtuoso initialization data directory. After starting, use sparql to query Virtuoso to check whether the data is imported successfully.

[0021] Preferably, in step 3, the paper id of the database is used to associate the title and abstract of the paper, as well as some relevant information that needs to be frequently queried, to construct the schema required for the ElasticSearch database, and the design is carried out in combination with the content to be queried.

[0022] Preferably, in step 4, create the schema designed above in ElasticSearch and import the data into the ElasticSearch database in the schema format.

[0023] Preferably, in step 5, the backend query module includes at least a data acquisition unit and a data cache unit;

[0024] The data acquisition unit obtains data through the above two databases, uses the Virtuoso database when performing queries related to attributes or many-to-many relationships, and uses the ElasticSearch database when performing document-related queries;

[0025] The data cache unit caches the above data in memory for subsequent searches of the module.

[0026] The present invention also provides a construction system for a query software system based on an academic knowledge graph, including the following modules:

[0027] Module M1: Design the schema of the RDF database;

[0028] Module M2: According to the designed schema, export the corresponding data from the database and store it in the RDF database Virtuoso;

[0029] Module M3: Design the schema of the document database;

[0030] Module M4: According to the designed schema, export the relevant document data and some meta-information of the paper from the database and store it in the document database ElasticSearch;

[0031] Module M5: Construct a backend query module according to the query capabilities provided by the cooperation of the above two databases;

[0032] Module M6: Complete the visual display of relevant functions at the front end according to the interface provided by the backend.

[0033] Preferably, in the module M4, the relevant document data and some meta-information exported from the database for the papers include: the abstract of the paper and the title of the paper.

[0034] In the module M6, the visual display of relevant functions is completed at the front end, including: the paper aggregation function and the content query function.

[0035] Preferably, in the module M1, based on the existing paper attributes in the database, the associated relationship of triples is constructed and designed in combination with the content to be queried.

[0036] The module M2 includes the following modules:

[0037] Module M2.1: Use the D2RQ tool to generate and modify the corresponding ttl file according to the designed schema. The ttl file describes the attributes of the nodes in the schema and the relationships between the nodes.

[0038] Module M2.2: According to the ttl file, convert the data in the database into an.nt format RDF triple file.

[0039] Module M2.3: Use docker to install the Virtuoso image.

[0040] Module M2.4: Start the Virtuoso image. When starting, mount the above-mentioned.nt file in the Virtuoso initialization data directory. After starting, use sparql to query Virtuoso to check whether the data is imported successfully.

[0041] In the module M3, use the paper id in the database to associate the title and abstract of the paper, as well as some relevant information that needs to be frequently queried, to construct the schema required by the ElasticSearch database and design it in combination with the content to be queried.

[0042] In the module M4, create the above-designed schema in ElasticSearch and import the data into the ElasticSearch database in the schema format.

[0043] In the module M5, the backend query module includes at least a data acquisition unit and a data cache unit.

[0044] The data acquisition unit obtains data through the above two databases. When performing queries related to attributes or many-to-many relationships, it uses the Virtuoso database. When performing document-related queries, it uses the ElasticSearch database.

[0045] The above data caching unit caches the above data in the memory for subsequent lookup by the module.

[0046] Compared with the prior art, the present invention has the following beneficial effects:

[0047] 1. The present invention significantly reduces the software development cost and development cycle. The system architecture is simple and clear, and the usage process is fast and convenient.

[0048] 2. The present invention has a relatively wide range of applications and can provide an effective way for the development of upper-layer applications of academic knowledge graphs.

[0049] 3. The construction method of the query software system based on the academic knowledge graph of the present invention provides a general method for developing an academic knowledge graph query system, and a displayable query system can be quickly iterated according to this method. Brief Description of the Drawings

[0050] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:

[0051] Figure 1 It is a flowchart of the construction method of the query software system based on the academic knowledge graph of the present invention;

[0052] Figure 2 It is a system architecture diagram of the construction system of the query software system based on the academic knowledge graph of the present invention;

[0053] Figure 3 It is a front-end page display diagram of the construction system of the query software system based on the academic knowledge graph of the present invention;

[0054] Figure 4 It is a partial triple relationship display diagram of the TTL file. Detailed Description of the Embodiments

[0055] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0056] Example 1:

[0057] As Figure 1 shown, this embodiment provides a construction method of a query software system based on an academic knowledge graph, including the following steps:

[0058] Step 1: Design the schema of the RDF database; construct the association relationship of triples based on the existing paper attributes in the database, and design in combination with the content to be queried.

[0059] Step 2: Export the corresponding data from the database according to the designed schema and store it in the RDF database Virtuoso; Step 2 includes the following steps:

[0060] Step 2.1: Use the D2RQ tool to generate and modify the corresponding ttl file according to the designed schema. The ttl file describes the attributes of the nodes in the schema and the relationships between nodes and nodes.

[0061] Step 2.2: Convert the data in the database into an RDF triple file in.nt format according to the ttl file.

[0062] Step 2.3: Use docker to install the Virtuoso image.

[0063] Step 2.4: Start the Virtuoso image. When starting, mount the above.nt file in the Virtuoso initialization data directory. After starting, use sparql to query Virtuoso to check whether the data is imported successfully.

[0064] Step 3: Design the schema of the document database; use the paper id in the database to associate the title and abstract of the paper, as well as some relevant information that needs to be frequently queried, to construct the schema required by the ElasticSearch database, and design in combination with the content to be queried.

[0065] Step 4: Export the relevant document data and some meta-information of the papers from the database and store them in the document database ElasticSearch; Exporting the relevant document data and some meta-information of the papers from the database includes: the abstract of the paper, the title of the paper; create the above-designed schema in ElasticSearch and import the data into the ElasticSearch database according to the schema format.

[0066] Step 5: Construct the backend query module according to the query capabilities provided by the cooperation of the above two databases; The backend query module includes at least a data acquisition unit and a data cache unit.

[0067] The data acquisition unit obtains data through the above two databases. When performing queries related to attributes or many-to-many relationships, the Virtuoso database is used. When performing document-related queries, the ElasticSearch database is used; the data cache unit caches the above data in memory for subsequent lookups by the module.

[0068] Step 6: According to the interfaces provided by the backend, complete the visual display of relevant functions on the frontend; Completing the visual display of relevant functions on the frontend includes: paper aggregation function, content query function.

[0069] Example 2:

[0070] This embodiment provides a construction system for a query software system based on an academic knowledge graph, including the following modules:

[0071] Module M1: Design the schema of the RDF database; According to the existing paper attributes in the database, construct the association relationship of triples and design in combination with the content to be queried.

[0072] Module M2: According to the designed schema, export the corresponding data from the database and store it in the RDF database Virtuoso; Module M2 includes the following modules:

[0073] Module M2.1: Use the D2RQ tool to generate and modify the corresponding ttl file according to the designed schema. The ttl file describes the attributes of the nodes in the schema and the relationships between the nodes.

[0074] Module M2.2: According to the ttl file, convert the data in the database into an.nt format RDF triple file.

[0075] Module M2.3: Use docker to install the Virtuoso image.

[0076] Module M2.4: Start the Virtuoso image. When starting, mount the above.nt file in the Virtuoso initialization data directory. After starting, use sparql to query Virtuoso to check whether the data is imported successfully.

[0077] Module M3: Design the schema of the document database; Use the paper id in the database to associate the title and abstract of the paper, as well as some relevant information that needs to be frequently queried, to construct the schema required by the ElasticSearch database and design in combination with the content to be queried.

[0078] Module M4: Export the relevant document data and some meta-information of the thesis from the database according to the designed schema, and store them in the document database ElasticSearch; The relevant document data and some meta-information of the thesis exported from the database include: the abstract of the thesis, the title of the thesis; Create the above-designed schema in ElasticSearch, and import the data into the ElasticSearch database according to the schema format.

[0079] Module M5: Build a backend query module according to the query capabilities provided by the cooperation of the above two databases; The backend query module at least includes a data acquisition unit and a data caching unit;

[0080] The data acquisition unit obtains data through the above two databases, uses the Virtuoso database when performing queries related to attributes or many-to-many relationships, and uses the ElasticSearch database when performing document-related queries; The data caching unit caches the above data in memory for subsequent searches of the module.

[0081] Module M6: Complete the visual display of relevant functions on the front end according to the interfaces provided by the backend; The visual display of relevant functions on the front end includes: thesis aggregation function, content query function.

[0082] Example 3:

[0083] Those skilled in the art can understand this embodiment as a more specific description of Embodiment 1 and Embodiment 2.

[0084] This embodiment provides a construction method of a query software system based on an academic knowledge graph, including:

[0085] Step S1: Design the schema of the RDF database;

[0086] Step S2: Export the corresponding data from the database according to the designed schema, and store them in the RDF database Virtuoso;

[0087] Step S3: Design the schema of the document database;

[0088] Step S4: Export the corresponding data from the database according to the designed schema, and store them in the document database ElasticSearch;

[0089] Step S5: Build the backend query logic according to the query capabilities provided by the cooperation of the above two databases;

[0090] Step S6: According to the interfaces provided by the backend, complete the visual display of functions such as paper aggregation and content query on the front end.

[0091] The said step S1 includes:

[0092] Step S1.1: According to the existing paper attributes in the database, construct the association relationship of triples. For example, if a paper has the author attribute, then the triple (paper, is_written_by, author) needs to be constructed. This step needs to be designed in combination with the system requirements, that is, the content to be queried.

[0093] The said step S2 includes:

[0094] Step S2.1: Use the D2RQ tool to generate and modify the corresponding ttl file according to the designed schema. The ttl file describes the attributes of the nodes in the schema and the relationships between the nodes;

[0095] Step S2.2: According to the ttl file, convert the data in the database into an RDF triple file in.nt format;

[0096] Step S2.3: Use docker to install the Virtuoso image;

[0097] Step S2.4: Start the Virtuoso image. When starting, mount the above.nt file in the Virtuoso initialization data directory. After starting, use sparql to query Virtuoso to check whether the data is imported successfully.

[0098] The said step S3 includes:

[0099] Step S3.1: Use the paper id in the database to associate the title and abstract of the paper, as well as some relevant information that needs to be frequently queried, to construct the schema required by the ElasticSearch database. This step needs to be designed in combination with the system requirements, that is, the content to be queried.

[0100] The said step S4 includes:

[0101] Step S4.1: Create the above-designed schema in ElasticSearch and import the data into the ElasticSearch database in the schema format. The above several steps refer to Figure 2 the data support layer.

[0102] The backend query module of the said step S5 should at least include a data acquisition unit and a data caching unit, such as Figure 2As shown in the backend logic layer. The data acquisition unit obtains data through the above two databases. In particular, the Virtuoso database should be used when performing queries related to attributes or many-to-many relationships, while the ElasticSearch database should be used when document-related queries are required, such as querying titles using keywords. The data caching unit caches the above data in memory to improve the efficiency of subsequent lookups in the module.

[0103] Step S6 includes: According to the interface provided by the backend, complete the visual display of functions such as paper aggregation and content query on the frontend. As Figure 2 shown in the frontend display layer.

[0104] Example 4:

[0105] This embodiment provides a construction system for a query software system based on an academic knowledge graph, including the following modules:

[0106] Module M1: Design the schema of the RDF database;

[0107] Module M2: According to the designed schema, export the corresponding data from the database and store it in the RDF database virtuoso;

[0108] Module M3: Design the schema of the document database;

[0109] Module M4: According to the designed schema, export document data such as abstracts and titles of papers and some meta-information from the database and store them in the document database ElasticSearch;

[0110] Module M5: Build a backend query module according to the query capabilities provided by the cooperation of the above two databases;

[0111] Module M6: According to the interface provided by the backend, complete the visual display of functions such as paper aggregation and content query on the frontend.

[0112] The module M1 includes:

[0113] Module M1.1: Build the association relationship of triples according to the existing paper attributes in the database. For example, if a paper has an author attribute, then the triple (paper, is_written_by, author) needs to be built.

[0114] The module M2 includes:

[0115] Module M2.1: Use the D2RQ tool to generate and modify the corresponding ttl file according to the designed schema. The ttl file describes the attributes of the nodes in the schema and the relationships between nodes and nodes.

[0116] Module M2.2: Convert the data in the database into an RDF triple file in.nt format according to the ttl file.

[0117] Module M2.3: Use docker to install the Virtuoso image.

[0118] Module M2.4: Start the Virtuoso image and mount the above.nt file in the Virtuoso initialization data directory when starting. After starting, use sparql to query Virtuoso to check whether the data is imported successfully.

[0119] The said module M3 includes:

[0120] Step M3.1: Use the paper id in the database to associate the title and abstract of the paper, as well as some information that needs to be frequently queried, including the author, the journal to which it belongs, the location, etc., to construct the schema required for the ElasticSearch database.

[0121] The said module M4 includes:

[0122] Module M4.1: Create the above-designed schema in ElasticSearch and import the data into the ElasticSearch database in the schema format. The above several modules refer to Figure 2 the data support layer.

[0123] The backend query module of the said module M5 should at least include a data acquisition unit and a data cache unit, as shown in Figure 2 the backend logic layer. The data acquisition unit obtains data through the above two databases. In particular, when performing queries related to attributes or many-to-many relationships, the Virtuoso database should be used, while when performing document-related queries, such as querying the title using keywords, the ElasticSearch database should be used. The said data cache unit caches the above data in memory to improve the efficiency of subsequent searches of the module.

[0124] The said module M6 includes: Visualize functions such as paper aggregation and content query on the front end according to the interfaces provided by the backend. As shown in Figure 2 the front-end display layer.

[0125] Example 5:

[0126] Those skilled in the art can understand this embodiment as a more specific illustration of Embodiment 3.

[0127] Taking the GAKG database composed of 1,122,094 papers in the geoscience field and related attributes such as authors, institutions, and topics as the data source, this embodiment provides a query system based on the geoscience academic knowledge graph, which involves designing the schema of the RDF database; exporting corresponding data from the GAKG database and storing it in the RDF database Virtuoso; designing the schema of the document database; exporting document data such as abstracts and titles of papers and some meta-information from the GAKG database and storing them in the document database ElasticSearch; constructing a backend query module according to the query capabilities provided by the cooperation of the above two databases, including a paper aggregation module and a content Q&A module, and each module includes a data cache unit and a data query unit; and completing the visual display of functions such as paper aggregation and content query on the front end according to the interfaces provided by the backend. Specifically, as Figure 1 shown, it includes the following steps:

[0128] Step S1: Design the schema of the RDF database;

[0129] Step S2: According to the designed schema, export corresponding data from the GAKG database and store it in the RDF database Virtuoso;

[0130] Step S3: Design the schema of the document database;

[0131] Step S4: According to the designed schema, export document data such as abstracts and titles of papers and some meta-information from the GAKG database and store them in the document database ElasticSearch;

[0132] Step S5: According to the query capabilities provided by the cooperation of the above two databases, construct a backend query module, including a paper aggregation module and a content Q&A module, and each module includes a data cache unit and a data query unit;

[0133] Step S6: According to the interfaces provided by the backend, complete the visual display of functions such as paper aggregation and content query on the front end.

[0134] Step S1 includes: Designing the schema of the RDF database. Specifically:

[0135] Step S101: Each entity in the DDE database corresponds to a table, such as a paper table, an institution table, a scenario table (institution, journal), etc. When constructing triples, it can be divided into two types. One is the relationship between entities. For example, if a paper is cited by another paper, the triple is (paper, is_cited_by, paper); if a paper is written by a certain author, the triple is (paper, is_written_by, author), representing entity-relationship-entity. The other is the relationship between an entity and an attribute. For example, the publication year of a paper (paper, year, integer), representing entity-attribute-value type. When designing RDF triples, a unique identifier needs to be designed for each entity, and the data source ACEKG of DDE is used as the unified prefix for this. The specific designed schema triples are as follows:

[0136] Triples of entity and entity:

[0137] (paper, is_written_by, author)

[0138] (paper, is_cited_by, paper)

[0139] (paper, has_illustration, illustration)

[0140] (paper, has_table, papertable)

[0141] (paper, on_the_topic_of, topic)

[0142] (paper, on_the_topic_of, concept)

[0143] (paper, is_published_in, journal)

[0144] (paper, mention_location, location)

[0145] (paper, mention_timescale, timescale)

[0146] (author, is_last_known_in, affiliation)

[0147] (timescale, in_the_period_of, timescale)

[0148] (timescale, before, timescale)

[0149] (paper, has_theme, concept)

[0150] (paper, has_developed, concept)

[0151] (paper, has_designed, concept)

[0152] (paper, has_concluded, concept)

[0153] (paper, learn_in_the_way_of, concept)

[0154] (location, has_geohash, location)

[0155] (affiliation, is_located_in, location)

[0156] (illustration, has_geohash, location)

[0157] Triple of entity and attribute:

[0158] (country, alphacode, integer)

[0159] (country, name, string)

[0160] (country, official_name, string)

[0161] (country, label, string)

[0162] (affiliation, homepage, string)

[0163] (affiliation, abbreviation, string)

[0164] (affiliation, gridcode, string)

[0165] (affiliation, label, string)

[0166] (affiliation, introduction, string)

[0167] (author, last_paper_date, date)

[0168] (author, name, string)

[0169] (author, label, string)

[0170] (location, label, string)

[0171] (location, name, string)

[0172] (location, latittude, double)

[0173] (location, longitude, double)

[0174] (timescale, label, string)

[0175] (illustration, dpi, integer)

[0176] (illustration, label, string)

[0177] (illustration, tag, string)

[0178] (illustration, caption, string)

[0179] (journal, label, string)

[0180] (journal, homepage, string)

[0181] (journal, name, string)

[0182] (journal, ISSN, string)

[0183] (concept, definition, string)

[0184] (concept, source, string)

[0185] (paper, page, integer)

[0186] (paper, url, string)

[0187] (paper, year, string)

[0188] (paper, first_page, integer)

[0189] (paper, issue, string)

[0190] (paper, last_page, string)

[0191] (paper, volume, integer)

[0192] (paper, label, string)

[0193] (paper, publish_date, date)

[0194] (paper, doi, string)

[0195] (papertable, dpi, integer)

[0196] (papertable, label, string)

[0197] (papertable, tag, string)

[0198] (papertable, caption, string)

[0199] (topic, introduction, string)

[0200] (topic, definition, string)

[0201] (topic, label, string)

[0202] Step S2 includes: Using the D2RQ tool to generate and modify the corresponding ttl file according to the designed schema. The ttl file describes the attributes of the nodes in the schema and the relationships between the nodes; According to the ttl file, convert the data in the database into an RDF triple file in.nt format; Use docker to install the Virtuoso image; Start the Virtuoso image, and mount the above.nt file in the Virtuoso initialization data directory when starting. After starting, use sparql to query Virtuoso to check whether the data is imported successfully. Specifically:

[0203] Step S201: Use the D2RQ tool to generate and modify the corresponding ttl file according to the designed schema. The ttl file describes the attributes of the nodes in the schema and the relationships between the nodes. The content of the ttl file can be referred to Figure 4 ;

[0204] Step S202: According to the ttl file, use the D2RQ tool to convert the data in the database into an RDF triple file in.nt format;

[0205] Step S2.3: Use docker to install the Virtuoso image;

[0206] Step S2.4: Start the Virtuoso image, and mount the above.nt file in the Virtuoso initialization data directory when starting. After starting, use sparql to query Virtuoso to check whether the data is imported successfully.

[0207] Step 3 includes: Use the paper id in the database to associate the title and abstract of the paper, as well as the relevant information that the geological query system needs to query frequently, to construct the schema required by the ElasticSearch database.

[0208] Step S301: Specifically, design the following document schema:

[0209]

[0210]

[0211]

[0212]

[0213]

[0214] Step S4 includes: According to the designed schema, export the document data such as the abstract and title of the papers, as well as some meta-information from the GAKG database, and store them in the document database ElasticSearch;

[0215] Step S5 includes: Based on the query capabilities provided by the cooperation of the above two databases, construct a backend query module, including a paper aggregation module and a content Q&A module, and each module contains a data cache unit and a data query unit;

[0216] Step S501: Construct a general data cache unit. Write an LFU module through code. Use the query api + parameters as the key, store the query result as the value in the memory, only keep the first one thousand high-frequency key-value pairs, and set a timeout to prevent excessive memory occupation.

[0217] Step S502: Construct a general query unit. Use the sparql language for querying the Virtuoso database, and use the AsyncElasticsearch module of python for ElasticSearch, and call the corresponding interface for querying.

[0218] Step S503: Construct a paper aggregation module. Specifically, by giving keywords, query for matching papers in the paper title and abstract, and aggregate the authors, journals, countries, institutions, and the research period of the papers, aiming to display options for the front-end home page. Since we performed text parsing on the title and abstract when constructing the document database schema, text matching retrieval can be performed based on the inverted index provided by ElasticSearch.

[0219] Step S504: Construct a content Q&A module: Map to the corresponding sparql query statement on the backend through a given artificial template. For example, if the question is "the paper that first proposed the concept concept1", the backend will construct the following sparql query statement:

[0220] SELECT?s_title?year WHERE{

[0221] ?s rdf:type ace:paper.

[0222] ?s acer:is_in_the_field_of?c1.

[0223] ?c1 skos:prefLabel?c1_label.

[0224] ?s acep:title?s_title.

[0225] ?s acep:year?year.

[0226] FILTER(REGEX(str(?c1_label), 'concept1')).

[0227] }

[0228] ORDER BY ASC(?year)

[0229] LIMIT 10

[0230] A total of 16 Q&A templates are provided as follows:

[0231] (1) Articles / images that involve both entities concept1 and concept2

[0232] (2) Images in the papers where author1 studies concept1

[0233] (3) The paper that first proposed the concept concept1

[0234] (4) The author who first proposed the concept concept1

[0235] (5) The journal / conference where author1 first published a paper

[0236] (6) The knowledge concepts obtained / designed / developed / owned / carried forward by article paper1

[0237] (7) Institutions that study concept1

[0238] (8) Authors who study concept1

[0239] (9) Articles published in the field of concept1 in year1

[0240] (10) Articles published by author1 in year1 and before

[0241] (11) Papers that contain the knowledge concept concept1

[0242] (12) Images that contain the knowledge concept concept1

[0243] (12) The locations studied in paper1

[0244] (14) The locations where author1 conducts research

[0245] (15) The locations where concept1 is studied

[0246] Images / tables related to concept1

[0247] Step S6 includes: according to the interfaces provided by the back end, visual displays of functions such as thesis aggregation and content query are completed on the front end, as Figure 3 shown.

[0248] First, the construction method of the query software system based on the academic knowledge graph provides a general method for developing an academic knowledge graph query system, and a set of displayable query systems can be quickly iterated according to this method. Secondly, in this embodiment, the visualization of various functions is performed on the front end, enabling users to intuitively query the content they want to see from complex data. Finally, using this construction method, a query system based on the geoscience academic knowledge graph is specifically developed, demonstrating the specific schema design process and the paradigm of front-end and back-end development, and proving the effectiveness of this construction method.

[0249] The present invention significantly reduces the development cost and cycle of the software. The system architecture is simple and clear, the usage process is fast and convenient, and the applicable scope is relatively wide, which can provide an effective way for the upper-layer application development of the academic knowledge graph.

[0250] Those skilled in the art know that in addition to implementing the system, its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system, its various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same functions. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a kind of hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structures within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as either software modules for implementing the method or structures within the hardware component.

[0251] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A construction method of a query software system based on an academic knowledge graph, characterized in that, It includes the following steps: Step 1: Design the schema of the RDF database; Step 2: According to the designed schema, export the corresponding data from the database and store it in the RDF database Virtuoso; Step 3: Design the schema of the document database; Step 4: According to the designed schema, export the relevant document data and some meta-information of the papers from the database and store it in the document database ElasticSearch; Step 5: Build the backend query module according to the query capabilities provided by the cooperation of the above two databases; Step 6: Complete the visual display of relevant functions on the front end according to the interfaces provided by the backend; In Step 1, according to the existing paper attributes in the database, construct the association relationship of triples and design in combination with the content to be queried; Step 2 includes the following steps: Step 2.1: Use the D2RQ tool to generate and modify the corresponding ttl file according to the designed schema. The ttl file describes the attributes of the nodes in the schema and the relationships between nodes and nodes; Step 2.2: According to the ttl file, convert the data in the database into an RDF triple file in.nt format; Step 2.3: Use docker to install the Virtuoso image; Step 2.4: Start the Virtuoso image. When starting, mount the above.nt format RDF triple file in the Virtuoso initialization data directory. After starting, use sparql to query Virtuoso to check whether the data is imported successfully; In Step 3, use the paper id in the database to associate the title and abstract of the paper and some relevant information that needs to be frequently queried, construct the schema required by the ElasticSearch database, and design in combination with the content to be queried.

2. The construction method of the query software system based on the academic knowledge graph according to claim 1, characterized in that, In Step 4, the relevant document data and some meta-information of the papers exported from the database include: the abstract of the paper, the title of the paper; In Step 6, the visual display of relevant functions on the front end includes: paper aggregation function, content query function.

3. The construction method of the query software system based on the academic knowledge graph according to claim 1, characterized in that, In Step 4, create the above-designed schema in ElasticSearch and import the data into the ElasticSearch database in the schema format.

4. The construction method of the query software system based on the academic knowledge graph according to claim 1, characterized in that, In Step 5, the backend query module at least includes a data acquisition unit and a data cache unit; The data acquisition unit obtains data through the above two databases, uses the Virtuoso database when performing queries related to attributes or many-to-many relationships, and uses the ElasticSearch database when performing document-related queries; The data cache unit caches the above data in memory for subsequent searches of the module.

5. A construction system for a query software system based on an academic knowledge graph, characterized in that, It includes the following modules: Module M1: Design the schema of the RDF database; Module M2: According to the designed schema, export the corresponding data from the database and store it in the RDF database Virtuoso; Module M3: Design the schema of the document database; Module M4: Export the relevant document data and partial meta-information of the thesis from the database according to the designed schema, and store them in the document database ElasticSearch; Module M5: Build a backend query module according to the query capabilities provided by the cooperation of the above two databases; Module M6: Complete the visual display of relevant functions on the front end according to the interfaces provided by the backend; In the said Module M1, according to the existing thesis attributes in the database, construct the association relationship of triples and design in combination with the content to be queried; The said Module M2 includes the following modules: Module M2.1: Use the D2RQ tool to generate and modify the corresponding ttl file according to the designed schema. The ttl file describes the attributes of the nodes in the schema and the relationships between nodes and nodes; Module M2.2: Convert the data in the database into an.nt format RDF triple file according to the ttl file; Module M2.3: Use docker to install the Virtuoso image; Module M2.4: Start the Virtuoso image. When starting, mount the above.nt format RDF triple file in the Virtuoso initialization data directory. After starting, use sparql to query Virtuoso to check whether the data is imported successfully; In the said Module M3, use the thesis id in the database to associate the title and abstract of the thesis, as well as some relevant information that needs to be frequently queried, construct the schema required by the ElasticSearch database, and design in combination with the content to be queried; In the said Module M4, create the above-designed schema in ElasticSearch and import the data into the ElasticSearch database in the schema format; In the said Module M5, the backend query module at least includes a data acquisition unit and a data cache unit; The said data acquisition unit obtains data through the above two databases, uses the Virtuoso database when performing queries related to attributes or many-to-many relationships, and uses the ElasticSearch database when performing document-related queries; The said data cache unit caches the above data in memory for subsequent searches of the module; 6. The construction system of the query software system based on the academic knowledge graph according to claim 5, characterized in that, In the said Module M4, exporting the relevant document data and partial meta-information of the thesis from the database includes: the abstract of the thesis, the title of the thesis; In the said Module M6, completing the visual display of relevant functions on the front end includes: thesis aggregation function, content query function.

Citation Information

Patent Citations

  • Disease diagnosis method, device and equipment based on knowledge reasoning of medical knowledge graph

    CN112820400A

  • Knowledge graph construction and query method based on power enterprise

    CN110929042A

  • Semantic knowledge base

    US20190311003A1