Knowledge graph-based scientific research project retrieval method

Through the scientific research project search method based on knowledge graph, the problem that traditional information search methods cannot understand semantics and deal with ambiguousness and synonyms is solved, and efficient and accurate scientific research project search is achieved, which is especially suitable for the processing of large-scale text data.

CN119961462APending Publication Date: 2025-05-09SICHUAN COMPUTER RES INST
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202411979816.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Traditional information retrieval methods cannot accurately understand semantics, process ambiguousness and synonyms, resulting in low efficiency and accuracy in finding information related to user query in massive text data.

Method used

Using a scientific research project search method based on knowledge graph, efficient semantic retrieval of scientific research projects is achieved through steps such as data collection and preprocessing, entity recognition and relationship extraction, knowledge graph construction and optimization, vector space model construction, path generation algorithm application, semantic matching and search result sorting.

Benefits of technology

This method can effectively solve the problems of semantic understanding, ambiguousness and synonyms, improve the accuracy and efficiency of information retrieval, and especially has high scalability and flexibility when processing large-scale text data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961462A_ABST
    Figure CN119961462A_ABST
Patent Text Reader

Abstract

The invention provides a scientific research project retrieval method based on a knowledge graph. The method solves a plurality of problems in the prior art. Firstly, the efficiency and the accuracy of information retrieval are improved; by constructing the knowledge graph, information in various data sources is integrated, and more comprehensive and accurate information is provided for users. Secondly, complex query and cross-domain integration are supported; the entities and relationships in the mapping knowledge domain can span different fields and themes, so that a user can perform cross-field query and integration. In addition, the method also has the capability of real-time updating. Along with continuous addition of new data, the knowledge graph can be dynamically updated, and timeliness and accuracy of information are kept. Finally, the method provides a rich interaction mode and a user feedback mechanism, so that the user can conveniently obtain the required information and participate in the optimization of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention specifically relates to a scientific research project retrieval method based on knowledge graph. Background Art

[0002] The Vector Space Model (VSM) is an important technique in information retrieval. Originating from the mathematical processing of text data, it aims to convert text data into vector form for efficient similarity calculation and retrieval. In the context of information retrieval, the VSM treats both documents and queries as vectors, where each dimension represents a term or feature, and the value of the vector typically represents the weight of that term in the document or query.

[0003] Information retrieval is the process of finding information relevant to a user's query within massive amounts of text data. While traditional information retrieval methods, such as keyword matching and Boolean search, are simple and intuitive, they suffer from limitations such as an inability to accurately understand semantics and handle polysemy and synonyms. These limitations limit their effectiveness and accuracy.

[0004] To sum up, this application proposes a scientific research project retrieval method based on knowledge graph. Summary of the Invention

[0005] The purpose of the present invention is to provide a scientific research project retrieval method based on knowledge graph to address the deficiencies of the existing technology. The scientific research project retrieval method based on knowledge graph can well solve the above problems.

[0006] To achieve the above requirements, the technical solution adopted by the present invention is to provide a scientific research project retrieval method based on knowledge graph, which includes the following steps:

[0007] S1: Data collection and preprocessing steps. First, data is collected from various data sources, including academic papers, project reports, and patent information. This data often exists in the form of unstructured text and requires preprocessing, including noise removal, word segmentation, and stop word removal to ensure data quality and consistency. This is done using efficient text processing techniques, namely lexical analysis and syntactic analysis in natural language processing.

[0008] S2: Perform entity recognition and relationship extraction. In the preprocessed text data, various entities and the relationships between them need to be identified. This is achieved by using named entity recognition and relationship extraction techniques in NLP. NER identifies entities in the text, while relationship extraction can determine the associations between these entities. You can use spaCy's NLP library for entity recognition and relationship extraction.

[0009] S3: Steps for constructing a knowledge graph: Based on the identified entities and relationships, a knowledge graph is constructed. During the construction process, a unique identifier is assigned to each entity and relationship, and graph database technology is used for storage and query.

[0010] S4: Steps for knowledge graph optimization. The constructed knowledge graph contains redundant information or noise and needs to be optimized. The optimization steps include deduplication, merging similar entities, and deleting invalid relationships. Through the reasoning ability of the knowledge graph, new knowledge and patterns can be discovered from existing knowledge to further enrich the graph content.

[0011] S5: Steps for constructing a vector space model. To improve retrieval efficiency, a vector space model is used to represent entities and relationships in the knowledge graph. The vector space model evaluates the similarity between the query vector and the document vector by calculating the cosine similarity between them. The formula is:

[0012] D(q,d)=∫S(q(t),d(t))dt+C;

[0013] in:

[0014] D(q,d) represents the similarity between the query vector q and the document vector d;

[0015] S(q(t),d(t)) represents the local similarity function between the query vector q and the document vector d at time t;

[0016] ∫ represents the integral operation, which is used to calculate the accumulation of local similarities in the entire time or space range, represented by t here, to capture the similarity characteristics between vectors that change with time or space;

[0017] dt represents a small change in t;

[0018] C is a constant term used to adjust the calculation results of similarity;

[0019] S6: Apply the path generation algorithm. The goal of path generation is to map the entities, concepts, and relationships in the user's query to the knowledge graph and generate a query path that matches the entities, concepts, and relationships on the knowledge graph. This is done using a traversal algorithm, which performs a depth-first or breadth-first search on the knowledge graph until matching entities, concepts, and relationships are found.

[0020] S7: Semantic matching step. After generating the query path, semantic matching is required. Semantic matching matches the user query with the entities, relationships, and concepts in the knowledge graph to find the most relevant results. Graph neural network technology is used to capture the complex interaction patterns between nodes and edges in the graph, thereby revealing the underlying patterns and connections hidden behind the data.

[0021] S8: the step of sorting the search results; sorting the search results based on the similarity between the query and the results and the importance of the results, using a ranking learning machine learning algorithm to sort the search results to ensure that the most relevant results are ranked first;

[0022] S9: The step of displaying and interacting with the results; displaying the sorted search results to the user and providing interactive functions.

[0023] The advantages of this knowledge graph-based scientific research project retrieval method are as follows:

[0024] (1) Semantic understanding problem: Traditional information retrieval methods are mainly based on text matching and cannot understand the semantic content of text. However, the vector space model converts text into vector form and can use the similarity between vectors to calculate the semantic similarity between texts, thus solving the problem of semantic understanding to a certain extent.

[0025] (2) Polysemy and synonymy: In natural language processing, the same word can have multiple meanings (polysemy), and different words can represent the same concept (synonyms). These problems pose challenges to information retrieval. The vector space model can address these issues to a certain extent and improve retrieval accuracy by calculating the weight of words in documents and combining vector similarity calculations.

[0026] (3) Efficient retrieval: Traditional information retrieval methods are inefficient when processing large-scale text data. However, the vector space model can achieve fast information retrieval by converting text into vector form and utilizing efficient vector operation algorithms and indexing techniques.

[0027] (4) Scalability and flexibility: The vector space model has good scalability and flexibility. It can easily introduce new vocabulary and feature items to adapt to the ever-changing text data and query requirements. At the same time, it can also be combined with other information retrieval technologies, such as natural language processing and machine learning, to improve the effectiveness and accuracy of retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The same reference numerals are used in these drawings to represent the same or similar parts. The exemplary embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0029] Figure 1 A schematic diagram of a scientific research project retrieval method based on a knowledge graph according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0030] In order to make the objectives, technical solutions and advantages of this application clearer, this application is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0031] In the following description, references to "one embodiment," "an embodiment," "an example," "an example," etc. indicate that the embodiment or example described may include certain features, structures, characteristics, properties, elements, or limitations, but not every embodiment or example necessarily includes the certain features, structures, characteristics, properties, elements, or limitations. In addition, repeated use of the phrase "according to one embodiment of the present application" may refer to the same embodiment, but does not necessarily refer to the same embodiment.

[0032] For the sake of simplicity, certain technical features well known to those skilled in the art are omitted in the following description.

[0033] According to one embodiment of the present application, a scientific research project retrieval method based on knowledge graph is provided, such as Figure 1 As shown, the following steps are included:

[0034] Step 1: Data collection and preprocessing

[0035] In a research project retrieval system, data must first be collected from various sources, including academic papers, project reports, and patent information. This data often exists in the form of unstructured text and requires preprocessing, such as noise removal, word segmentation, and stop word removal. This step is the foundation for the subsequent construction of the knowledge graph, and ensuring data quality and consistency is crucial. The quality of data preprocessing directly affects the effectiveness of subsequent steps, so efficient text processing technologies, such as lexical analysis and syntactic analysis in natural language processing (NLP), are essential.

[0036] Step 2: Entity Recognition and Relation Extraction

[0037] In preprocessed text data, it is necessary to identify various entities (such as research project names, researchers, institutions, etc.) and the relationships between them (such as collaborative relationships, relationships between projects and achievements, etc.). This can be achieved by using named entity recognition (NER) and relationship extraction techniques in NLP. NER can identify entities in text, while relationship extraction can determine the relationships between these entities. For example, NLP libraries such as spaCy can be used for entity recognition and relationship extraction.

[0038] Step 3: Knowledge graph construction

[0039] Based on the identified entities and relationships, a knowledge graph can be constructed. A knowledge graph is a semantic network consisting of nodes (entities, concepts) and edges (relationships). During the construction process, each entity and relationship must be assigned a unique identifier and stored and queried using technologies such as graph databases. This step requires reference to existing domain knowledge and professional background to ensure the accuracy and completeness of the knowledge graph.

[0040] Step 4: Knowledge Graph Optimization

[0041] A constructed knowledge graph may contain redundant information or noise, requiring optimization. Optimization steps include deduplication, merging similar entities, and removing invalid relationships. Furthermore, the reasoning capabilities of the knowledge graph can be leveraged to discover new insights and patterns from existing knowledge, further enriching the graph's content.

[0042] Step 5: Vector Space Model Construction

[0043] Steps for constructing a vector space model; In order to improve retrieval efficiency, a vector space model is used to represent entities and relationships in the knowledge graph. The vector space model evaluates the similarity between the query vector and the document vector by calculating the cosine similarity between them. The formula is:

[0044] D(q,d)=∫S(q(t),d(t))dt+C;

[0045] in:

[0046] D(q,d) represents the similarity between the query vector q and the document vector d;

[0047] S(q(t),d(t)) represents the local similarity function between the query vector q and the document vector d at time t;

[0048] ∫ represents the integral operation, which is used to calculate the accumulation of local similarities in the entire time or space range, represented by t here, to capture the similarity characteristics between vectors that change with time or space;

[0049] dt represents a small change in t;

[0050] C is a constant term used to adjust the calculation results of similarity;

[0051] In specific implementations, the local similarity function S(q(t), d(t)) is defined based on the specific characteristics of the query vector and document vector. For example, if the query vector and document vector are both time series data, then S(q(t), d(t)) can be defined as a function of the difference or ratio of the two time series at time t. If the query vector and document vector are point sets in space, then S(q(t), d(t)) can be defined as a function of the Euclidean distance or cosine similarity between the two point sets at time t (or a certain spatial position).

[0052] In addition, the value of constant C can be adjusted according to actual needs. For example, if you want to increase the weight of global similarity, you can increase the value of C; if you want to pay more attention to local similarity, you can decrease the value of C.

[0053] By introducing calculus and constant terms, the improved vector space model similarity calculation formula can more comprehensively capture the similarity characteristics between query vectors and document vectors, improving the accuracy and robustness of similarity calculations. At the same time, this formula also provides more possibilities and flexibility for subsequent optimization and expansion.

[0054] Step 6: Application of path generation algorithm

[0055] Path generation is a critical issue in knowledge graph-based information retrieval. The goal of path generation is to map the entities, concepts, and relationships in a user's query statement onto the knowledge graph and generate a query path that matches the entities, concepts, and relationships on the knowledge graph. Commonly used path generation algorithms include traversal and pruning. Traversal algorithms perform a depth-first or breadth-first search on the knowledge graph until matching entities, concepts, and relationships are found. Pruning algorithms, on the other hand, restrict and prune the search path during the search process using constraints and heuristics to reduce the search space and accelerate path generation.

[0056] Step 7: Semantic Matching

[0057] After generating the query path, semantic matching is required. This involves matching the user query with the entities, relationships, and concepts in the knowledge graph to find the most relevant results. This step can leverage technologies such as graph neural networks (GNNs) to capture the complex interaction patterns between nodes and edges in the graph, thereby revealing the underlying patterns and connections hidden in the data.

[0058] Step 8: Sorting search results

[0059] To improve the user experience, search results need to be sorted. This sorting can be based on factors such as the similarity between the query and the results, or the importance of the results. Machine learning algorithms, such as learning to rank (LTO), can be used to sort search results, ensuring that the most relevant results appear first.

[0060] Step 9: Result display and interaction

[0061] The sorted search results are displayed to the user, and interactive functions are provided, such as result filtering, detailed viewing, etc. This step requires the design of a friendly user interface and interactive methods so that users can easily obtain the required information.

[0062] Step 10: Feedback and Optimization

[0063] When using a search system, users may provide feedback, such as satisfaction ratings and query intent. This feedback can be used to optimize the search system, improving search performance and user satisfaction. For example, path generation algorithms and semantic matching algorithms can be adjusted based on user feedback.

[0064] According to one embodiment of the present application, step S3 of the scientific research project retrieval method based on the knowledge graph specifically includes the following steps:

[0065] System construction and preparation: Determine the data model of the knowledge graph based on the data content, data type, and data characteristics of the knowledge graph. At the same time, establish the data model in the form of a large set based on the data content, data type, and data characteristics, denoted as set N. Establish sub-sets within the sets corresponding to the data content, data type, and data characteristics, denoted as set n, so that the overall framework of the data model is hierarchical by the sets composed of it, where set N contains set n.

[0066] Obtaining data information of the knowledge graph: After the data model is established, all specific data information that constitutes the knowledge graph is collected to form a large data set that contains all the data information that constitutes the knowledge graph;

[0067] Data information processing: pre-processing the collected data information;

[0068] Substitute data information according to the model: After the data set and the corresponding data information within it are processed, the processed data information is substituted into the corresponding sets in the data model according to the established data model. Specifically, the search set in the data model is substituted into set N, and the data information of the corresponding category in the search set is substituted into set n;

[0069] Forming a knowledge graph: After the data information is substituted into the data model, the data model that integrates the data information becomes the search knowledge graph. At this point, the knowledge graph has been constructed and can be used for subsequent search operations.

[0070] According to one embodiment of the present application, step S8 of the scientific research project retrieval method based on the knowledge graph specifically includes:

[0071] Get search statement: Enter the search statement through the system interface, and the system obtains the search statement as input;

[0072] Transforming search statements: The system transforms search statements into multiple similar search expressions using natural language processing techniques, including word segmentation, part-of-speech tagging, and synonym replacement. Furthermore, the system assigns a frequency weight to each search expression based on search history, helping the system more accurately understand the user's search intent and prioritize results related to the user's historical searches.

[0073] Forming search control instructions: The system forms search control instructions according to each search expression and its corresponding frequency weight; setting each search control instruction to terminate during the search process helps control the depth and breadth of the search and avoid excessive consumption of system resources due to excessive searching;

[0074] Execute retrieval operations: The query module controls the query depth of the retrieval expression in the knowledge graph library based on the retrieval control instructions; the system queries the knowledge results associated with the retrieval expression in the knowledge graph library according to the retrieval expression; the termination instruction is used to determine whether to terminate the query based on the credibility of the knowledge result or the number of knowledge results. When the query result meets the preset conditions, the system terminates the query and returns the result.

[0075] According to one embodiment of the present application, step S9 of the scientific research project retrieval method based on the knowledge graph specifically includes:

[0076] Choose an appropriate display method: Based on the index category of the retrieved node, the characteristics of the data itself, and the need to efficiently and intuitively present the search results to users, choose an appropriate display method, such as key-value pairs, tables, images, and knowledge graphs.

[0077] Display search results: Display the acquired data in the corresponding display components, including equipment attributes as key-value pairs, equipment performance in table format, equipment photos as images, and relationships between equipment as knowledge graphs;

[0078] Link to related knowledge: Click on a node displayed in the knowledge graph to link to the child or parent node page of the search node. Use the knowledge graph to obtain more search results and related information;

[0079] Result processing and feedback: Further processing or operation of search results, including saving, sharing, and exporting. The system continuously optimizes the search algorithm and result display method based on user feedback and behavior data to improve search accuracy and user experience.

[0080] The above-described embodiments merely represent several implementations of the present invention. While the descriptions are relatively specific and detailed, they are not to be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, and such modifications and improvements fall within the scope of the present invention. Therefore, the scope of the present invention shall be determined by the claims.

Claims

1. A scientific research project retrieval method based on knowledge graph, characterized in that: The steps include: S1: Data collection and preprocessing steps; First, data is collected from various data sources, including academic papers, project reports, and patent information. These data often exist in the form of unstructured text and need to be preprocessed, including noise removal, word segmentation, and stop word removal to ensure data quality and consistency. This is done using efficient text processing technology, namely lexical analysis and syntactic analysis in natural language processing. S2: Steps for entity recognition and relationship extraction; In the preprocessed text data, various entities and the relationships between them need to be identified. This can be achieved by using named entity recognition and relation extraction techniques in NLP. NER identifies entities in the text, while relation extraction can determine the associations between these entities. You can use spaCy's NLP library for entity recognition and relation extraction. S3: Steps for constructing the knowledge graph; Based on the identified entities and relationships, a knowledge graph is constructed. During the construction process, each entity and relationship is assigned a unique identifier, and stored and queried using graph database technology; S4: Steps for knowledge graph optimization; The constructed knowledge graph contains redundant information or noise, which needs to be optimized. The optimization steps include deduplication, merging similar entities, and deleting invalid relationships. Through the reasoning ability of the knowledge graph, new knowledge and rules can be discovered from existing knowledge to further enrich the graph content. S5: Steps for constructing a vector space model; In order to improve retrieval efficiency, the vector space model is used to represent entities and relationships in the knowledge graph. The vector space model evaluates the similarity between the query vector and the document vector by calculating the cosine similarity between them. The formula is: D(q,d)=∫S(q(t),d(t))dt+C; in: D(q, d) represents the similarity between the query vector q and the document vector d; S(q(t), d(t)) represents the local similarity function between the query vector q and the document vector d at time t; ∫ represents the integral operation, which is used to calculate the accumulation of local similarities in the entire time or space range, represented by t here, to capture the similarity characteristics between vectors that change over time or space; dt represents a small change in t; C is a constant term used to adjust the calculation results of similarity; S6: Step of applying the path generation algorithm; The purpose of path generation is to map the entities, concepts, and relationships in the user's query statement to the knowledge graph and generate a query path that can match the entities, concepts, and relationships on the knowledge graph. This is done using a traversal algorithm, which performs a depth-first or breadth-first search on the knowledge graph until matching entities, concepts, and relationships are found. S7: step of performing semantic matching; After the query path is generated, semantic matching is required. Semantic matching is to match the user query with the entities, relationships, and concepts in the knowledge graph to find the most relevant results. The graph neural network technology is used to capture the complex interaction patterns between nodes and edges in the graph, thereby revealing the potential rules and associations hidden behind the data. S8: a step of sorting the search results; Sort the search results based on the similarity between the query and the results and the importance of the results. Use the learning-to-rank machine learning algorithm to sort the search results to ensure that the most relevant results are ranked first. S9: Steps for result display and interaction; The sorted search results are displayed to the user and interactive functions are provided.

2. The scientific research project retrieval method based on knowledge graph according to claim 1 is characterized in that: It also includes S10: a step of performing feedback and optimization; When using the retrieval system, users will provide some feedback, which is used to optimize the retrieval system and improve retrieval results and user satisfaction.

3. The scientific research project retrieval method based on knowledge graph according to claim 1 is characterized in that: Step S3 specifically includes the following steps: System construction and preparation: Determine the data model of the knowledge graph, and determine the data model based on the data content, data type and data characteristics of the knowledge graph; at the same time, establish the data model in the form of a large set based on the data content, data type and data characteristics, recorded as set N, and establish a sub-set in the set corresponding to the data content, data type and data characteristics, recorded as set n, so that the overall framework of the data model is graded by the sets composed of it, where set N contains set n; Obtaining data information of the knowledge graph: After the data model is established, all specific data information constituting the knowledge graph is collected to form a huge data set, which contains all the data information constituting the knowledge graph; Data information processing: pre-process the collected data information; Substitute data information according to the model: After the data set and the corresponding data information in it are processed, the processed data information is substituted into the corresponding sets in the data model according to the established data model. Specifically, the search set in the data model is substituted into set N, and the data information of the corresponding category in the search set is substituted into set n; Forming a knowledge graph: After the data information is substituted into the data model, the data model that integrates the data information is the search knowledge graph. At this point, the knowledge graph has been constructed and can be used for subsequent retrieval operations.

4. The scientific research project retrieval method based on knowledge graph according to claim 1 is characterized in that: Step S8 specifically includes: Get search statement: input the search statement through the system interface, and the system gets the search statement as input; Transform search statements: The system transforms search statements into multiple similar search expressions through natural language processing technology, including word segmentation, part-of-speech tagging, and synonym replacement. At the same time, the frequency weight of each search expression is configured based on the search history, which helps the system understand the user's search intent more accurately and give priority to displaying results related to the user's historical searches. Forming search control instructions: The system forms search control instructions according to each search expression and its corresponding frequency weight; setting each search control instruction to terminate the instruction during the search process helps control the depth and breadth of the search and avoid excessive consumption of system resources due to excessive search; Execute retrieval operations: The query module controls the query depth of the retrieval expression in the knowledge graph library based on the retrieval control instructions; the system queries the knowledge results associated with the retrieval expression in the knowledge graph library according to the retrieval expression; the termination instruction is used to determine whether to terminate the query based on the credibility of the knowledge results or the number of knowledge results. When the query results meet the preset conditions, the system terminates the query and returns the results.

5. The scientific research project retrieval method based on knowledge graph according to claim 1 is characterized in that: Step S9 specifically includes: Choose the appropriate display method: According to the index category to which the obtained node belongs, the characteristics of the data itself, and the need to efficiently and intuitively display the search results to users, choose the appropriate display method, including key-value pairs, tables, pictures, and knowledge graphs; Display search results: Display the acquired data in the corresponding display components, including equipment attributes as key-value pairs, equipment performance in table form, equipment photos as pictures, and relationships between equipment in the form of knowledge graphs; Link to related knowledge: Click on the node displayed in the knowledge graph to link to the child node or parent node page of the search node. Use the knowledge graph to obtain more search results and related information; Result processing and feedback: Further processing or operation of the search results, including saving, sharing, and exporting. The system continuously optimizes the search algorithm and result display method based on user feedback and behavior data to improve the accuracy of the search and user experience.

Citation Information

Cited By

  • Electric power scientific research information text retrieval optimization method and system

    CN120508607A

  • Science and technology project information management method and system based on big data

    CN121010216A

  • Big data-based scientific and technological project information management method and system

    CN121010216B

  • Retrieval and tracking method and device for scientific research information

    CN121233697A

  • Inspection knowledge inference analysis method based on combination of large language model and knowledge graph

    CN121480667A