Scientific and technological service field knowledge base construction method and system based on knowledge graph and RAG
By adopting methods based on knowledge graphs and RAG technology in the field of science and technology services, we build scientific and personal knowledge graphs, and combining RAG models for knowledge retrieval and generation, the shortcomings of the traditional knowledge base in terms of knowledge representation, search efficiency, dynamic update and personalized services are solved, and efficient, accurate and personalized knowledge services are achieved.
Patent Information
- Application Number
- CN202510015821.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-06
AI Technical Summary
In the field of scientific and technological services, the traditional knowledge base has problems such as incomplete knowledge representation, low retrieval efficiency, difficulty in dynamic updates and lack of personalized services.
Using a method based on knowledge graph and RAG technology, we use scientific and technological knowledge graphs and personal knowledge graphs to search and generate knowledge with RAG models to achieve efficient knowledge representation and personalized services.
It significantly improves the intelligence level and service quality of the knowledge base, improves the accuracy and personalization of the search results, supports complex queries and reasoning, and meets users' needs for efficient and accurate services.
Smart Images

Figure CN119940500A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and knowledge engineering technology, and in particular to a method and system for constructing a knowledge base in the field of scientific and technological services based on a knowledge graph and RAG. Background Art
[0002] In the field of scientific and technological services, the construction and management of knowledge bases are crucial for technology transformation, innovation support, and service optimization. However, traditional knowledge base construction methods have the following defects:
[0003] 1. Incomplete knowledge representation
[0004] Traditional knowledge bases usually use structured databases or simple text storage methods, which are difficult to fully represent complex scientific and technological knowledge. Knowledge in the field of science and technology is highly professional and complex, involving a large amount of entities, relationships and their attribute information. For example, in patent analysis, a technology may involve multiple inventors, affiliated institutions, technical fields, application scenarios and other multi-dimensional information. Due to the lack of structured representation of knowledge, traditional knowledge bases are difficult to effectively capture these complex relationships, resulting in incomplete knowledge representation and limiting the in-depth application of knowledge.
[0005] 2. Low retrieval efficiency
[0006] The retrieval function of existing knowledge bases mostly relies on keyword matching or simple rule engines, which have a low level of intelligence and are difficult to meet users' needs for efficient and accurate retrieval. For example, in technical consultation, users often need to filter out content that is highly relevant to their needs from a large amount of literature and data, while traditional retrieval tools often return a large number of irrelevant or low-quality results due to their lack of understanding of semantics and context, causing users to spend a lot of time on secondary screening. In addition, existing retrieval systems are difficult to support complex queries and reasoning, such as predicting the development trend or potential application scenarios of technologies by analyzing the relationship chain between technologies.
[0007] 3. Difficulty in dynamic updates
[0008] Knowledge in the field of science and technology is updated at an extremely fast pace, and emerging technologies emerge in an endless stream. Traditional knowledge bases, due to the lack of a dynamic update mechanism, are unable to keep up with the pace of knowledge updates. For example, in technology transfer, the commercial prospects and market application information of new technologies need to be updated in real time, while traditional knowledge bases mostly use regular batch updates, which cannot meet real-time requirements. In addition, due to the wide and scattered sources of data, traditional knowledge bases often face problems such as inconsistent data formats and data redundancy when integrating new data, which further increases the difficulty of dynamic updates.
[0009] 4. Lack of personalized service
[0010] With the continuous deepening of scientific and technological innovation, users' demands for scientific and technological services are becoming increasingly diversified and personalized. For example, researchers may need in-depth analysis of specific technical fields, while enterprises may pay more attention to the commercialization prospects and market applications of technologies. However, the existing knowledge bases mostly adopt a "one-size-fits-all" approach, which is difficult to meet the personalized needs of different users. In addition, users have higher requirements for the timeliness and accuracy of services, and traditional knowledge bases are difficult to provide high-quality customized services in a short period of time.
[0011] In summary, traditional knowledge base construction methods have significant deficiencies in knowledge representation, retrieval efficiency, dynamic updates, etc., and are difficult to meet the growing needs of the science and technology service field. In the existing technology, there is a lack of an intelligent knowledge base construction method that can combine knowledge graphs and RAG technology. As a structured knowledge representation method, knowledge graphs can effectively capture entities, relationships and their attribute information, and support complex queries and reasoning; while RAG technology can combine external knowledge sources to generate high-quality natural language responses. By combining the two, the intelligence level and service quality of the knowledge base can be significantly improved. Summary of the invention
[0012] (I) Purpose of the invention
[0013] The present invention aims to provide a method for constructing a knowledge base in the field of scientific and technological services based on knowledge graph and RAG, realize the structured representation of knowledge through knowledge graph, combine RAG technology to realize efficient knowledge retrieval and generation, and improve the intelligence level and service capabilities of the knowledge base.
[0014] (II) Technical solution
[0015] To achieve the purpose of the invention, the present invention adopts the following technical solutions:
[0016] The first invention object of the present invention is to provide a method for constructing a knowledge base in the field of scientific and technological services based on a knowledge graph and RAG, comprising the following steps:
[0017] Step S1: data collection and preprocessing: data cleaning, deduplication, labeling and standardization;
[0018] Step S2: knowledge graph construction, constructing a scientific and technological knowledge graph and a personal knowledge graph respectively, wherein the scientific and technological knowledge graph contains entity-relationship-entity-attribute quadruple information;
[0019] Step S3 RAG model construction;
[0020] Step S4: knowledge base construction and dynamic update;
[0021] Step S5: Knowledge retrieval and generation. Based on the user query information, the RAG model is used in combination with the user's personal knowledge graph to retrieve relevant scientific and technological achievements from the scientific and technological knowledge graph. The generation module generates a natural language response based on the search results.
[0022] Further, step S1 data collection and preprocessing includes the following steps:
[0023] Step S1.1 collects multi-source data in the field of science and technology services, including data resources such as scientific and technological achievements, technical requirements, technical documents, specifications and standards, and science and technology policies, as well as relevant business knowledge, business procedures, and service content such as scientific and technological innovation, scientific and technological achievement evaluation, scientific and technological assessment, paper retrieval, high-tech enterprise cultivation, science and technology finance, technology contract registration, intellectual property rights, and inspection and testing;
[0024] Step S1.2: cleaning, deduplication, labeling and standardization of the collected data;
[0025] Step S1.3 converts unstructured data (such as text, image) into structured data;
[0026] Furthermore, step S2 of knowledge graph construction includes the following steps:
[0027] Step S2.1 extracts literature materials from the science and technology service database, parses the text and generates scientific and technological achievement text information;
[0028] Step S2.2 uses entity recognition on the corresponding text to extract entity-relationship-entity-attribute quadruplets and store them in the scientific and technological knowledge graph. Through ontology mapping, it supports intelligent matching of the attributes of ontology objects with the fields of data sets, and determines the value conversion logic from data sets to graph data in the form of configuration mapping rules. Through the knowledge graph, unique identification information is created for the data, and the function of rapid positioning is performed to realize that all nodes in the database have sources and basis;
[0029] Step S2.3 constructs a personal knowledge graph for the user based on the user's uploaded results, user needs, user tags, user historical questions, user concerns, etc. The personal knowledge graph can integrate and represent personalized information of various types of users.
[0030] Further, step S3 RAG model construction includes the following steps:
[0031] Step S3.1 designs a RAG model, including a retrieval module and a generation module; the retrieval module is responsible for retrieving knowledge fragments related to the user's question from the knowledge graph, and the input is the user input text + user knowledge graph. Based on the user input text and the user knowledge graph, a user entity-relationship-entity-attribute quadruple is generated, and similarity matching is performed with the science and technology knowledge graph quadruple. The output is knowledge fragments related to the input, such as entities, relationships, and attributes; the generation module generates a reply in natural language form based on the retrieved knowledge fragments, and the input is the user input text + retrieved knowledge fragments, and the output is the generated reply text;
[0032] Step S3.2 initializes the generation module using the pre-trained language model;
[0033] Select a pre-trained language model suitable for the generation task, load the weights of the pre-trained model, initialize the parameters of the generation module, design the input format, splice the user input and the retrieved knowledge fragments as the input of the generation module, design the output format, and ensure that the generated response meets the task requirements;
[0034] Step S3.3: training the retrieval module so that it can efficiently retrieve relevant knowledge from the knowledge graph;
[0035] Step S3.3.1 constructs a training data set, including user input, relevant knowledge fragments and their relevance labels, samples knowledge fragments from the science and technology knowledge graph, and annotates their relevance to the question;
[0036] Step S3.3.2 uses the Transformer-based dual-tower model to design model input and output to support the retrieval of knowledge fragments from the scientific and technological knowledge graph;
[0037] Step S3.3.3 uses contrastive learning combined with cross entropy loss to train the retrieval module to maximize the retrieval probability of relevant knowledge fragments and minimize the retrieval probability of irrelevant knowledge fragments;
[0038] The loss function is as follows:
[0039] L=C1*L InfoNCE +C2*L CE +C3
[0040]
[0041] Where N is the number of samples, z i is the feature representation of the anchor sample, z i + is the feature representation of the positive sample, z j is the feature representation of negative samples, τ is the temperature parameter (temperature), which is used to control the sharpness of the distribution, K is the number of negative samples, C is the number of categories, and yi,c is the true label of sample i. If sample i belongs to category c, then y i,c =1, otherwise y i,c =0, p i,c is the probability that the model predicts that sample i belongs to category c, C1 and C2 are loss weights, and C1+C2=1, and C3 is the adjustment coefficient.
[0042] Step S3.3.4 designs an efficient query algorithm to support rapid retrieval of relevant knowledge fragments from the knowledge graph and uses the graph database Neo4j to optimize query performance.
[0043] Step S3.4 uses a joint training strategy to optimize the retrieval module and the generation module simultaneously, and uses reinforcement learning to optimize the response quality of the generation module.
[0044] Further, step S4 of knowledge base construction and dynamic update includes the following steps:
[0045] Step S4.1 combines the knowledge graph with the RAG model to build a knowledge base in the field of science and technology services;
[0046] Step S4.2 implements dynamic updating of the knowledge base, including real-time access to new data and incremental updating of the knowledge graph;
[0047] Step S4.3 uses the RAG model to optimize the knowledge base to improve retrieval and generation efficiency.
[0048] Further, step S5 of knowledge retrieval and generation includes the following steps:
[0049] Step S5.1 receives a user query and uses the RAG model combined with the user's personal knowledge graph to retrieve relevant scientific and technological achievements from the scientific and technological knowledge graph;
[0050] Step S5.2: the generation module generates a reply in natural language according to the search results;
[0051] Step S5.3 provides multimodal output (such as text, charts) to meet the diverse needs of users.
[0052] Furthermore, after step S5, step S6 feedback and optimization is also included, and specifically includes the following steps:
[0053] Step S6.1 collects user feedback, including query satisfaction and answer accuracy;
[0054] Step S6.2 uses the feedback data to iteratively optimize the RAG model and knowledge graph;
[0055] Step S6.3: Update the knowledge base regularly to keep the knowledge current and accurate.
[0056] The second invention object of the present invention is to provide a knowledge base construction system in the field of science and technology services based on knowledge graph and RAG, which is used to implement the method described above.
[0057] (III) Technical Effect
[0058] Compared with the prior art, the method and system for constructing a knowledge base in the field of science and technology services based on knowledge graph and RAG of the present invention have the following beneficial and significant technical effects:
[0059] (1) Intelligent improvement of information retrieval and consulting services
[0060] In the traditional field of science and technology services, information retrieval and consulting services mostly rely on keyword matching and simple rule engines, which have a low level of intelligence and are difficult to accurately capture the user's true intentions, resulting in insufficient relevance and practicality of the search results. For example, when a user searches for "the application of artificial intelligence in medicine", traditional search tools may only return documents containing the keywords "artificial intelligence" and "medical care", but cannot understand the user's specific needs (such as technical details, application cases or commercial prospects). This keyword-based search method lacks understanding of semantics and context, and often returns a large number of irrelevant or low-quality results. Users need to spend a lot of time on secondary screening, which seriously affects service efficiency and user experience.
[0061] To solve this problem, the present invention significantly improves the accuracy and personalization of search results by combining the scientific and technological knowledge graph with the personal knowledge graph and intelligently matching it with the search information. Specifically, the scientific and technological knowledge graph is a structured knowledge base built based on massive scientific and technological literature, patents, papers and other data, covering entities such as technical fields, scientific research teams, innovative achievements and their interrelationships. The personal knowledge graph is built based on personalized data such as the user's historical behavior, interest tags, and concerns, and can accurately portray the user's needs and preferences. By combining the two, the system can simultaneously understand the professional knowledge in the field of science and technology and the personalized needs of users, providing a solid foundation for subsequent retrieval and services.
[0062] When a user makes a search request, the system will not only analyze the keywords entered by the user, but also combine the scientific and technological knowledge graph and the personal knowledge graph to perform contextual understanding and semantic reasoning. For example, when a user searches for "the application of artificial intelligence in medicine", the system will not only return relevant technical literature, but also provide more targeted search results based on the user's professional background (such as whether he is a medical industry practitioner) and interest preferences (such as whether he is concerned about the commercialization of technology). This intelligent matching mechanism significantly improves the accuracy and practicality of the search results, and meets the user's needs for efficient and accurate services.
[0063] In addition, by combining knowledge graphs and intelligent retrieval technology, the system can flexibly respond to the needs of different users and provide customized service solutions. For example, for scientific researchers, the system will give priority to returning technological frontiers and academic papers; while for corporate users, the system will focus on technology application cases and market analysis reports. This personalized service not only improves the accuracy of retrieval results, but also greatly improves user satisfaction.
[0064] (2) Knowledge graph expression of quadruple structure
[0065] Traditional knowledge graphs mostly use triples (entity-relationship-entity) structures. Although they can represent basic semantic relationships, they have the problem of insufficient information expression in complex scenarios. For example, when expressing "a certain scientific research team developed a certain technology", the triple can only express the "research and development" relationship between "scientific research team" and "technology", but cannot further describe attribute information such as research and development time and technology maturity. This limitation of information expression restricts the application of knowledge graphs in complex scenarios.
[0066] To solve this problem, the present invention innovatively introduces a quadruple structure (entity-relationship-entity-attribute) to represent nodes and edges in the scientific and technological knowledge graph, which significantly improves the expressiveness and practicality of the knowledge graph. The quadruple structure can not only represent the relationship between entities, but also describe the attribute information of the relationship. For example, when expressing "a scientific research team has developed a certain technology", the quadruple can further describe attribute information such as R&D time and technology maturity. This structure enables the knowledge graph to express complex scientific and technological knowledge more comprehensively.
[0067] When constructing the scientific and technological knowledge graph, the system extracts entities, relationships, and attributes from data sources such as scientific literature, patents, and papers, and constructs quadruple. By using ontology mapping technology, the quadruple is intelligently matched with the data set fields to ensure data consistency and scalability. For example, when integrating patent data from different sources, the system can automatically identify and unify different expressions of the same entity to avoid data redundancy and conflict.
[0068] By introducing the quadruple structure, the present invention significantly improves the structuring of the knowledge graph, facilitating knowledge management and retrieval. For example, in technical consultation, the system can quickly retrieve technologies, teams, and application cases related to user needs, and provide detailed attribute information. In addition, the quadruple structure also supports complex queries and reasoning, such as predicting the development trend or potential application scenarios of technologies by analyzing the relationship chain between technologies.
[0069] (3) Combined training of contrastive learning and cross entropy loss
[0070] When training the retrieval module, the traditional cross-entropy loss function can effectively measure the difference between the model prediction and the true label, but its generalization ability is insufficient and it is easy to cause the model to overfit. For example, in the retrieval task, the cross-entropy loss function may make the model overly dependent on specific patterns in the training data and fail to generalize well to unseen data. On the other hand, contrastive learning learns feature representation by bringing positive sample pairs closer and pushing negative sample pairs away, which can effectively improve the generalization ability of the model, but its training process is complex and difficult to converge.
[0071] To solve this problem, the present invention uses contrastive learning combined with cross-entropy loss to train the retrieval module, which not only overcomes the problem of insufficient generalization ability of cross-entropy loss, but also solves the problem of contrastive learning being difficult to converge. Specifically, contrastive learning enables the model to learn more discriminative feature representations by constructing positive sample pairs and negative sample pairs. For example, in a retrieval task, the positive sample pair can be a knowledge fragment that is highly relevant to the user's question, while the negative sample pair is a knowledge fragment that is irrelevant to the user's question. Through contrastive learning, the model can better distinguish between relevant and irrelevant knowledge fragments, thereby improving the accuracy of retrieval.
[0072] During the training process, the cross-entropy loss function is used to measure the difference between the model prediction and the true label, ensuring that the model can accurately predict the similarity of positive sample pairs. For example, in the retrieval task, the cross-entropy loss function can be used to measure the correlation between the knowledge fragment predicted by the model and the user's question. By combining contrastive learning and cross-entropy loss, the model can not only learn discriminative feature representations, but also accurately predict the similarity of positive sample pairs, thereby significantly improving the performance of the retrieval module.
[0073] In addition, by introducing temperature parameters and boundary parameters, the present invention further optimizes the combined training process of contrastive learning and cross entropy loss. The temperature parameter is used to control the sharpness of the similarity distribution, and the boundary parameter is used to control the degree of push of negative sample pairs. By adjusting these parameters, the model can better balance the generalization ability and convergence speed, thereby improving training efficiency and retrieval performance.
[0074] By matching the scientific and technological knowledge graph with personal knowledge graph and retrieval information, introducing the quadruple structure, and combining contrastive learning with cross entropy loss training, the present invention significantly improves the intelligence level and service quality of the knowledge base in the field of scientific and technological services. These innovative technologies provide strong support for the intelligent transformation of the scientific and technological service industry and have broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 It is a schematic diagram of the overall architecture of a specific embodiment of the present invention. DETAILED DESCRIPTION
[0076] In order to better understand the present invention, the content of the present invention is further explained in conjunction with the embodiments below. In the accompanying drawings, the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be understood as limitations on the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present invention. The structure and technical solution of the present invention are further described in detail below in conjunction with the accompanying drawings, and the following embodiments of the present invention are given.
[0077] Example 1
[0078] like Figure 1 As shown, a method for constructing a knowledge base in the field of scientific and technological services based on knowledge graph and RAG includes the following steps:
[0079] Step S1: data collection and preprocessing: data cleaning, deduplication, labeling and standardization;
[0080] Step S2: knowledge graph construction, constructing a scientific and technological knowledge graph and a personal knowledge graph respectively, wherein the scientific and technological knowledge graph contains entity-relationship-entity-attribute quadruple information;
[0081] Step S3 RAG model construction;
[0082] Step S4: knowledge base construction and dynamic update;
[0083] Step S5: Knowledge retrieval and generation. Based on the user query information, the RAG model is used in combination with the user's personal knowledge graph to retrieve relevant scientific and technological achievements from the scientific and technological knowledge graph. The generation module generates a natural language response based on the search results.
[0084] Further, step S1 data collection and preprocessing includes the following steps:
[0085] Step S1.1 collects multi-source data in the field of science and technology services, including data resources such as scientific and technological achievements, technical requirements, technical documents, specifications and standards, and science and technology policies, as well as relevant business knowledge, business procedures, and service content such as scientific and technological innovation, scientific and technological achievement evaluation, scientific and technological assessment, paper retrieval, high-tech enterprise cultivation, science and technology finance, technology contract registration, intellectual property rights, and inspection and testing;
[0086] Step S1.2: cleaning, deduplication, labeling and standardization of the collected data;
[0087] Specifically, remove duplicate data records from the collected data to ensure the uniqueness of the data, fill in missing values (such as using the mean, median or mode) or delete them, detect and process outliers (such as using box plots or the 3σ principle), and unify the data format; manually annotate some data, such as technical classification, demand category, sentiment label, etc.
[0088] Perform word segmentation, stop word removal, and stem extraction on text data. For example, use NLTK or Jieba to perform Chinese word segmentation. Perform normalization or standardization on numerical data, such as Min-Max normalization or Z-score standardization. Perform One-Hot Encoding or Label Encoding on categorical data.
[0089] Step S1.3 converts unstructured data (such as text, image) into structured data;
[0090] Use pre-trained language models (such as BERT, Word2Vec) to convert text into vector representations, for example, use BERT to generate embedded representations of text; use convolutional neural networks (CNN) to convert images into vector representations, for example, use ResNet to extract image features; use multi-layer perceptrons (MLP) or embedding layers to convert tabular data into vector representations, for example, concatenate the numerical and categorical features in the table into vectors.
[0091] Through the above steps, multi-source data in the field of scientific and technological services is collected, covering scientific and technological achievements, demand-side data and related business knowledge; the data is cleaned, deduplicated, labeled and standardized to ensure data quality. Finally, unstructured data (such as text, images) is converted into structured data, laying the foundation for subsequent feature extraction and modeling.
[0092] Furthermore, step S2 of knowledge graph construction includes the following steps:
[0093] Step S2.1 extracts literature materials from the science and technology service database, parses the text and generates scientific and technological achievement text information;
[0094] Step S2.2 uses entity recognition on the corresponding text to extract entity-relationship-entity-attribute quadruplets and store them in the scientific and technological knowledge graph. Through ontology mapping, it supports intelligent matching of the attributes of ontology objects with the fields of data sets, and determines the value conversion logic from data sets to graph data in the form of configuration mapping rules. Through the knowledge graph, unique identification information is created for the data, and the function of rapid positioning is performed to realize that all nodes in the database have sources and basis;
[0095] Specifically, a named entity recognition (NER) model is used to extract entities from text, such as technology names, organization names, personnel names, etc., and a relationship extraction model is used to identify the relationships between entities, such as "development", "application", etc., and to extract entity attributes from text, such as the maturity of the technology, application fields, technical descriptions, etc. For example, ("deep learning", "applied to", "image recognition", "technical description: an image classification method based on convolutional neural networks.").
[0096] Through ontology mapping, entities, relationships, and attributes are matched with ontology objects in the knowledge graph; for example, "artificial intelligence algorithm" is mapped to the "technology" node in the knowledge graph; the mapping rules from dataset fields to knowledge graph attributes are configured, and the value conversion logic is determined. For example, the "technology maturity" field is mapped to the "maturity" attribute in the knowledge graph, and unique identification information (such as URI) is created for each node in the knowledge graph to support fast positioning and query.
[0097] The above steps introduce a quadruple structure (entity-relationship-entity-attribute) to represent nodes and edges in the scientific and technological knowledge graph, which significantly improves the expressiveness and practicality of the knowledge graph. The quadruple structure can not only represent the relationship between entities, but also describe the attribute information of the relationship. This structure enables the knowledge graph to express complex scientific and technological knowledge more comprehensively. By introducing the quadruple structure, the present invention significantly improves the structuring of the knowledge graph and facilitates the management and retrieval of knowledge. In addition, the quadruple structure also supports complex queries and reasoning, such as predicting the development trend or potential application scenarios of technology by analyzing the relationship chain between technologies.
[0098] Step S2.3 constructs a personal knowledge graph for the user based on the user's uploaded results, user needs, user tags, user historical questions, user concerns, etc. The personal knowledge graph can integrate and represent personalized information of various types of users.
[0099] Specifically, entities are extracted from user data, such as user names, technology names, organization names, etc., and the relationships between entities are identified, such as "user A follows technology B"; attributes of entities are extracted, such as user technology preferences, historical behaviors, etc. For example, the technologies that users follow are associated with technology nodes in the knowledge graph to build a personal knowledge graph.
[0100] The above steps can generate a user portrait based on the personal knowledge graph to characterize the user's personalized information, such as the user's technical preferences, historical behaviors, areas of interest, etc., and use the personal knowledge graph to provide users with personalized recommendations for scientific and technological achievements.
[0101] Further, step S3 RAG model construction includes the following steps:
[0102] Step S3.1 designs a RAG model, including a retrieval module and a generation module; the retrieval module is responsible for retrieving knowledge fragments related to the user's question from the knowledge graph, and the input is the user input text + user knowledge graph. Based on the user input text and the user knowledge graph, a user entity-relationship-entity-attribute quadruple is generated, and similarity matching is performed with the science and technology knowledge graph quadruple. The output is knowledge fragments related to the input, such as entities, relationships, and attributes; the generation module generates a reply in natural language form based on the retrieved knowledge fragments, and the input is the user input text + retrieved knowledge fragments, and the output is the generated reply text;
[0103] When a user makes a search request, the system will not only analyze the keywords entered by the user, but also combine the personal knowledge graph to perform contextual understanding and semantic reasoning. For example, when a user searches for "application of artificial intelligence in medicine", the system will not only return relevant technical literature, but also provide more targeted search results based on the user's professional background (such as whether he is a medical practitioner) and interest preferences (such as whether he is concerned about the commercialization of technology). This intelligent matching mechanism significantly improves the accuracy and practicality of search results, and meets the user's needs for efficient and accurate services.
[0104] Step S3.2 initializes the generation module using the pre-trained language model;
[0105] Select a pre-trained language model suitable for the generation task, load the weights of the pre-trained model, initialize the parameters of the generation module, design the input format, splice the user input and the retrieved knowledge fragments as the input of the generation module, design the output format, and ensure that the generated response meets the task requirements;
[0106] Step S3.3: training the retrieval module so that it can efficiently retrieve relevant knowledge from the knowledge graph;
[0107] Step S3.3.1 constructs a training data set, including user input, relevant knowledge fragments and their relevance labels, samples knowledge fragments from the science and technology knowledge graph, and annotates their relevance to the question;
[0108] Step S3.3.2 uses the Transformer-based dual-tower model to design model input and output to support the retrieval of knowledge fragments from the scientific and technological knowledge graph;
[0109] Step S3.3.3 uses contrastive learning combined with cross entropy loss to train the retrieval module to maximize the retrieval probability of relevant knowledge fragments and minimize the retrieval probability of irrelevant knowledge fragments;
[0110] The loss function is as follows:
[0111] L=C1*L InfoNCE +C2*L CE +C3
[0112]
[0113] Where N is the number of samples, z i is the feature representation of the anchor sample, z i + is the feature representation of the positive sample, z j is the feature representation of negative samples, τ is the temperature parameter (temperature), which is used to control the sharpness of the distribution, K is the number of negative samples, C is the number of categories, and y i,c is the true label of sample i. If sample i belongs to category c, then y i,c =1, otherwise y i,c =0, p i,c is the probability that the model predicts that sample i belongs to category c, C1 and C2 are loss weights, and C1+C2=1, and C3 is the adjustment coefficient.
[0114] Step S3.3.4 designs an efficient query algorithm to support rapid retrieval of relevant knowledge fragments from the knowledge graph and uses the graph database Neo4j to optimize query performance.
[0115] Step S3.4 uses a joint training strategy to optimize the retrieval module and the generation module at the same time, and uses reinforcement learning to optimize the answer quality of the generation module;
[0116] The above steps use contrastive learning combined with cross-entropy loss to train the retrieval module, which not only overcomes the problem of insufficient generalization ability of cross-entropy loss, but also solves the problem of contrastive learning being difficult to converge. Specifically, contrastive learning enables the model to learn more discriminative feature representations by constructing positive and negative sample pairs; the cross-entropy loss function is used to measure the difference between the model prediction and the true label, ensuring that the model can accurately predict the similarity of positive sample pairs. By combining contrastive learning and cross-entropy loss, the model can not only learn discriminative feature representations, but also accurately predict the similarity of positive sample pairs, thereby significantly improving the performance of the retrieval module.
[0117] Further, step S4 of knowledge base construction and dynamic update includes the following steps:
[0118] Step S4.1 combines the knowledge graph with the RAG model to build a knowledge base in the field of science and technology services;
[0119] Step S4.2 implements dynamic updating of the knowledge base, including real-time access to new data and incremental updating of the knowledge graph;
[0120] Step S4.3 uses the RAG model to optimize the knowledge base to improve retrieval and generation efficiency.
[0121] The above steps can ensure the timeliness and accuracy of the knowledge base, support real-time access to new data and incremental updates of the knowledge graph; improve the retrieval and generation efficiency of the knowledge base, enhance the user experience, and at the same time, by optimizing the RAG model, reduce the retrieval time and improve the accuracy and fluency of the generated answers.
[0122] Further, step S5 of knowledge retrieval and generation includes the following steps:
[0123] Step S5.1 receives a user query and uses the RAG model combined with the user's personal knowledge graph to retrieve relevant scientific and technological achievements from the scientific and technological knowledge graph;
[0124] Step S5.2: the generation module generates a reply in natural language according to the search results;
[0125] Step S5.3 provides multimodal output (such as text, charts) to meet the diverse needs of users.
[0126] Through the above steps, we can combine user personalized information to provide accurate search results, and provide multimodal output (such as text, charts, images) to enhance the readability and practicality of the response, meet user needs for information in different forms, and improve user experience.
[0127] Furthermore, after step S5, step S6 feedback and optimization is also included, and specifically includes the following steps:
[0128] Step S6.1 collects user feedback, including query satisfaction and answer accuracy;
[0129] Step S6.2 uses the feedback data to iteratively optimize the RAG model and knowledge graph;
[0130] Step S6.3: Update the knowledge base regularly to keep the knowledge current and accurate.
[0131] Through the above steps, the RAG model and knowledge graph can be optimized according to user feedback, improving the accuracy of the system and user experience. Iterative optimization enables the system to continuously adapt to new needs and scenarios, and ensures that the knowledge in the knowledge base is always up-to-date and accurate, avoiding the provision of outdated or erroneous information.
[0132] Example 2
[0133] A knowledge base construction system in the field of scientific and technological services based on knowledge graph and RAG, which is used to implement the method described above.
[0134] Through the above embodiments, the purpose of the present invention is fully and effectively achieved. Those skilled in the art can understand that the present invention includes but is not limited to the contents described in the drawings and the above specific embodiments. Although the present invention has been described with respect to the most practical and preferred embodiments currently considered, it should be understood that the present invention is not limited to the disclosed embodiments, and any modification that does not deviate from the functional and structural principles of the present invention will be included in the scope of the claims.
Claims
1. A method for constructing a knowledge base in the field of scientific and technological services based on knowledge graph and RAG, characterized in that: The following steps are involved: Step S1: data collection and preprocessing: data cleaning, deduplication, labeling and standardization; Step S2: knowledge graph construction, constructing a scientific and technological knowledge graph and a personal knowledge graph respectively, wherein the scientific and technological knowledge graph contains entity-relationship-entity-attribute quadruple information; Step S3 RAG model construction; Step S4: knowledge base construction and dynamic update; Step S5: Knowledge retrieval and generation. Based on the user query information, the RAG model is used in combination with the user's personal knowledge graph to retrieve relevant scientific and technological achievements from the scientific and technological knowledge graph. The generation module generates a natural language response based on the search results.
2. The method for constructing a knowledge base in the field of scientific and technological services based on knowledge graph and RAG according to claim 1, characterized in that: Step S1: Data collection and preprocessing includes the following steps: Step S1.1 collects multi-source data in the field of science and technology services, including data resources such as scientific and technological achievements, technical requirements, technical documents, specifications and standards, and science and technology policies, as well as relevant business knowledge, business procedures, and service content such as scientific and technological innovation, scientific and technological achievement evaluation, scientific and technological assessment, paper retrieval, high-tech enterprise cultivation, science and technology finance, technology contract registration, intellectual property rights, and inspection and testing; Step S1.2: cleaning, deduplication, labeling and standardization of the collected data; Step S1.3 converts the unstructured data into structured data.
3. The method for constructing a knowledge base in the field of scientific and technological services based on knowledge graph and RAG according to claim 1 is characterized in that: Step S2: Knowledge graph construction includes the following steps: Step S2.1 extracts literature materials from the science and technology service database, parses the text and generates scientific and technological achievement text information; Step S2.2 uses entity recognition on the corresponding text to extract entity-relationship-entity-attribute quadruplets and store them in the scientific and technological knowledge graph. Through ontology mapping, it supports intelligent matching of the attributes of ontology objects with the fields of data sets, and determines the value conversion logic from data sets to graph data in the form of configuration mapping rules. Through the knowledge graph, unique identification information is created for the data, and the function of rapid positioning is performed to realize that all nodes in the database have sources and basis; Step S2.3 constructs a personal knowledge graph for the user based on the user's uploaded results, user needs, user tags, user historical questions, user concerns, etc. The personal knowledge graph can integrate and represent various types of user personalized information. Through the knowledge graph system, in the subsequent retrieval generation steps, the content that is directly or indirectly related to the current user input content can be quickly obtained.
4. The method for constructing a knowledge base in the field of scientific and technological services based on knowledge graph and RAG according to claim 1, characterized in that: Step S3 RAG model construction includes the following steps: Step S3.1 designs a RAG model, including a retrieval module and a generation module; the retrieval module is responsible for retrieving knowledge fragments related to the user's question from the knowledge graph, and the input is the user input text + user knowledge graph. Based on the user input text and the user knowledge graph, a user entity-relationship-entity-attribute quadruple is generated, and similarity matching is performed with the science and technology knowledge graph quadruple. The output is knowledge fragments related to the input, such as entities, relationships, and attributes; the generation module generates a reply in natural language form based on the retrieved knowledge fragments, and the input is the user input text + retrieved knowledge fragments, and the output is the generated reply text; Step S3.2 initializes the generation module using the pre-trained language model; Select a pre-trained language model suitable for the generation task, load the weights of the pre-trained model, initialize the parameters of the generation module, design the input format, splice the user input and the retrieved knowledge fragments as the input of the generation module, design the output format, and ensure that the generated response meets the task requirements; Step S3.3: training the retrieval module so that it can efficiently retrieve relevant knowledge from the knowledge graph; Step S3.4 uses a joint training strategy to optimize the retrieval module and the generation module simultaneously, and uses reinforcement learning to optimize the response quality of the generation module.
5. The method for constructing a knowledge base in the field of scientific and technological services based on knowledge graph and RAG according to claim 4 is characterized in that: Step S3.3: Training the retrieval module to enable it to efficiently retrieve relevant knowledge from the knowledge graph includes the following steps: Step S3.3.1 constructs a training data set, including user input, relevant knowledge fragments and their relevance labels, samples knowledge fragments from the science and technology knowledge graph, and annotates their relevance to the question; Step S3.3.2 uses the Transformer-based dual-tower model to design model input and output to support the retrieval of knowledge fragments from the scientific and technological knowledge graph; Step S3.3.3 uses contrastive learning combined with cross entropy loss to train the retrieval module to maximize the retrieval probability of relevant knowledge fragments and minimize the retrieval probability of irrelevant knowledge fragments; The loss function is as follows: L=C1*L InfoNCE +C2*L CE +C3 Where N is the number of samples, z i is the feature representation of the anchor sample, z i + is the feature representation of the positive sample, z j is the characteristic representation of negative samples, τ is the temperature parameter used to control the sharpness of the distribution, K is the number of negative samples, C is the number of categories, and y i,c is the true label of sample i. If sample i belongs to category c, then y i,c =1, otherwise y i,c =0, p i,c is the probability that the model predicts that sample i belongs to category c, C1 and C2 are loss weights, and C1+C2=1, C3 is the adjustment coefficient; Step S3.3.4 designs an efficient query algorithm to support rapid retrieval of relevant knowledge fragments from the knowledge graph and uses the graph database Neo4j to optimize query performance.
6. The method for constructing a knowledge base in the field of scientific and technological services based on knowledge graph and RAG according to claim 1, characterized in that: Step S4: knowledge base construction and dynamic update includes the following steps: Step S4.1 combines the knowledge graph with the RAG model to build a knowledge base in the field of science and technology services; Step S4.2 implements dynamic updating of the knowledge base, including real-time access to new data and incremental updating of the knowledge graph; Step S4.3 uses the RAG model to optimize the knowledge base to improve retrieval and generation efficiency.
7. The method for constructing a knowledge base in the field of scientific and technological services based on knowledge graph and RAG according to claim 1, characterized in that: Step S5: Knowledge retrieval and generation includes the following steps: Step S5.1 receives a user query and uses the RAG model combined with the user's personal knowledge graph to retrieve relevant scientific and technological achievements from the scientific and technological knowledge graph; Step S5.2: the generation module generates a reply in natural language according to the search results; Step S5.3 provides multimodal output such as text and charts to meet the diverse needs of users.
8. The method for constructing a knowledge base in the field of scientific and technological services based on knowledge graph and RAG according to claim 1, characterized in that: Step S5 is followed by step S6 of feedback and optimization, which specifically includes the following steps: Step S6.1 collects user feedback, including query satisfaction and answer accuracy; Step S6.2 uses the feedback data to iteratively optimize the RAG model and knowledge graph; Step S6.3: Update the knowledge base regularly to keep the knowledge current and accurate.
9. A knowledge base construction system for science and technology service field based on knowledge graph and RAG, characterized in that: Used to implement the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Scientific and technological character knowledge graph construction method and device based on deep learning model, and terminal
CN113254667A
Knowledge graph construction, retrieval and visualization method and system for science and technology service
CN115309885A
Psychological counseling dialogue generation method and device, equipment and storage medium
CN118553412A
Response enhancement method and device based on RAG technology
CN119005330A
Method and apparatus for processing data based on knowledge graph, electronic device and medium
US20220350829A1
Cited By
Artificial intelligence system for multi-domain questions and answers
CN120144726A
Scientific and technological event time sequence prediction method based on RAG
CN120316233A
Enterprise scientific and technological achievement adaptation method based on big data accurate retrieval and query
CN120336546A
Water conservancy project knowledge graph large model construction method and device
CN120372019A
Intelligent question learning method and system based on large language model
CN120541193A