Retrieval generation system and method based on RAG and ontology driving battery material knowledge graph
By introducing a knowledge graph retrieval system based on RAG and ontology drive in the field of battery materials, the problems of data dispersion, low retrieval efficiency, lagging knowledge update and insufficient semantic understanding in the existing search methods are solved, efficient and accurate knowledge retrieval and management are achieved, and research and innovation efficiency is improved.
Patent Information
- Application Number
- CN202411739594.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-09
AI Technical Summary
The existing battery material retrieval methods face the problems of data dispersion, low retrieval efficiency, lagging knowledge updates and insufficient semantic understanding, which limits research efficiency and innovation speed.
A search and generation system for battery material knowledge graphs based on RAG and ontology driven is designed. Through data acquisition, preprocessing, entity recognition, relationship extraction, knowledge graph construction and query engine module integration, efficient knowledge management and intelligent retrieval are achieved.
It improves the integration of data and the accuracy of retrieval, supports dynamic expansion and real-time update of multimodal data, solves the problems of data dispersion, low retrieval efficiency, lagging knowledge update and insufficient semantic understanding, and improves the research efficiency and innovation capabilities in the field of battery materials.
Smart Images

Figure CN119964689A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a retrieval generation system and method based on RAG and ontology-driven battery material knowledge graph. Background Art
[0002] Searches in the field of battery materials have important academic and industrial significance, such as promoting scientific research and innovation and improving R&D efficiency.
[0003] Existing retrieval methods include manual processing, keyword search, single knowledge graph, and single retrieval enhancement. Among them, the information sources related to battery materials are extensive, covering academic papers, patents, industry reports, experimental data and other forms. Most of these data lack unified formats and standards, making it difficult to collect and process information automatically. Researchers need to manually screen and analyze a large amount of literature and data, which not only consumes time and energy, but also greatly prolongs the R&D cycle and affects the speed of innovation. The keyword search method requires users to accurately enter relevant keywords to obtain information. This method has limited processing capabilities for multi-source heterogeneous data, especially when facing complex or ambiguous queries. The retrieval effect is poor. Research in the field of battery materials often requires information based on context and semantic levels, and traditional keyword searches often cannot meet this demand, resulting in inaccurate or inappropriate retrieval results. Although knowledge graphs can provide structured information queries, existing technologies still rely on a large number of manual operations in the construction and update of knowledge graphs, which makes the update speed of graphs unable to keep up with the rapid development of the field of battery materials. In addition, due to the lack of support from domain knowledge ontology, knowledge graphs often lack semantic consistency when facing new data, affecting their stability and reusability. Traditional knowledge graphs often have difficulty providing accurate contextual answers when responding to natural language queries, especially when it comes to complex semantic understanding and reasoning. If the current RAG (Retrieval-Augmented Generation) technology is not combined with domain knowledge ontology, the answers it generates may lack professionalism and accuracy, making it difficult to meet the demand for accurate and in-depth information in the field of battery technology.
[0004] In summary, existing retrieval methods face problems such as data dispersion, low retrieval efficiency, lagging knowledge update and insufficient semantic understanding, which limit the research efficiency and innovation speed in the field of battery materials.
[0005] Therefore, it is necessary to design a new system to solve the problems faced by existing retrieval methods, such as data dispersion, low retrieval efficiency, lagging knowledge update and insufficient semantic understanding. Summary of the invention
[0006] The purpose of the present invention is to overcome the defects of the prior art and provide a retrieval generation system and method based on RAG and ontology-driven battery material knowledge graph.
[0007] To achieve the above-mentioned purpose, the present invention adopts the following technical scheme: a retrieval generation system based on RAG and ontology-driven battery material knowledge graph, comprising: a data acquisition unit, a data processing unit, an ontology engine, a knowledge graph construction unit, a data storage unit, a query engine unit and an application interface;
[0008] The data acquisition unit is used to provide a data acquisition interface and collect relevant data of various battery materials;
[0009] The data processing unit is used to perform preprocessing and NLP processing on the collected data to obtain processing results and vector data;
[0010] The ontology engine is used to construct an ontology library in the battery field;
[0011] The knowledge graph construction unit is used to perform entity recognition and relationship extraction in the battery field by fine-tuning BERT and its variant models, and to achieve relationship fusion and alignment in combination with the battery field ontology library to obtain a knowledge graph in the battery field;
[0012] The data storage unit is used to store the knowledge graph and the vector data to form a knowledge graph database and a vector database;
[0013] The application interface is used to input a query request;
[0014] The query engine unit is used to call the database in the data storage unit to perform retrieval according to the query request, and return the retrieval result to the application interface.
[0015] Its further technical solution is: the relevant data of various types of battery materials include academic papers, patent documents, experimental data, network data and non-public information.
[0016] Its further technical solution is: the data processing unit is used to process the collected data into structured and unstructured data using regex, OCR model, and pdf recognition series, and use the spaCy tool to perform word segmentation and part-of-speech tagging to obtain processing results and vector data.
[0017] Its further technical solution is: the ontology engine is used to define the concepts, attributes and relationships in the battery field, and to construct a preliminary ontology, and to design an ontology reasoning engine for the ontology, and to use the ontology reasoning engine to map, disambiguate and fuse the ontology to build an ontology library in the battery field.
[0018] Its further technical solution is: the knowledge graph construction unit is used to perform entity recognition and relationship extraction in the battery field on the processing results by fine-tuning BERT and its variant models, and map, disambiguate and fuse the entities in combination with the battery field ontology library to obtain a knowledge graph in the battery field.
[0019] Its further technical solution is: the knowledge graph construction unit is used to achieve entity matching through rules, and the automatic mapping uses similarity calculation or machine learning for matching; the entity mapping is determined by analyzing the context, attributes and semantic similarity; and entity fusion is performed by merging similar or identical entities.
[0020] A further technical solution is: the query engine unit includes a GraphRAG module, and the GraphRAG module is used to perform relational reasoning from the knowledge graph database according to the query request to obtain a search result;
[0021] The query engine unit further includes a VectorRAG module, and the VectorRAG module is used to perform a similarity query in the vector data according to the query request to obtain a retrieval result.
[0022] Its further technical solution is: the query engine unit also includes a HybridRAG module, which is used to input the query request into the vector database through the VectorRAG module for similarity retrieval to obtain documents or fragments similar to the query request; transmit the documents or fragments similar to the query request to the GraphRAG module for reasoning and knowledge fusion to obtain a fusion result; and fuse the documents or fragments similar to the query request and the fusion result to obtain a retrieval result.
[0023] Its further technical solution is: the query engine unit also includes a FusionRAG module, which is used to directly search from the knowledge graph data and the vector database, fuse the retrieved results, and use the generative model to generate the final answer based on the fused results.
[0024] The present invention also provides a retrieval method based on RAG and the knowledge graph of the ontology-driven battery material, including:
[0025] Get the query request;
[0026] The corresponding query engine calls the database to perform a search according to the query request, and returns the search result;
[0027] Among them, the database includes a knowledge graph database and a vector database. The knowledge graph database is obtained by collecting relevant data of various battery materials, preprocessing and NLP processing the collected data to obtain processing results, and performing entity recognition and relationship extraction in the battery field by fine-tuning BERT and its variant models, and combining the battery field ontology library built by the ontology engine to achieve relationship fusion and alignment to obtain the knowledge graph in the battery field and store the formed database; the vector database is obtained by collecting relevant data of various battery materials, preprocessing and NLP processing the collected data, and storing the formed database.
[0028] Compared with the prior art, the present invention has the following beneficial effects: the present invention realizes efficient management and intelligent retrieval of knowledge in the field of battery materials by integrating functions such as data collection, processing, ontology construction, knowledge graph generation, storage and query; the RAG technology and ontology-driven method are used to improve the integration of data and the accuracy of retrieval, and support the dynamic expansion and real-time update of multimodal data; users can input query requests through the application interface, and the query engine will call the knowledge graph and vector data in the storage unit for efficient retrieval and return accurate results. The overall solution is to solve the problems faced by the existing retrieval methods, such as data dispersion, low retrieval efficiency, delayed knowledge update and insufficient semantic understanding.
[0029] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0031] Figure 1 A schematic block diagram of a retrieval and generation system for a knowledge graph of battery materials driven by RAG and ontology provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0033] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0034] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0035] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0036] Retrieval in the field of battery materials faces the problems of scattered data and inconsistent formats, which makes it difficult to collect and process information. Existing retrieval methods such as manual screening, keyword search, knowledge graphs and RAG technology have problems of low efficiency and insufficient accuracy. Traditional keyword retrieval has difficulty handling complex or fuzzy queries, especially lacking semantic understanding. The construction of knowledge graphs relies on a lot of manual work, has a slow update speed and lacks domain ontology support, which affects its stability and reusability. When RAG technology is not combined with a knowledge graph based on domain knowledge ontology mapping, it is difficult to provide accurate and professional answers. In general, existing retrieval methods restrict the scientific research efficiency and innovation process in the field of battery materials.
[0037] To this end, an embodiment of the present invention provides a retrieval generation system based on RAG and ontology-driven battery material knowledge graph to solve the problems faced by existing retrieval methods, such as data dispersion, low retrieval efficiency, delayed knowledge update and insufficient semantic understanding.
[0038] Specifically, the system collects multi-source heterogeneous data such as academic papers, patents, industry reports, and experimental data through the data collection unit to form a unified data management. NLP technology (such as spaCy, regex, OCR and other tools) is used to structure the collected data and generate vector data to improve the processing efficiency of retrieval data. By building an ontology library in the battery field and using the ontology reasoning engine for entity alignment and relationship fusion, the domain knowledge is standardized and refined. Combined with fine-tuning BERT and its variant models, entity recognition and relationship extraction are performed on the data to generate a knowledge graph in the battery field, which improves the accuracy of retrieval and the real-time nature of knowledge update. The query engine combines graph reasoning and vector similarity query through GraphRAG and VectorRAG modules to provide a multi-dimensional query method to improve the retrieval effect. Through automated entity fusion technology, ambiguity is eliminated and the semantic understanding of battery material knowledge is enhanced, making the retrieval results more accurate. Neo4j and Milvus are used to store knowledge graphs and vector data to improve query efficiency and support rapid retrieval of large-scale data. The query engine fuses the knowledge graph with vector data through the HybridRAG and FusionRAG modules, and provides comprehensive retrieval results based on the generative model, thereby improving the accuracy and response speed of knowledge retrieval.
[0039] See also Figure 1 , Figure 1 A schematic block diagram of a retrieval and generation system for a knowledge graph of battery materials driven by RAG and ontology provided in an embodiment of the present invention.
[0040] The knowledge graph is designed based on a knowledge ontology designed by experts in the field of battery technology, aiming to provide a top-down structured framework for the field of battery materials. This framework clarifies key entities, attributes, and relationships between battery components, performance parameters, and production processes. By using the guidance of the domain ontology, the semantic structure of the constructed knowledge graph can be ensured to facilitate data integration and intelligent analysis.
[0041] The system of this embodiment ensures that the knowledge graph is highly structured and consistent during the construction process, thereby achieving data discoverability, accessibility, interoperability and reusability, in line with the FAIR principle. Through this structured data design, human users can understand it intuitively, and computer systems can process and analyze it efficiently, ensuring the stability of knowledge when multi-source data is integrated. In addition, the CLEAR principle, a special data processing principle for the battery field, is also crucial. CLEAR stands for the following meanings:
[0042] Contextual: Ensure that the data has complete contextual information in battery materials and related fields to facilitate users to understand the background and meaning of the data.
[0043] Linkable: Data can be seamlessly linked with other related data and knowledge graphs to promote the sharing and integration of cross-domain information.
[0044] Explorable: Data should be highly explorable to facilitate users to query, retrieve and discover relevant information.
[0045] Accessible: Ensure that data access rights are clear and easily accessible, so that users can obtain the information they need without complicated procedures or formalities.
[0046] Reusable: Data should be designed to be easy to reuse in different projects and studies to reduce duplication of work.
[0047] In one embodiment, see Figure 1 , the above-mentioned retrieval generation system based on RAG and ontology-driven battery material knowledge graph includes: a data acquisition unit 10, a data processing unit 20, an ontology engine 30, a knowledge graph construction unit 40, a data storage unit 50, a query engine unit 60 and an application interface 70;
[0048] The data acquisition unit 10 is used to provide a data acquisition interface and collect relevant data of various battery materials;
[0049] The data processing unit 20 is used to perform preprocessing and NLP processing on the collected data to obtain processing results and vector data;
[0050] An ontology engine 30, used to construct an ontology library in the battery field;
[0051] The knowledge graph construction unit 40 is used to perform entity recognition and relationship extraction in the battery field by fine-tuning BERT and its variant models, and to achieve relationship fusion and alignment in combination with the battery field ontology library to obtain a knowledge graph in the battery field;
[0052] A data storage unit 50, used to store knowledge graphs and vector data to form a knowledge graph database and a vector database;
[0053] Application interface 70, used to input query requests;
[0054] The query engine unit 60 is used to call the database in the data storage unit 50 to perform retrieval according to the query request, and return the retrieval result to the application interface 70.
[0055] Specifically, the relevant data of various battery materials include academic papers, patent documents, experimental data, network data and non-public information.
[0056] The data acquisition unit 10 is mainly designed with a unified data acquisition interface.
[0057] Connect to databases such as Scopus and Web of Science through API to obtain academic papers; connect to the patent office API and patent database to obtain patent documents; establish a standardized experimental data upload interface to obtain experimental data; develop a web crawler system, collect and update it regularly, and obtain network data; experts edit materials from industry conferences / seminars to obtain non-public information.
[0058] In one embodiment, the data processing unit 20 is used to process structured and unstructured data using regex, OCR model, and PDF recognition series on the collected data, and to perform word segmentation and part-of-speech tagging using spaCy tools to obtain processing results and vector data.
[0059] In one embodiment, the ontology engine 30 is used to define concepts, attributes and relationships in the battery field, construct a preliminary ontology, and design an ontology reasoning engine for the ontology. The ontology reasoning engine is used to map, disambiguate and merge the ontology to build an ontology library in the battery field.
[0060] Specifically, ontology construction is the basis for the ontology engine 30 to be implemented, and is intended to define and organize core concepts in a domain and the relationships between them.
[0061] Before ontology construction, we first need to identify the core concepts in the battery field. For example, battery materials, performance parameters, charging cycles, energy density, etc. Ontology construction tools can help create and manage the structure of the ontology. Here, Protégé is used, which is an open source ontology editor that is widely used in ontology construction and management.
[0062] Define the class and properties:
[0063] Classes: The battery domain may include classes such as BatteryMaterial, BatteryType, BatteryPerformance, etc.
[0064] Properties: Properties are characteristics or properties of a class, such as hasComposition (chemical composition), hasCapacity (capacity), hasCycleLife (service life), etc.
[0065] Relations: The relationship between classes, such as isTypeOf (belongs to a certain type), isComponentOf (is part of a certain component), etc.
[0066] Analyze the existing ontology to understand its class hierarchy, attributes, relationships and semantics. This step helps us identify the similarities and differences in the ontology and provides a basis for subsequent alignment and fusion work.
[0067] Ontology reasoning engine is a tool used to reason and infer knowledge. It can reason according to defined ontology rules and automatically infer new knowledge.
[0068] Ontology alignment refers to the process of matching and connecting concepts, attributes, and relationships in two or more ontologies to ensure that knowledge in different ontologies can be shared and interoperable. It includes mapping, disambiguation, and fusion.
[0069] Ontology Mapping: The core task of ontology mapping is to align similar concepts in different ontologies to ensure semantic consistency between them. For example, aligning BatteryType in one battery domain ontology with BatteryModel in another ontology.
[0070] Ontology disambiguation: In an ontology, if the same word or term has different meanings in different contexts, disambiguation is required. For example, the word "battery" may refer to the battery itself or the chemical reaction process of the battery. Its actual meaning needs to be determined by the context.
[0071] Ontology fusion refers to merging multiple ontologies into a unified ontology, eliminating redundancy and resolving conflicts during the merging process to ensure the consistency and integrity of the ontology.
[0072] For example, different types of batteries (such as power batteries, consumer electronics batteries, and energy storage batteries) can be merged to retain commonalities while deleting duplicate parts to ensure that cross-domain data can be seamlessly integrated.
[0073] In order to achieve efficient and accurate ontology alignment, in the ontology engine 30, some algorithms and rules are needed to measure and evaluate the similarity between ontologies.
[0074] Text similarity calculation: Use algorithms such as Jaccard similarity and cosine similarity to measure the text similarity of concepts in different ontologies. For example, in the battery field, "lithium battery" and "Li-ion Battery" may have a high text similarity, so they can be considered to be aligned.
[0075] Structural alignment: Compare the hierarchical structures and relationships between ontologies and quantify the structural similarity through algorithms such as tree edit distance. For ontologies with different hierarchical structures, the tree edit distance algorithm can be used to evaluate their similarities at the concept level to determine how to align them.
[0076] Ontology matching tools: Use specialized tools to automate the ontology alignment task. These tools can help us identify relevant concepts in different ontologies and match them, thus reducing manual intervention.
[0077] Protégé Alignment Plug-in: As a plug-in for Protégé, it helps users align ontologies through a graphical interface.
[0078] AgreementMaker: An automatic ontology alignment tool that supports multiple algorithms to achieve alignment between ontologies.
[0079] OntoMatcher: Another ontology alignment tool that supports the alignment of ontologies in different fields.
[0080] Once the ontology alignment, disambiguation and fusion are completed, the ontology reasoning engine can be used to reason and integrate knowledge. This process will extract more useful information from the aligned ontology through reasoning rules.
[0081] The ontology reasoning engine can automatically generate new knowledge based on the classes, attributes and relationships in the ontology. Through reasoning, some implicit relationships or attributes can be obtained to further enrich the content of the ontology. Through the collaboration of ontology alignment and reasoning engine, it can ensure that data from different sources can be seamlessly integrated, shared and interoperable. This is crucial for multi-domain collaboration and the construction of large-scale knowledge bases.
[0082] The purpose of the ontology engine 30 is to ensure that ontologies in different fields can communicate with each other and improve the efficiency and quality of knowledge integration.
[0083] In one embodiment, the above-mentioned knowledge graph construction unit 40 is used to perform entity recognition and relationship extraction in the battery field on the processing results by fine-tuning BERT and its variant models, and map, disambiguate and fuse the entities in combination with the battery field ontology library to obtain a knowledge graph in the battery field.
[0084] In one embodiment, the above-mentioned knowledge graph construction unit 40 is used to achieve entity matching through rules, and automatic mapping uses similarity calculation or machine learning for matching; determines entity disambiguation by analyzing context, attributes and semantic similarity; and performs entity fusion by merging similar or identical entities.
[0085] Specifically, through natural language processing technology (such as entity recognition, relationship extraction, coreference resolution, etc.), key entities and their relationships related to battery technology are automatically extracted from various text resources (such as academic papers, patents, industry reports, encyclopedias, forums, and expert blogs, etc.) to form preliminary structured data. Use the knowledge ontology in the battery field for semantic mapping and fusion, and integrate data from different sources into a consistent knowledge graph. This process is based on the classification and semantic rules defined by the ontology to ensure the structuring of knowledge and unified semantic standards, eliminate redundancy and fill in information gaps, and ensure the integrity and accuracy of the knowledge graph.
[0086] In one embodiment, the core task of the knowledge graph construction unit 40 is to extract structured information from unstructured or semi-structured text data and construct a domain knowledge base. This process includes three key steps: entity recognition, relationship extraction, and knowledge fusion. The following is a specific implementation method for each step:
[0087] The goal of entity recognition is to identify entities with specific meanings from text. To achieve this goal, the following measures can be taken:
[0088] BERT-based sequence labeling: Use BERT or its variant models (such as RoBERTa, DistilBERT, etc.) for sequence labeling, and train the model to identify entity boundaries and types in the text. The model outputs the probability distribution of the label sequence, and uses conditional random fields (CRF) or other decoding strategies to optimize the coherence of the sequence labeling.
[0089] Combine domain dictionaries and rules: Build a dictionary in the field of battery materials, including information such as material names, properties, and uses, to assist the model in improving recognition accuracy. This helps the model better understand domain-specific terms.
[0090] Relation extraction aims to identify the relationship between entities and provide a basis for the construction of knowledge graphs. Specific implementation methods include:
[0091] Use open source tools such as REBEL: Design specific relationship models and use these tools to automatically extract relationships between entities from text.
[0092] The knowledge fusion module is responsible for solving the problems of entity mapping, disambiguation and fusion, ensuring that entities in different data sources can be correctly associated and merged, thereby improving the consistency and completeness of the knowledge base.
[0093] Entity mapping involves matching the identified entities with entities in a known entity library or ontology. Common methods include:
[0094] Rule specification: including naming mapping rules, attribute mapping rules, context mapping rules, etc. Specifically, based on the naming rules of entities, entities with the same or similar names are mapped. For example, "Li-ionBattery" and "lithium battery" are mapped to the same entity. "ElectricVehicleBattery" and "EV battery" are mapped to the same entity.
[0095] If the names of two entities are exactly the same or match after lexical processing (such as ignoring case, removing spaces, etc.), they are mapped to the same entity. Alternatively, the similarity of the strings (such as edit distance) is calculated to determine whether the two entities are mapped to the same entity.
[0096] Mapping is done based on the attributes of the entities. For example, if two entities have very similar attribute values, they can also be considered the same entity. If the attribute values of two entities differ within a certain threshold, they are considered to belong to the same entity. Or, based on the same or similar attribute types, such as the relationship between "battery type" and "battery model".
[0097] Context-based mapping rules: By analyzing the context in which the entities appear, determine whether they belong to the same entity.
[0098] If two entities appear in the same context and describe the same type of thing, they can be mapped to the same entity. Alternatively, using semantic similarity evaluation, if two entities have similar meanings in the context, they can be mapped to the same entity.
[0099] Mapping implementation: Use automated mapping methods, combining rules and machine learning models.
[0100] Mapping quality evaluation: The mapping effect is evaluated through indicators such as accuracy, recall, F1 value, and mapping coverage.
[0101] Maintenance and update: Use tools such as OWLAPI for ontology management and version control.
[0102] Among them, accuracy (Precision) is the proportion of correctly mapped entities to all mapping attempts.
[0103] Recall: The proportion of all correct entities that are successfully mapped.
[0104] F1-score: The weighted average of precision and recall, used to comprehensively evaluate the effect of entity mapping.
[0105] Mapping coverage: The ratio of successfully mapped entities to the number of entities to be mapped.
[0106] Use OWLAPI for ontology management, support version control and update of ontology, and define the rules and process for ontology update.
[0107] Monitor the usage of mappings, collect user feedback, and identify possible mapping issues. Regularly evaluate the effectiveness of mappings and update mapping rules and algorithms to adapt to changes in domain knowledge.
[0108] Based on the evaluation results, the mapping algorithm or rules are optimized to gradually improve the mapping accuracy. For example, if the mapping accuracy of some entities is found to be low, specific rules can be strengthened or more context information can be used to improve the mapping.
[0109] Perform version control on the ontology and mapping, record each modification and its reason for easy tracing and restoration, and ensure that all modifications are reflected in related applications and systems.
[0110] Record mapping rules, processes, decisions, and results to form systematic documentation to facilitate subsequent understanding and maintenance.
[0111] Regularly organize domain experts to review the mapping results, especially for complex entity mapping problems that are difficult for machines to accurately judge. After the expert review, update the mapping rules and optimize them based on the expert feedback.
[0112] Entity disambiguation aims to solve entity homonymy or polysemy problems and ensure that each entity is mapped to the correct concept. Disambiguation methods include:
[0113] Context-based disambiguation rules: Consider the contextual information in which the entity appears. Including:
[0114] Contextual disambiguation rules: Determine the correct meaning of an entity in a specific scenario based on the contextual information of the entity in a sentence or paragraph. For example, "Apple" may refer to "Apple Company" or "Apple Fruit" in different contexts, and the choice needs to be made based on the context.
[0115] Temporal disambiguation rules: Consider temporal information to distinguish entities with the same name. For example, "battery" may refer to different technologies or products at different points in time.
[0116] Spatial disambiguation rules: Disambiguation is done by taking geographic location information into account. Judgment must be made based on specific geographic location or industry background.
[0117] Attribute-based disambiguation rules: use the attribute values or features of entities. They include:
[0118] Attribute value disambiguation rules: When multiple entities have the same name but different attributes, they can be distinguished by the specific value of the attribute. For example, "Li-ion Battery" and "lithium battery" are distinguished under different voltages or capacities.
[0119] Category disambiguation rules: Consider the category information of the entity and select entities that match a specific category. For example, "battery" has different disambiguation results in different types of devices (such as mobile phones and electric cars).
[0120] Disambiguation rules based on semantic similarity: Distinguish entities by calculating semantic similarity. Including: Word vector similarity disambiguation rules: Determine the most appropriate matching entity by calculating the word vector similarity between the entity and the candidate entity.
[0121] Knowledge graph disambiguation rules: Disambiguation is performed based on the existing relationships and entity features in the knowledge graph, and the entity that best matches the upstream and downstream relationships is selected.
[0122] Disambiguation implementation: A combination of automation and manual review is used. Specifically, automated disambiguation: Entity disambiguation is automatically performed using rules and models (such as machine learning models or deep learning models). Combining the contextual information of the text, the attributes of the entity, the category of the entity, and other information, rule-based matching methods or deep learning-based models are used for disambiguation.
[0123] Manual review: Manual intervention is performed for complex situations that are difficult to automatically resolve. Experts can make resolution judgments based on their knowledge and experience.
[0124] Disambiguation verification: Manual verification is performed on the disambiguation results to ensure the accuracy of the disambiguation and avoid errors in subsequent analysis and reasoning due to incorrect disambiguation.
[0125] Disambiguation quality evaluation: It is also evaluated by indicators such as precision, recall, F1 value and disambiguation coverage. Specifically, precision is the ratio of the number of correctly disambiguated entities to all disambiguation attempts.
[0126] Calculation formula: Among them, TP represents the number of correctly disambiguated entities, and FP represents the number of incorrectly disambiguated entities.
[0127] Recall: The proportion of entities that are successfully disambiguated among all entities that need to be disambiguated. Calculation formula: Where FN represents the number of entities that are missed for disambiguation.
[0128] F1-score: The weighted average of precision and recall, which comprehensively evaluates the disambiguation effect. Calculation formula:
[0129] Disambiguation coverage: The ratio of successfully disambiguated entities to the entities to be disambiguated.
[0130] As domain knowledge develops, disambiguation rules are updated regularly to ensure that new types of entities can be correctly disambiguated. By using a feedback system, the disambiguation model is continuously optimized to gradually improve the disambiguation accuracy. Feedback from domain experts is collected to correct errors in the automatic disambiguation system.
[0131] Entity fusion is to merge duplicate or similar entities, reduce redundancy and keep the knowledge base clean. Fusion methods include:
[0132] Attribute-based fusion rules: When entities have similar or overlapping attributes, they are merged. Including: Attribute value merging rules: If some attribute values of two entities are very similar or equal, they can be directly merged.
[0133] Attribute conflict handling rules: When the attributes of two entities conflict, a conflict handling mechanism needs to be defined. Usually, a more credible data source is selected, or domain knowledge is combined to decide which attribute to keep.
[0134] Context-based fusion rules: When entities appear in the same context, they are fused. Including: Context consistency fusion rules: When two entities appear in the same context or application scenario, they can be fused into one entity.
[0135] Historical record fusion rule: Based on historical record information, the most frequently occurring entity is selected as the merging target.
[0136] Semantic-based fusion rules: use semantic similarity evaluation to decide whether to merge. Including: Similarity threshold fusion rule: when the semantic similarity of two entities exceeds a certain threshold, they are considered to represent the same object.
[0137] Relationship matching rule: Determine whether two entities should be merged based on the relationship network between the entities.
[0138] Fusion implementation: using a combination of automated fusion and manual intervention. Including:
[0139] Automatic fusion: Use predefined fusion rules to automatically perform entity fusion by combining entity attributes, contextual information, and semantic information.
[0140] Machine learning methods or deep learning methods can be used for entity fusion, using models to predict which entities can be merged.
[0141] Manual intervention and review: Manual intervention is performed on complex fusion situations, and experts judge whether fusion is possible based on domain knowledge, especially when polysemous words and ambiguous entities are involved.
[0142] Fusion verification: Manual or automatic verification of the fusion results to ensure the accuracy of entity fusion and avoid erroneous fusion.
[0143] Fusion quality assessment: It is evaluated through indicators such as accuracy, recall, F1 value, and fusion coverage rate.
[0144] As the domain knowledge changes, regularly update the entity fusion rules to ensure that new entities can be correctly fused. Continuously adjust and optimize the entity fusion strategy through domain expert or user feedback to solve complex fusion situations.
[0145] Combine the identified entities and their relationships, and fuse data from different sources through ontology reasoning rules to ensure the unity and consistency of knowledge. Based on the feedback data and expert review results, formulate new update frequencies and standards, set the threshold T. When the mean satisfaction MF < T, increase the update frequency, record and evaluate the effect of each update, and continuously optimize the update strategy. Through the above steps and examples, a cyclic self-enhancing ontology system can be realized, improving user satisfaction and ensuring the dynamic update and accuracy of the knowledge base.
[0146] Through the above steps, knowledge can be effectively extracted and integrated from the text data processed by the data processing unit 20, and a dynamic, accurate, and consistent domain knowledge base can be constructed. In addition, regularly updating and optimizing the mapping, disambiguation, and fusion rules, combined with user feedback and expert review, can further improve the system performance and user experience. This process not only improves the quality of the knowledge base but also provides a solid foundation for subsequent knowledge application and services.
[0147] In one embodiment, please refer to Figure 1 , the above data storage unit 50 is used to store the knowledge graph using Neo4j and store vector data using Milvus to form a knowledge graph database and a vector database.
[0148] In addition, the above data storage unit 50 is also used to store the corresponding documents using ElasticSearch, Mysql, etc., such as the documents processed by the data processing unit 20, as well as store the meta-information of the data, etc.
[0149] In this embodiment, regarding the formation of the vector database, first convert various unstructured texts (such as academic papers, industry reports, etc.) into vector representations, and select a pre-trained embedding model suitable for the battery field for domain-specific fine-tuning. Generate semantic vectors of the text through the embedding model to improve the performance of specific tasks.
[0150] Store these generated semantic vectors in an efficient vector database and optimize the retrieval efficiency through an indexing mechanism. The vector database should support high-concurrency similarity retrieval to ensure fast response and return relevant texts.
[0151] Use similarity search algorithms (such as cosine similarity, Euclidean distance, etc.) to match user queries with text snippets. To enhance the search effect, combine domain keyword matching with similarity fusion strategy to ensure that the returned content is not only semantically relevant, but also accurately matches professional terms and concepts. In addition, multiple rounds of iterative search can be used to optimize queries and improve the relevance and contextual consistency of search results.
[0152] In terms of retrieval enhancement based on knowledge graph, an index of the battery material knowledge graph is constructed based on the knowledge ontology of the battery field, so that knowledge retrieval can be performed based on entities and relationships, thereby improving the accuracy of retrieval.
[0153] The system can extract relevant subgraph content based on the query through the entity relationship structure in the knowledge graph. For example, when a user queries the performance parameters of battery materials, the system will return relevant material entities and their attribute values, production processes and other information, and provide a context-rich structured subgraph. Path query technology can also be used to find indirect or adjacent entities related to the query to ensure that multi-level associated knowledge is provided.
[0154] In order to improve the retrieval effect of the system, semantic expansion and relationship fusion are introduced. By expanding the semantic rules defined in the ontology (such as synonyms, hyponyms, and hyponyms), the system can integrate the user's questions with related concepts and entities during the query, thereby expanding the coverage of the knowledge graph and ensuring that users obtain more comprehensive information.
[0155] In one embodiment, see Figure 1 , the query engine unit 60 includes a GraphRAG module, which is used to perform relational reasoning from the knowledge graph database according to the query request to obtain a retrieval result;
[0156] The query engine unit 60 further includes a VectorRAG module, which is used to perform similarity query in vector data according to a query request to obtain a retrieval result.
[0157] In one embodiment, see Figure 1 The query engine unit 60 further includes a HybridRAG module, which is used to input the query request into the vector database through the VectorRAG module for similarity retrieval to obtain documents or fragments similar to the query request; transmit the documents or fragments similar to the query request to the GraphRAG module for reasoning and knowledge fusion to obtain fusion results; and fuse the documents or fragments similar to the query request and the fusion results to obtain retrieval results.
[0158] In one embodiment, see Figure 1The query engine unit 60 further includes a FusionRAG module, which is used to directly search from the knowledge graph data and the vector database, fuse the retrieved results, and generate a final answer based on the fused results using a generative model.
[0159] In this embodiment, the GraphRAG module is good at handling tasks that require complex logic and deep information mining, but this also leads to relatively low processing speed, especially when faced with large amounts of data or complex relationship networks. It is able to provide responses containing a lot of context and deep information, which is very useful for questions that require detailed explanations or background knowledge.
[0160] Graph neural networks are used for subgraph retrieval, and the representation in the graph is learned by updating the node representation H(L+1)=σ(AH(L)W(L)), where A is the adjacency matrix, W(l) is the weight matrix, and σ is the activation function.
[0161] The VectorRAG module is based on a vector database (such as FAISS / Milvus). VectorRAG can quickly perform similarity queries and is particularly suitable for processing real-time queries in large-scale data sets.
[0162] When processing simple queries, it can significantly reduce computing time and resource consumption and improve the system's response speed.
[0163] By calculating the angular difference between the vectors Evaluate the semantic similarity of two vectors.
[0164] Suitable for scenarios that require fast response and efficient processing.
[0165] The HybridRAG module combines the speed and efficiency of the VectorRAG module with the deep reasoning capabilities of GraphRAG to form a system that enables both fast retrieval and deep analysis.
[0166] Use a variety of methods (such as splicing, summary generation, or advanced fusion algorithms) to integrate information from different sources into a coherent response.
[0167] It is particularly suitable for fields that require both quick response and high-quality, high-precision answers, such as professional consulting, scientific research, etc.
[0168] The FusionRAG module directly retrieves information from knowledge graphs and vector databases and generates the final answer through a large language model (LLM), simplifying the traditional RAG process.
[0169] For problems that do not require complex reasoning, FusionRAG provides a faster and more direct solution.
[0170] Ability to effectively handle a wide range of situations, from simple queries to problems of moderate complexity.
[0171] Applicable to most information retrieval tasks, especially those that can be solved by direct data integration and simple semantic understanding, such as educational tutoring, product recommendations, etc.
[0172] Each module has its own unique advantages and most suitable application scenarios. When choosing a suitable model, it is necessary to consider the specific requirements of the task, including factors such as the requirements for response speed, the requirements for information depth, and acceptable computational costs. For example, for applications in the battery field, if the goal is to provide fast market information queries, VectorRAG may be the best choice; while for problems that require in-depth analysis of battery performance and material science, HybridRAG or GraphRAG are more suitable.
[0173] In this embodiment, the GraphRAG and VectorRAG methods are combined to provide users with intelligent knowledge retrieval and answer generation. The GraphRAG method retrieves structured data from the knowledge graph and accurately extracts subgraph content related to entity relationships; the VectorRAG method performs semantic retrieval from unstructured data and flexibly responds to fuzzy and context-dependent questions.
[0174] Through the HybridRAG method, the system combines the advantages of GraphRAG and VectorRAG, integrates the information of structured knowledge graph and unstructured text, and generates accurate and coherent answers. This method can effectively improve the accuracy, contextual consistency and semantic accuracy of answers, ensuring that users can obtain comprehensive and in-depth knowledge of battery materials.
[0175] In one embodiment, the design and implementation of the above application interface is based on underlying support and general technology, aiming to provide a series of basic support functions for the application layer, including but not limited to:
[0176] RESTful API interface design: provides standardized interface design, supports efficient data exchange between clients and servers, and constitutes the core communication method of the application.
[0177] WebSocket enables real-time query: It is suitable for scenarios that require high real-time performance, such as real-time data streaming and dynamic visualization updates, and effectively solves the delay problem of traditional HTTP polling.
[0178] Permission management and access control: Ensure system security and maintain data security and consistency of business rules through sophisticated permission settings and access control measures.
[0179] Visual component development: Develop a friendly and intuitive front-end interactive interface to improve user experience, serving as the main window for function display and business interaction.
[0180] In one embodiment, the above-mentioned application interface includes: a knowledge retrieval interface, an intelligent question and answer interface, and a visual analysis interface.
[0181] Specifically, the knowledge retrieval interface provides efficient knowledge base retrieval functions, supporting users to quickly find knowledge content related to the query through keyword retrieval, semantic retrieval, and relationship retrieval. It has the following features:
[0182] Multi-level retrieval: supports multiple retrieval methods to meet users' diverse query needs.
[0183] Fuzzy search: enables queries in cases of incomplete matches, improving search flexibility and accuracy.
[0184] Contextual retrieval: Combines the user's query history and context information to provide personalized retrieval results.
[0185] Smart Sorting: Automatically sort search results based on factors such as relevance, timeliness and quality.
[0186] Advanced filtering: supports filtering based on time, source, type and other conditions to help users find information more accurately.
[0187] This interface can help R&D teams make more scientific decisions in material selection, technology comparison and process verification.
[0188] Intelligent question-answering interface: An intelligent question-answering system built on a knowledge base that supports users to ask questions in natural language and provides accurate answers. It has the following features:
[0189] Natural language processing: Able to understand and parse users' natural language questions and accurately capture user intent.
[0190] Contextual understanding: supports multi-round conversations and provides a coherent question-and-answer experience.
[0191] Answer generation: Combine the RAG model and knowledge base content to automatically generate clear and accurate answers.
[0192] Question and answer recommendation: Automatically recommend answers to frequently asked questions to improve query efficiency.
[0193] Knowledge Supplement: When a direct answer is not possible, provide additional suggestions for relevant information.
[0194] This interface can help R&D personnel quickly answer performance issues, technical questions, process parameters and other issues, and speed up R&D progress.
[0195] Visual analysis interface: Provides a visual display tool based on knowledge graph, supporting users to view and explore complex knowledge systems through a graphical interface. It has the following features:
[0196] Relationship network display: Display entities and relationships in the knowledge graph in a graphical way, supporting multi-level display.
[0197] Multi-dimensional analysis: allows users to filter and focus information based on specific dimensions.
[0198] Dynamic interaction: supports operations such as clicking, dragging, and zooming, providing a flexible browsing experience.
[0199] Data filtering and aggregation: Filter and aggregate data according to specific conditions to help users focus on key information.
[0200] Time series display: supports the timeline function to show the changing trend of the knowledge graph over time.
[0201] This interface can be used to assist users in understanding material properties and their applications. It can also help users track the development history of technology or materials and support R&D decisions.
[0202] In addition, the interface is integrated with the graph database to achieve efficient large-scale graph data management. Advanced front-end visualization tools are used to improve display effects and user experience. Graph algorithms such as community discovery, path analysis, and centrality calculation are supported to help users deeply explore the value of knowledge.
[0203] The method of this embodiment can ensure the standardization and interoperability of data by defining a clear domain ontology, thereby supporting cross-platform data sharing and integration; it emphasizes the contextualized, linked, explorable and reliable characteristics of data, which is particularly suitable for data management and analysis in professional fields. For example, in the study of battery materials, the value and availability of data can be greatly improved by associating data with specific experimental conditions, test methods and physical and chemical properties. The ability to extract and fuse information from multiple data sources, including scientific literature, experimental reports, patent documents, etc. Through automated data mapping and integration technology, scattered data can be concentrated into a unified knowledge graph, which not only enhances the comprehensiveness and timeliness of the data, but also provides users with richer and deeper channels for knowledge acquisition. Using RAG (Retrieval-Augmented Generation) technology, combined with pre-trained language models and knowledge bases in specific fields, accurate understanding and response generation of user queries can be achieved. This method not only improves the relevance and accuracy of information retrieval, but also generates personalized answers based on the specific needs of users. The HybridRAG method performs well in dealing with complex technical problems by combining the advantages of retrieval and generation.
[0204] The system has the ability to self-learn and optimize. When users are dissatisfied with the search results, they can collect this information through the feedback mechanism and use GEN AI (Generative AI) technology to optimize and expand the existing domain ontology. This continuous improvement process helps to maintain the advancement and practicality of the system and ensure that it serves the needs of scientific research and industrial development in the long term.
[0205] Researchers can use the system to quickly obtain the latest research results and technological trends, thereby accelerating the discovery of new materials and optimizing the performance of existing materials. This helps shorten the R&D cycle, reduce costs, and improve competitiveness. Enterprises and research institutions can use the system's powerful patent analysis capabilities to identify potential patent opportunities, avoid infringement risks, and develop reasonable intellectual property strategies. This is of great significance for protecting technological innovation achievements and supporting sustainable development. By promoting the open sharing of data, the system can strengthen exchanges and cooperation among various entities in the industry, form synergies, and jointly promote the advancement and development of battery technology. This plays a positive role in building a healthy industrial ecology and achieving effective allocation and utilization of resources.
[0206] In short, through the application of the above-mentioned technical solutions, the scientific research efficiency and innovation capabilities in the field of battery materials can be significantly improved, while providing strong support for the company's intellectual property management and the coordinated development of the industry.
[0207] The above-mentioned retrieval generation system based on RAG and ontology-driven battery material knowledge graph realizes efficient management and intelligent retrieval of knowledge in the field of battery materials by integrating functions such as data collection, processing, ontology construction, knowledge graph generation, storage and query. It uses RAG technology and ontology-driven methods to improve data integration and retrieval accuracy, and supports dynamic expansion and real-time update of multimodal data. Users can input query requests through application interface 70, and the query engine will call the knowledge graph and vector data in the storage unit for efficient retrieval and return accurate results. It comprehensively solves the problems faced by existing retrieval methods, such as data dispersion, low retrieval efficiency, delayed knowledge update and insufficient semantic understanding.
[0208] The above-mentioned retrieval and generation system based on RAG and ontology-driven battery material knowledge graph can be implemented in the form of a computer program, which can be run on a computer device.
[0209] In one embodiment, a retrieval method based on RAG and the knowledge graph of the ontology-driven battery material is also provided, including:
[0210] Get the query request;
[0211] The corresponding query engine calls the database to perform a search according to the query request, and returns the search result;
[0212] Among them, the database includes a knowledge graph database and a vector database. The knowledge graph database is obtained by collecting relevant data of various battery materials, preprocessing and NLP processing the collected data to obtain processing results, and performing entity recognition and relationship extraction in the battery field by fine-tuning BERT and its variant models, and combining the battery field ontology library constructed by the ontology engine 30 to achieve the fusion and alignment of entities and relationships to obtain a knowledge graph in the battery field, and store the formed database; the vector database is obtained by collecting relevant data of various battery materials, preprocessing and vectorizing the collected data, and storing the formed database.
[0213] It should be noted that technical personnel in the relevant field can clearly understand that the specific implementation process of the above-mentioned retrieval method based on RAG and the knowledge graph of the main driving battery material can refer to the corresponding description in the aforementioned system embodiment, and for the convenience and conciseness of the description, it will not be repeated here.
[0214] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0215] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0216] The steps in the method of the embodiment of the present invention can be adjusted in order, combined and deleted according to actual needs. The units in the system of the embodiment of the present invention can be combined, divided and deleted according to actual needs. In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0217] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, terminal, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.
[0218] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A retrieval and generation system based on RAG and ontology-driven battery material knowledge graph, characterized by: include: Data collection unit, data processing unit, ontology engine, knowledge graph construction unit, data storage unit, query engine unit and application interface; The data acquisition unit is used to provide a data acquisition interface and collect relevant data of various battery materials; The data processing unit is used to perform preprocessing and NLP processing on the collected data to obtain processing results and vector data; The ontology engine is used to construct an ontology library in the battery field; The knowledge graph construction unit is used to perform entity recognition and relationship extraction in the battery field by fine-tuning BERT and its variant models, and to achieve relationship fusion and alignment in combination with the battery field ontology library to obtain a knowledge graph in the battery field; The data storage unit is used to store the knowledge graph and the vector data to form a knowledge graph database and a vector database; The application interface is used to input a query request; The query engine unit is used to call the database in the data storage unit to perform retrieval according to the query request, and return the retrieval result to the application interface.
2. The retrieval generation system based on RAG and ontology-driven battery material knowledge graph according to claim 1 is characterized in that: The relevant data of various types of battery materials include academic papers, patent documents, experimental data, network data and non-public information.
3. The retrieval generation system based on RAG and ontology-driven battery material knowledge graph according to claim 1 is characterized in that: The data processing unit is used to process structured and unstructured data using regex, OCR model, and PDF recognition series on the collected data, and to perform word segmentation and part-of-speech tagging using the spaCy tool to obtain processing results and vector data.
4. The retrieval generation system based on RAG and ontology-driven battery material knowledge graph according to claim 1 is characterized in that: The ontology engine is used to define concepts, attributes and relationships in the battery field, construct a preliminary ontology, and design an ontology reasoning engine for the ontology. The ontology reasoning engine is used to map, disambiguate and merge the ontology to build an ontology library in the battery field.
5. The retrieval generation system based on RAG and ontology-driven battery material knowledge graph according to claim 1 is characterized in that: The knowledge graph construction unit is used to perform entity recognition and relationship extraction in the battery field on the processing results by fine-tuning BERT and its variant models, and map, disambiguate and fuse the entities in combination with the battery field ontology library to obtain a knowledge graph in the battery field.
6. The retrieval generation system based on RAG and ontology-driven battery material knowledge graph according to claim 5 is characterized in that: The knowledge graph construction unit is used to achieve entity matching through rules, and automatic mapping uses similarity calculation or machine learning for matching; determines entity mapping by analyzing context, attributes and semantic similarity; and performs entity fusion by merging similar or identical entities.
7. The retrieval and generation system based on RAG and ontology-driven battery material knowledge graph according to claim 1 is characterized in that: The query engine unit includes a GraphRAG module, and the GraphRAG module is used to perform relational reasoning from the knowledge graph database according to a query request to obtain a search result; The query engine unit further includes a VectorRAG module, and the VectorRAG module is used to perform a similarity query in the vector data according to the query request to obtain a retrieval result.
8. The retrieval generation system based on RAG and ontology-driven battery material knowledge graph according to claim 7 is characterized in that: The query engine unit further includes a HybridRAG module, which is used to input the query request into the vector database through the VectorRAG module for similarity retrieval to obtain documents or fragments similar to the query request; transmit the documents or fragments similar to the query request to the GraphRAG module for reasoning and knowledge fusion to obtain a fusion result; The documents or fragments similar to the query request and the fusion result are fused to obtain a retrieval result.
9. The retrieval generation system based on RAG and ontology-driven battery material knowledge graph according to claim 8 is characterized in that: The query engine unit also includes a FusionRAG module, which is used to directly search from the knowledge graph data and the vector database, fuse the retrieved results, and generate a final answer based on the fused results using a generation model.
10. A retrieval method based on RAG and the knowledge graph of ontology-driven battery materials, characterized in that: include: Get the query request; The corresponding query engine calls the database to perform a search according to the query request, and returns the search result; The database includes a knowledge graph database and a vector database. The knowledge graph database is obtained by collecting relevant data of various battery materials, preprocessing and NLP processing the collected data to obtain processing results, and fine-tuning BERT and its variant models to perform entity recognition and relationship extraction in the battery field, and combining the battery field ontology library built by the ontology engine to achieve relationship fusion and alignment to obtain a knowledge graph in the battery field and store the formed database; The vector database is formed by collecting relevant data of various battery materials, preprocessing and NLP processing the collected data, and storing the vector data.
Citation Information
Cited By
Electrochemical energy storage battery overcharge safety early warning method based on large model
CN120974311A
A Safety Early Warning Method for Overcharge of Electrochemical Energy Storage Batteries Based on a Large Model
CN120974311B