A knowledge graph and digital object hybrid method for scientific discovery

By performing structured parsing and concept disambiguation on the literature dataset, an unambiguous concept dataset is generated, which solves the mismatch problem caused by concept ambiguity in knowledge graphs and achieves efficient and accurate knowledge retrieval.

CN121524372BActive Publication Date: 2026-04-14PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-19
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies for constructing knowledge graphs suffer from problems such as mismatches and inaccurate search results due to conceptual ambiguity, which affect the search efficiency and effectiveness of researchers.

Method used

By performing structured parsing on the literature dataset, extracting concept names and explanatory data, identifying and merging concept names with the same conceptual connotations, generating an unambiguous concept dataset, and establishing knowledge graph data, we can ensure a clear mapping between concept names and literature data.

Benefits of technology

It improves the accuracy and efficiency of knowledge retrieval, ensuring that searchers can quickly locate target concepts in the knowledge graph and trace back to relevant literature fragments, thereby enhancing the search efficiency and reliability of researchers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121524372B_ABST
    Figure CN121524372B_ABST
Patent Text Reader

Abstract

The application discloses a knowledge graph and digital object hybrid method for scientific discovery, and belongs to the field of data retrieval, and comprises the following steps: analyzing a literature data set to obtain an initial concept data set of the literature data set; comparing the concept explanation data of each concept data in the initial concept data set, fusing the concept data corresponding to the concept names representing the same concept connotation, and obtaining a concept data set after ambiguity elimination; and generating knowledge graph data of the literature data set according to the concept data set after ambiguity elimination. The application realizes the layer-by-layer conversion from original literature to structured knowledge by the core idea of joint storage-concept disambiguation, and effectively improves the problems of semantic retrieval mismatching and the lack of accurate literature corresponding to the knowledge graph, thereby improving the accuracy and efficiency of the retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data retrieval, and specifically relates to a method, apparatus, device and storage medium for the integration of knowledge graphs and digital objects for scientific discovery. Background Technology

[0002] With the ever-growing scale of scientific research data, researchers need to quickly extract knowledge related to their research topics from a large number of papers, experimental reports, and patent documents. Different documents often contain information such as concepts, experimental conditions, and performance results, and these information have complex semantic relationships. To support cross-document knowledge fusion and scientific discovery, there is an urgent need for a data retrieval method that can structurally express conceptual relationships and maintain a close connection with the original documents. This would help researchers accurately locate target concepts in knowledge graphs and trace them back to the corresponding literature data.

[0003] Existing technologies have combined document identification systems with full-text search engines to build knowledge graphs, enabling document-level indexing and querying. These technologies can support, to some extent, the querying of conceptual relationships and the fusion of knowledge across documents, and improve the relevance of searches through semantic retrieval or vectorized matching.

[0004] However, semantic retrieval methods are prone to mismatches when dealing with conceptual ambiguity, resulting in chaotic and incomplete search results. At the same time, concept nodes in knowledge graphs often lack precise correspondence with specific document paragraphs or data fragments, all of which affect the retrieval efficiency of searchers. Summary of the Invention

[0005] This application aims to provide a method, apparatus, device, and storage medium for combining knowledge graphs and digital objects in scientific discovery, at least addressing the problems of insufficient accuracy and efficiency in knowledge retrieval.

[0006] In a first aspect, embodiments of this application disclose a method for blending knowledge graphs and digital objects for scientific discovery, including:

[0007] The literature dataset is parsed to obtain an initial concept dataset; the literature dataset contains at least one data document; the concept dataset contains multiple concept data; each concept data is used to characterize the corresponding concept name and concept explanation data in the literature dataset, as well as the relationship between the concept names;

[0008] The conceptual explanation data for each concept data in the initial concept dataset is compared to the conceptual data corresponding to concept names that are characterized as having the same conceptual connotation, so as to merge the conceptual data after disambiguation;

[0009] Based on the unambiguous concept dataset, a knowledge graph data of the literature dataset is generated; the knowledge graph data is used to characterize the relationships between various concept names appearing in the literature dataset, as well as the data literature corresponding to each concept name.

[0010] Secondly, embodiments of this application also disclose a knowledge graph and digital object hybrid device for scientific discovery, comprising:

[0011] A parsing storage module is used to parse the literature dataset to obtain an initial concept dataset of the literature dataset; the literature dataset contains at least one data document; the concept dataset contains multiple concept data; each concept data is used to characterize the corresponding concept name and concept explanation data in the literature dataset, as well as the relationship between the concept names;

[0012] The ambiguity resolution module is used to compare the concept explanation data of each concept data in the initial concept dataset, so as to merge the concept data corresponding to the concept names that represent the same concept connotation, so as to obtain the concept dataset after ambiguity resolution.

[0013] The knowledge graph building module is used to generate knowledge graph data of the literature dataset based on the unambiguous concept dataset; the knowledge graph data is used to represent the relationship between various concept names appearing in the literature dataset, and the data literature corresponding to each concept name.

[0014] Thirdly, embodiments of this application also disclose an electronic device, including a processor and a memory, wherein the memory stores a program or instructions that can run on the processor, and the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0015] Fourthly, embodiments of this application also disclose a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the method described in the first aspect.

[0016] In summary, in this embodiment, by performing structured parsing on the literature dataset, the concept names, concept explanations, and their relationships in the literature are extracted as an initial concept dataset. This allows knowledge units scattered across different documents to be stored and managed in a unified structured form, ensuring data organization and processability. While ensuring traceability in data retrieval, this provides a foundation for subsequent concept disambiguation and knowledge graph generation. Furthermore, by comparing the concept explanations, different concept names representing the same concept are identified and merged, reducing duplicate or incorrect matching problems caused by concept ambiguity. This lowers the risk of mismatches common in semantic retrieval, making the concept dataset more unified and accurate, providing semantically consistent input for knowledge graph construction. Finally, knowledge graph data is generated based on the disambiguated concept dataset, enabling each concept name to not only establish clear relationships with other concepts but also to establish a clear mapping with the corresponding literature data. This allows searchers to quickly locate target concepts in the graph and directly trace back to relevant literature fragments, thereby improving search efficiency and the reliability of results. Therefore, the method based on the embodiments of this application, through the core idea of ​​joint storage-concept disambiguation, realizes the layer-by-layer transformation from original documents to structured knowledge, thereby effectively improving the problems of semantic retrieval mismatch and lack of precise document correspondence in knowledge graphs, thus improving the accuracy and efficiency of retrieval and providing researchers with higher quality knowledge retrieval support. Attached Figure Description

[0017] In the attached diagram:

[0018] Figure 1 This is a flowchart illustrating the steps of a method for combining knowledge graphs and digital objects in scientific discovery, as provided in an embodiment of this application.

[0019] Figure 2 This is a flowchart illustrating another method for combining knowledge graphs and digital objects in scientific discovery, as provided in an embodiment of this application.

[0020] Figure 3 This is a workflow of a data storage and retrieval system provided in an embodiment of this application;

[0021] Figure 4 This is a block diagram of a knowledge graph and digital object hybrid device for scientific discovery provided in an embodiment of this application;

[0022] Figure 5 This is a block diagram of an electronic device provided in one embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, in this application, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.

[0025] like Figure 1 The image shown is a method for combining knowledge graphs and digital objects for scientific discovery, provided in an embodiment of this application.

[0026] The method may include the following steps:

[0027] Step 101: Parse the literature dataset to obtain the initial conceptual dataset of the literature dataset.

[0028] The literature dataset contains at least one data document; the concept dataset contains multiple concept data; each concept data is used to characterize the corresponding concept name and concept explanation data in the literature dataset, as well as the relationship between the concept names.

[0029] In some embodiments of this application, it is necessary to perform structured parsing on the original document dataset to extract the knowledge units from the documents into a unified concept dataset, thereby providing a foundation for subsequent disambiguation and knowledge graph generation. This involves parsing the document dataset to obtain an initial concept dataset. The document dataset contains at least one data document, and the concept dataset contains multiple concept data. Each concept data represents the corresponding concept name and explanation data in the document dataset, as well as the relationships between the concept names. The concept dataset refers to a structured collection formed by parsing the document dataset, containing concept names, concept explanation data, and the relationships between concepts, used for unified storage and management of knowledge units scattered across different documents. This transforms unstructured document content into a structured concept dataset, enabling subsequent concept disambiguation and knowledge graph generation to be performed on a unified data basis, thereby improving data organization and processability.

[0030] In a specific example, the system needs to process a dataset of literature containing chemical reaction mechanisms. The system first parses these documents, extracting concept names such as "catalyst," "reaction rate," and "activation energy," and combines this with explanatory data from the documents to form an initial concept dataset. Simultaneously, it establishes relationships between these concepts, such as the correlation between "catalyst" and "reaction rate." In this way, the system obtains a structured set containing multiple concept datasets, each corresponding to a concept name, explanatory data, and its relationships in the literature, providing clear input for subsequent concept disambiguation and knowledge graph generation.

[0031] In the embodiments of this application, the parsing process of the concept dataset can utilize Natural Language Processing (NLP), Named Entity Recognition (NER), or a Transformer-based text parsing model. These models can perform word segmentation, entity recognition, and relation extraction on document text, thereby generating a structured concept dataset. The training dataset used to train the above models may include a scientific literature corpus, a manually annotated set of concept names and explanations, and annotated data of concept relations. It should be noted that the specific training methods and implementation details of the above models have been fully disclosed and applied in the prior art, and will not be repeated here.

[0032] Step 102: Compare the concept explanation data of each concept data in the initial concept dataset to merge the concept data corresponding to concept names that are characterized as having the same conceptual connotation, so as to obtain the concept dataset after eliminating ambiguity.

[0033] In some embodiments of this application, it is necessary to address the semantic ambiguity caused by the same concept being represented by different names in different documents, thereby ensuring semantic consistency in the subsequent knowledge graph generation. This can be achieved by comparing the concept explanation data of each concept in the initial concept dataset, fusing the concept data corresponding to concept names that represent the same conceptual connotation, thus obtaining a disambiguated concept dataset. Concept explanation data refers to textual or structured information used to explain the specific meaning, attributes, or context corresponding to a concept name, and it may differ in different documents. This reduces the problem of duplicate or incorrect matching caused by concept ambiguity, making the concept dataset more unified and accurate, and providing semantically consistent input for subsequent knowledge graph construction.

[0034] In a specific example, when processing a biomedical literature dataset, researchers found that different documents used "myocardial infarction" and "heart attack" to describe the same medical concept. After parsing the documents, the system incorporated these two concept names and their corresponding conceptual explanations into the initial concept dataset. The system then compared the conceptual explanations of the two documents, identified their consistent connotations, and merged the two sets of conceptual data to generate an unambiguous concept dataset. The final result was a unified concept data entry represented as "myocardial infarction / heart attack," with a clear mapping to relevant literature, providing semantically consistent input for subsequent knowledge graph generation.

[0035] In the above embodiments, the concept disambiguation process can employ a Semantic Matching Model (SMM), a Word Embedding Model (WEM), or a Transformer-based semantic similarity calculation model. These models can semantically vectorize and calculate similarity for concept explanation data, thereby identifying whether different concept names have the same conceptual connotation. The training dataset used to train the above models can also include a scientific literature corpus, a manually annotated set of concept names and explanation data, and labeled data of concept semantic similarity. It should be noted that the specific training methods and implementation details of the models have been fully disclosed and applied in the prior art, and will not be repeated here.

[0036] Step 103: Generate knowledge graph data of the literature dataset based on the deambiguous concept dataset.

[0037] Among them, knowledge graph data is used to represent the relationships between various concept names appearing in the literature dataset, as well as the data literature corresponding to each concept name.

[0038] In some embodiments of this application, after eliminating conceptual ambiguity, the unified concept dataset needs to be transformed into a structured knowledge graph to intuitively represent the relationships between concepts and establish a mapping with data documents. At this point, knowledge graph data of the document dataset is generated based on the eliminated concept dataset. This knowledge graph data represents the relationships between various concept names appearing in the document dataset, as well as the corresponding data documents for each concept name. Knowledge graph data refers to a collection of knowledge units stored in a structured form such as graphs / tables. It contains node element data representing concept names and edge element data representing the relationships between concepts. A mapping relationship is established between the node element data and the document data to ensure effective tracing in the subsequent retrieval process. This allows searchers to quickly locate target concepts in the knowledge graph and directly trace back to relevant document fragments, thereby improving retrieval efficiency and the reliability of results.

[0039] In a specific example, when researchers are processing a literature dataset in the field of materials science, the system has already performed concept disambiguation, resulting in a unified concept dataset containing concepts such as "crystal structure," "defect type," and "conductivity." The system then generates a knowledge graph based on this concept data. Nodes in the graph represent the names of these concepts, and the edges between nodes represent the relationships between "crystal structure" and "conductivity," as well as the influence of "defect type" on "conductivity." Simultaneously, each node is mapped to corresponding literature data; for example, the "crystal structure" node is associated with a literature fragment describing crystal structures. This results in a knowledge graph that intuitively displays conceptual relationships and supports literature tracing, providing structured support for subsequent retrieval and reasoning.

[0040] In summary, in this embodiment, by performing structured parsing on the literature dataset, the concept names, concept explanations, and their relationships in the literature are extracted as an initial concept dataset. This allows knowledge units scattered across different documents to be stored and managed in a unified structured form, ensuring data organization and processability. While ensuring traceability in data retrieval, this provides a foundation for subsequent concept disambiguation and knowledge graph generation. Furthermore, by comparing the concept explanations, different concept names representing the same concept are identified and merged, reducing duplicate or incorrect matching problems caused by concept ambiguity. This lowers the risk of mismatches common in semantic retrieval, making the concept dataset more unified and accurate, providing semantically consistent input for knowledge graph construction. Finally, knowledge graph data is generated based on the disambiguated concept dataset, enabling each concept name to not only establish clear relationships with other concepts but also to establish a clear mapping with the corresponding literature data. This allows searchers to quickly locate target concepts in the graph and directly trace back to relevant literature fragments, thereby improving search efficiency and the reliability of results. Therefore, the method based on the embodiments of this application, through the core idea of ​​joint storage-concept disambiguation, realizes the layer-by-layer transformation from original documents to structured knowledge, thereby effectively improving the problems of semantic retrieval mismatch and lack of precise document correspondence in knowledge graphs, thus improving the accuracy and efficiency of retrieval and providing researchers with higher quality knowledge retrieval support.

[0041] Figure 2 This is another method for combining knowledge graphs and digital objects for scientific discovery, provided in the embodiments of this application.

[0042] The method may include the following steps:

[0043] Step 201: Parse the literature dataset to obtain the initial conceptual dataset of the literature dataset.

[0044] The literature dataset contains at least one data document; the concept dataset contains multiple concept data; each concept data is used to characterize the corresponding concept name and concept explanation data in the literature dataset, as well as the relationship between the concept names.

[0045] The method shown in this step has been explained in step 101 and will not be repeated here.

[0046] Optionally, step 201 includes the following sub-steps:

[0047] Sub-step 2011 involves performing structured parsing on each data document in the document dataset to obtain multiple sets of corresponding concept names and concept explanation data, and generating relational attribute data corresponding to each concept name.

[0048] Among them, the relational attribute data is used to describe at least one concept name that has an association with the corresponding concept name.

[0049] In some embodiments of this application, it is necessary to perform structured parsing on each data document in the literature dataset to transform unstructured text information into processable conceptual units, thereby providing a foundation for subsequent concept organization and knowledge graph generation. This can be achieved by performing structured parsing on each data document in the literature dataset to obtain multiple sets of corresponding concept names and concept explanation data, and generating relational attribute data corresponding to each concept name. Relational attribute data refers to at least one concept name that describes a relationship with the corresponding concept name, reflecting the logical connections and semantic dependencies between concepts. This forms a structured set containing concept names, explanation data, and their relational attributes, enabling the system to organize and manage concept data in a unified manner in subsequent steps, improving data organization and processability.

[0050] In a specific example, when processing a literature dataset in the field of chemistry, the system performs structured analysis on a paper on "catalytic reaction mechanism," extracting the concept names "catalyst," "reaction rate," and "activation energy," and generating corresponding concept explanation data based on the literature content, such as "catalyst: a substance that lowers the activation energy in a chemical reaction." Simultaneously, the system generates relational attribute data, such as a correlation between "catalyst" and "reaction rate," and a causal relationship between "activation energy" and "reaction rate." This results in a structured set containing concept names, concept explanation data, and relational attribute data, providing clear input for the subsequent formation of a three-dimensional data set.

[0051] Optionally, sub-step 2011 includes the following sub-steps:

[0052] Sub-step 20111 involves performing hierarchical parsing of the data documents based on the data types contained in each data document to obtain the concept names in each data document, and generating corresponding concept explanation data based on the contextual semantics of each concept name in the corresponding data document.

[0053] In some embodiments of this application, hierarchical parsing is required for different types of data documents to accurately extract concept names from the documents and generate corresponding concept explanation data based on their contextual semantics. Specifically, hierarchical parsing can be performed on the data documents according to the data types contained in each document to obtain the concept names in each document. Based on the contextual semantics of each concept name in the corresponding document, corresponding concept explanation data is generated. Concept explanation data refers to metadata entities used to describe the semantic meaning or attributes corresponding to the concept names. It helps the system achieve semantic recognition and concept disambiguation in subsequent steps. This transforms unstructured document content into a structured set of concept names and explanation data, enabling the system to have higher accuracy and scalability in subsequent processing.

[0054] In a specific example, when processing literature data in the field of chemistry, the system performs layered parsing based on the data type of the literature (such as experimental reports, review articles, and patent specifications). In experimental reports, the system extracts the concept name "catalyst" and generates the explanatory data "a substance that lowers the activation energy in a chemical reaction" based on contextual semantics; in review articles, the system extracts the concept name "reaction rate" and generates the explanatory data "the change in reactant concentration per unit time"; in patent specifications, the system extracts the concept name "activation energy" and generates the explanatory data "the minimum energy required for a chemical reaction to occur." This yields multiple concept names and their corresponding explanatory data, providing a foundation for subsequent relational attribute generation and three-element data set construction.

[0055] Optionally, sub-step 20111 includes the following sub-steps:

[0056] Sub-step 201111: In the case that the data document contains first structured data recorded as text data, perform a first extraction process on the first structured data to obtain multiple first concept names, and generate concept explanation data for each first concept name based on the first analysis.

[0057] The first extraction process includes sentence parsing extraction and / or semantic segmentation extraction; the first analysis includes at least one of lexical semantic analysis, syntactic structure analysis, and contextual logical relationship analysis.

[0058] In some embodiments of this application, it is necessary to perform refined parsing on the first structured data recorded as text data in the data document, in order to accurately extract multiple first concept names from the text and generate corresponding concept explanation data in combination with contextual semantics. Specifically, when the data document contains first structured data recorded as text data, a first extraction process can be performed on the first structured data to obtain multiple first concept names, and concept explanation data for each first concept name can be generated based on a first analysis. The first extraction process includes sentence parsing extraction and / or semantic segmentation extraction; the first analysis includes at least one of lexical semantic analysis, syntactic structure analysis, and contextual logical relationship analysis. This transforms the semantic information in the text-based structured data into a structured set of concept names and explanation data, thereby improving the accuracy and completeness of concept extraction.

[0059] In a specific example, when processing a chemistry lab report, the system identifies the first structured data as a text paragraph. In this sub-step, the system segments and parses this text paragraph, extracting the concept name "catalyst," and further generating the explanatory data "a substance that lowers the activation energy in a chemical reaction" through lexical semantic analysis and contextual logical relationship analysis. Simultaneously, the system extracts the concept name "reaction rate" from another text paragraph and generates the explanatory data "change in reactant concentration per unit time" through syntactic structure analysis. This yields multiple first concept names and their corresponding explanatory data, providing a foundation for subsequent relational attribute generation and three-dimensional data set construction.

[0060] In the embodiments of this application, the first extraction and first analysis processes can employ Natural Language Processing (NLP), Named Entity Recognition (NER), or a semantic parsing model based on the Transformer architecture. These models can perform sentence segmentation, semantic segmentation, and contextual logic analysis on textual structured data to generate concept explanation data. The training dataset used to train the above models may include a scientific literature corpus, a manually annotated set of concept names and explanation data, and annotated data of contextual semantic relationships. The manually annotated set refers to the data set that explicitly marks the correspondence between concept names and explanation data in the literature text, and the contextual semantic annotation data refers to the data set that marks the semantic relationship between concept names and context in the literature text. It should be noted that the specific training methods and implementation details of the models have been fully disclosed and applied in the prior art, and will not be repeated here.

[0061] Sub-step 201112: In the case that the data document contains second structured data recorded as chart data, perform a second extraction process on the second structured data to obtain multiple second concept names, and generate concept explanation data for each second concept name based on the second analysis.

[0062] The second extraction process includes table cell parsing and / or graphic element recognition and extraction; the second analysis includes at least one of numerical attribute analysis, graphic annotation analysis, and table logical relationship analysis.

[0063] In some embodiments of this application, it is necessary to parse the second structured data recorded as chart data in the data document in order to extract multiple second concept names from the tables or graphs and generate corresponding concept explanation data by combining numerical or annotation information. In this case, when the data document contains second structured data recorded as chart data, a second extraction process can be performed on the second structured data to obtain multiple second concept names, and concept explanation data for each second concept name can be generated based on a second analysis. The second extraction process includes table cell parsing and / or graphic element identification and extraction; the second analysis includes at least one of numerical attribute analysis, graphic annotation analysis, and table logical relationship analysis. This will transform the numerical and annotation information in the chart-type structured data into a structured set of concept names and explanation data, thereby improving the comprehensiveness and accuracy of concept extraction.

[0064] In a specific example, when researchers were processing charts and graphs in a chemistry lab report, the system identified the secondary structured data as experimental result tables and reaction rate curves. During table cell parsing and extraction, the system extracted the concept name "reaction temperature" and generated explanatory data "temperature variable under experimental conditions" through numerical attribute analysis. During graphical element identification and extraction, the system extracted the concept name "reaction rate curve" and generated explanatory data "trend of reaction rate with temperature" through graphical annotation analysis. Simultaneously, the system identified the correspondence between "reaction temperature" and "reaction rate" through table logical relationship analysis. This yielded multiple secondary concept names and their corresponding explanatory data, providing a foundation for subsequent relational attribute generation and three-dimensional data set construction.

[0065] In the embodiments of this application, the second extraction and analysis processes can employ multimodal analytical models (Transformers) based on the Transformer structure. These models are capable of recognizing table cells and graphical elements, and generating conceptual explanation data by combining numerical attributes, annotation information, and logical relationships. The training dataset used to train the above models may include scientific literature corpora, manually annotated sets of tables and graphs, and annotated data of numerical logical relationships. It should be noted that the specific training methods and implementation details of the models have been fully disclosed and applied in the prior art, and will not be repeated here.

[0066] Sub-step 201113: In the case that the data document contains unstructured data, perform third extraction processing on the unstructured data to obtain multiple third concept names, and generate concept explanation data for each third concept name based on the third analysis.

[0067] The third extraction process includes feature extraction and / or pattern recognition extraction; the third analysis includes at least one of semantic clustering analysis, context matching analysis, and statistical feature analysis.

[0068] In some embodiments of this application, it is necessary to process unstructured data contained in data documents to extract multiple third-party concept names from complex, unformatted text or symbolic information, and generate corresponding concept explanation data by combining semantic or statistical features. Specifically, when unstructured data is contained in the data document, third-party extraction processing can be performed on the unstructured data to obtain multiple third-party concept names, and concept explanation data for each third-party concept name can be generated based on third-party analysis. The third-party extraction processing includes feature extraction and / or pattern recognition extraction; the third-party analysis includes at least one of semantic clustering analysis, context matching analysis, and statistical feature analysis. This transforms the potential semantic information in unstructured data into a structured set of concept names and explanation data, thereby improving the coverage and accuracy of the system when processing complex documents.

[0069] In a specific example, when processing research notes in the field of chemistry, the system identifies the unstructured data contained within as free text descriptions and experimental logs. In this sub-step, the system identifies the concept name "reaction byproducts" through feature extraction and generates explanatory data "non-target compounds generated during the reaction" through semantic clustering analysis. Simultaneously, the system identifies the concept name "experimental anomalies" through pattern recognition extraction and generates explanatory data "phenomena caused by experimental conditions deviating from the preset range" through context matching analysis. Furthermore, the system identifies "frequently occurring chemical substance names" through statistical feature analysis and generates explanatory data "key chemical substances appearing multiple times in the experimental records." This yields multiple third-party concept names and their corresponding explanatory data, providing a foundation for subsequent relational attribute generation and three-party data set construction.

[0070] In the embodiments of this application, the third extraction and third analysis processes can employ an unstructured data parsing model (Transformer) based on the Transformer architecture. These models are capable of feature extraction, pattern recognition, and semantic clustering of unstructured data, thereby generating concept explanation data. The training dataset used to train the above models may include a scientific literature corpus, a manually annotated set of unstructured text, and annotated data of semantic clustering relationships. Here, the clustering annotation data refers to the data set that marks the semantic clustering relationships between concepts in unstructured text. It should be noted that the specific training methods and implementation details of the models have been fully disclosed and applied in the prior art, and will not be repeated here.

[0071] Sub-step 20112: Determine the semantic connections and / or citation relationships of each concept name in the corresponding data documents, and generate the corresponding relational attribute data based on the determined semantic connections and / or citation relationships of each concept name in the corresponding data documents.

[0072] In some embodiments of this application, after extracting concept names and explanatory data, it is necessary to further identify the semantic connections and / or citation relationships of each concept name in the corresponding data documents in order to generate relational attribute data that can characterize the logical relationships between concepts. Specifically, the semantic connections and / or citation relationships of each concept name in the corresponding data documents can be determined, and corresponding relational attribute data can be generated based on the determined semantic connections and / or citation relationships. Relational attribute data refers to structured information used to describe the logical dependencies, causal relationships, or citation relationships between concept names, which can support the construction of concept networks in subsequent steps. In this way, the implicit semantic relationships in the documents can be made explicit, so that the concept data not only contains a single explanation, but also has information on its association with other concepts, thereby improving the completeness and accuracy of knowledge extraction.

[0073] In a specific example, when parsing literature on "catalytic reaction mechanisms," the system has already extracted the concept names "catalyst," "reaction rate," and "activation energy." Then, when the system further identifies a semantic relationship between "catalyst" and "reaction rate" in the literature, characterized as "catalysts can affect reaction rates," and simultaneously identifies a causal relationship between "activation energy" and "reaction rate," characterized as "the magnitude of activation energy determines the reaction rate," the system generates relational attribute data, such as {catalyst—reaction rate, influence relationship} and {activation energy—reaction rate, causal relationship}. This yields multiple relational attribute data entries, providing clear input for the subsequent construction of the three-element data set.

[0074] Sub-step 2012 involves organizing the corresponding concept names, concept explanation data, and relational attribute data to form three-dimensional data sets, and then identifying each three-dimensional data set as a concept data in the initial concept dataset.

[0075] In some embodiments of this application, it is necessary to uniformly organize the concept names, concept explanation data, and relational attribute data obtained from the preceding parsing to ensure the integrity and processability of the concept data. This can be achieved by organizing the corresponding concept names, concept explanation data, and relational attribute data into three-dimensional data sets, with each three-dimensional data set designated as a concept data entry in the initial concept dataset. A three-dimensional data set is a structured unit composed of concept names, concept explanation data, and relational attribute data. Concept explanation data typically exists in the form of data entities or metadata, used to describe the semantics or attributes corresponding to the concept name, enabling the system to accurately understand the meaning of the concept. Concept names are generally represented using digital object encoding to ensure consistent identification and storage of the same concept across different documents. Relational attribute data describes the relationships between concept names and can be presented in the form of a node tree or node network, thus intuitively reflecting the logical connections and semantic dependencies between different concepts. This transforms the scattered parsing results into unified concept data entries, giving the initial concept dataset a clear structure and semantic relationships, providing stable input for subsequent concept disambiguation and knowledge graph generation.

[0076] In a specific example, within a chemical literature analysis scenario, the system has extracted the concept names "catalyst," "reaction rate," and "activation energy" from literature on "catalytic reaction mechanisms," generating corresponding concept explanation data and relational attribute data. The system then further organizes these analysis results into multiple three-element data sets, such as {catalyst, substance that lowers activation energy, reaction rate} and {activation energy, minimum energy required for a chemical reaction, reaction rate}. Here, "catalyst" and "activation energy" serve as concept names encoded as digital objects, "substance that lowers activation energy" and "minimum energy required for a chemical reaction" are concept explanation data in metadata form, and "reaction rate" serves as relational attribute data reflecting their logical connections. The system identifies these three-element data sets as multiple concept data points in the initial concept dataset. In this way, the system obtains a set containing multiple structured concept entries, providing clear input for subsequent concept disambiguation and knowledge graph generation.

[0077] Step 202: Compare the concept explanation data of each concept data in the initial concept dataset to merge the concept data corresponding to concept names that are characterized as having the same conceptual connotation, so as to obtain the concept dataset after disambiguation.

[0078] The method shown in this step has been explained in step 102 and will not be repeated here.

[0079] Optionally, step 202 includes the following sub-steps:

[0080] Sub-step 2021: Determine the similarity evaluation value between the various concept explanation data to identify concept names that are represented as having the same semantic connotation.

[0081] In some embodiments of this application, it is necessary to determine the semantic similarity between different concept explanation data through quantitative methods in order to identify concept names that are characterized by the same semantic connotation, thereby providing a basis for subsequent concept fusion. Specifically, a similarity evaluation value can be determined between each concept explanation data to identify concept names that are characterized by the same semantic connotation. The similarity evaluation value refers to a numerical value obtained through semantic matching or vector calculation methods, used to measure the semantic closeness between two concept explanation data. In this way, semantically consistent and semantically different concept names can be accurately distinguished in the initial concept dataset, thereby reducing the problem of duplicate or incorrect matching caused by semantic ambiguity.

[0082] In a specific example, when processing a biomedical literature dataset, the system found that the explanation for "myocardial infarction" in the initial concept dataset was "a pathological state of myocardial necrosis due to insufficient blood supply," while the explanation for "heart attack" was "a pathological state of cardiac dysfunction due to blood flow obstruction." In this sub-step, the system calculated the similarity between these two concept explanations. The similarity score exceeded a preset threshold, thus identifying that the two concept names represent the same semantic connotation. In this way, the system obtained a clear matching relationship, providing a basis for subsequent concept fusion.

[0083] Sub-step 2022 involves associating concept names with the same semantic connotations with term entries in a pre-defined domain term dataset.

[0084] In some embodiments of this application, it is necessary to map identified concept names with the same semantic connotations to unified domain terminology to ensure that concepts appearing in different documents maintain consistency at the data level. Specifically, concept names with the same semantic connotations can be associated with term entries in a pre-defined domain terminology dataset. A domain terminology dataset refers to a standardized set of terms pre-established in a specific discipline or research field, used to uniformly represent the semantic connotations of concept names. By associating with standard terms, semantic differences caused by different expressions can be reduced, making the concept dataset more unified and standardized in subsequent processing.

[0085] In a specific example, when processing a biomedical literature dataset, the system has identified that "myocardial infarction" and "heart attack" have the same semantic connotation. The system then associates these two concept names with the term "myocardial infarction" entry in a pre-defined domain terminology dataset. This results in a clear mapping relationship, where concept names from different documents are uniformly associated with the same standard terminology entry, providing a foundation for subsequent concept fusion.

[0086] Sub-step 2023 involves fusing the concept data corresponding to all concept names associated with the same term in the initial concept dataset to obtain a disambiguated concept dataset.

[0087] In some embodiments of this application, after associating concept names with domain terminology entries, it is necessary to fuse all concept data corresponding to the same terminology entry in the initial concept dataset to eliminate semantic duplication or ambiguity caused by different expressions. This involves fusing the concept data corresponding to all concept names associated with the same terminology entry in the initial concept dataset to obtain an ambiguous concept dataset. The ambiguous concept dataset refers to a dataset where a unified terminology entry is retained as the concept name during the fusion process, and its corresponding explanatory data and relational attribute data are integrated to form semantically consistent concept entries. This reduces the problem of duplicate matching caused by concept ambiguity, making the concept dataset more unified and accurate, and providing semantically consistent input for subsequent knowledge graph generation.

[0088] In a specific example, when processing a biomedical literature dataset, the system has already associated "myocardial infarction" and "heart attack" with the same term entry "myocardial infarction" in the domain terminology dataset. The system then further merges the conceptual data corresponding to these two names, uniformly retaining the term entry "myocardial infarction" as the conceptual name, and integrating its explanatory data "the pathological state of myocardial necrosis due to insufficient blood supply" and relational attribute data "semantic association with heart attack." This results in a disambiguated conceptual dataset where "myocardial infarction" serves as a unified conceptual entry, encompassing both semantic explanation and relationships with other concepts, providing semantically consistent input for subsequent knowledge graph generation.

[0089] Optionally, sub-step 2023 includes the following sub-steps:

[0090] Sub-step 20231 merges the concept explanation data and concept name of each concept data corresponding to the same term entry to obtain the explanation merged data.

[0091] In some embodiments of this application, after associating concept names with terminology entries, it is necessary to interpret and merge the names of multiple concept data corresponding to the same terminology entry to form unified interpretation-merged data. Specifically, the concept interpretation data and concept names of each concept data corresponding to the same terminology entry can be merged to obtain interpretation-merged data. Interpretation-merged data refers to a set of concept descriptions unified at the semantic level. It can integrate diverse descriptions of the same term in different documents and retain the digital object encoding form of concept names, thereby ensuring data consistency and integrity. This reduces semantic redundancy caused by different expressions, making concept interpretation more focused and unified, and providing clear input for subsequent attribute integration and concept updates.

[0092] In a specific example, when processing a biomedical literature dataset, the system has already associated "myocardial infarction" and "heart attack" with the same term entry, "myocardial infarction." The system can then merge the explanations for these two concepts, for example, "the pathological state of myocardial necrosis due to insufficient blood supply" and "the pathological state of cardiac dysfunction due to blood flow obstruction," and unify them into the merged explanation data "the pathological state of myocardial necrosis due to insufficient blood supply or blood flow obstruction." Simultaneously, the system retains the concept names "myocardial infarction" and "heart attack" in a unified digital object encoding format as the name set for the merged explanation data. This ultimately results in a semantically unified merged explanation data, providing a foundation for subsequent attribute integration.

[0093] In the embodiments of this application, the explanation merging process can utilize NLP or a semantic merging model based on the Transformer architecture. These models can semantically align and merge multiple concept explanation data under the same term, outputting unified explanation merged data. The training dataset used to train the above model may include a scientific literature corpus, a manually annotated set of concept explanation data, and semantically merged annotated data. The manually annotated set refers to a data set that explicitly marks the semantic correspondence between concept explanation data in the literature text, and the merged annotated data refers to a data set that marks the semantic fusion results in the concept explanation data set. It should be noted that the specific training methods and implementation details of the model have been fully disclosed and applied in the prior art, and will not be repeated here.

[0094] Sub-step 20232 involves integrating the relational attribute data of various conceptual data corresponding to the same term entry to generate attribute integration data.

[0095] In some embodiments of this application, after the interpretation and merging data generation is completed, it is necessary to integrate the relational attribute data of multiple conceptual data corresponding to the same term entry in order to form unified attribute integration data. Specifically, the relational attribute data of various conceptual data corresponding to the same term entry can be integrated by associating them to generate attribute integration data. Attribute integration data refers to a set of relations unified at the semantic level, which can integrate the conceptual relations involved in the same term in different documents and present them in the form of a node tree or node network. This reduces semantic redundancy caused by scattered relational attributes, making the conceptual relations more concentrated and unified, and providing clear input for the subsequent generation of concept update data.

[0096] In a specific example, when researchers are processing a biomedical literature dataset, the system has already associated "myocardial infarction" and "heart attack" with the same term entry, "myocardial infarction." The system will then integrate the relational attribute data of these two concepts, for example, "myocardial infarction—heart attack (semantic association)" and "myocardial infarction—coronary artery occlusion (causal relationship)," and uniformly generate integrated attribute data: {myocardial infarction, semantic association: heart attack; causal relationship: coronary artery occlusion}. This results in semantically unified integrated attribute data, providing a foundation for subsequent concept update data generation.

[0097] In the embodiments of this application, the relationship integration process can employ a relationship merging model based on the Transformer structure. These models can semantically align and integrate multiple relationship attribute data under the same term entry, outputting unified attribute integration data. The training dataset used to train the above model may include a scientific literature corpus, a manually annotated set of relationship attribute data, and annotated data for relationship integration. The manually annotated set refers to the data set that explicitly marks conceptual relationships in the literature text, and the integrated annotation data refers to the data set that marks the semantic fusion results in the conceptual relationship set. It should be noted that the specific training methods and implementation details of the model have been fully disclosed and applied in the prior art, and will not be repeated here.

[0098] Sub-step 20233 involves organizing the corresponding term entries, explanation merged data, and attribute integration data to form at least one concept update data set. All the concept update data sets are then added to the concept dataset as new concept data sets, and all target concept data sets in the concept dataset are deleted to obtain the disambiguated concept dataset.

[0099] The target concept data refers to the concept data corresponding to each concept data item in the added concept update data, which is the terminology in the concept update data.

[0100] In some embodiments of this application, after generating the interpretation-merged data and attribute-integrated data, it is necessary to organize them in a unified manner with the corresponding terminology entries to form new concept update data, replacing the original target concept data, thereby obtaining a disambiguated concept dataset. Specifically, the corresponding terminology entries, interpretation-merged data, and attribute-integrated data can be organized to form at least one concept update data, and all concept update data are added to the concept dataset as new concept data, while all target concept data in the concept dataset are deleted. Target concept data refers to the concept data corresponding to each concept data corresponding to the terminology entries in the added concept update data. This achieves semantic unification and redundancy elimination in the concept dataset, ensuring that the final concept dataset retains only the disambiguated concept update data, thereby improving data consistency and usability.

[0101] In a specific example, when processing a biomedical literature dataset, the system has already associated "myocardial infarction" and "heart attack" with the same term entry, "myocardial infarction," and generated explanatory merged data and attribute integrated data. The system then further organizes the term entry "myocardial infarction," the explanatory merged data "pathological state of myocardial necrosis due to insufficient blood supply or blood flow obstruction," and the attribute integrated data {semantic association: heart attack; causal relationship: coronary artery obstruction} into a concept update dataset. The system then adds this concept update data to the concept dataset and deletes the original target concept data "myocardial infarction" and "heart attack." This results in a disambiguated concept dataset where "myocardial infarction" serves as a unified concept entry, containing both semantic explanation and relationships with other concepts, providing semantically consistent input for subsequent knowledge graph generation.

[0102] Step 203: Generate knowledge graph data of the literature dataset based on the unambiguous concept dataset.

[0103] Among them, knowledge graph data is used to represent the relationships between various concept names appearing in the literature dataset, as well as the data literature corresponding to each concept name.

[0104] The method shown in this step has been explained in step 103 and will not be repeated here.

[0105] Optionally, step 203 includes the following sub-steps:

[0106] Sub-step 2031: Based on all the conceptual data in the deambiguous conceptual dataset, determine the conceptual relationship data used to characterize the literature dataset.

[0107] The concept relationship data contains multiple data nodes and data edges used to connect the multiple data nodes; each data node corresponds to a concept name and all the data documents corresponding to the corresponding concept name, and each data edge corresponds to two concept names that have a relationship.

[0108] In some embodiments of this application, after concept disambiguation, it is necessary to further determine concept relationship data that can represent the entire document dataset in order to provide structured input for knowledge graph generation. Specifically, based on all concept data in the disambiguated concept dataset, concept relationship data to represent the document dataset is determined. This concept relationship data includes multiple data nodes and data edges connecting these nodes. Each data node corresponds to a concept name and all the corresponding documents for that concept name, and each data edge corresponds to two related concept names. This allows for a node-based representation of concepts in the document dataset and their corresponding documents, and the data edges reveal the relationships between concepts, thereby giving the document dataset a visual, computable, and scalable structured feature.

[0109] In a specific example, when processing a literature dataset in the field of chemistry, the system has already eliminated ambiguity and standardized the conceptual names such as "catalyst," "reaction rate," and "activation energy." The system then identifies these conceptual names as data nodes and binds them to all the corresponding literature. For example, the node "catalyst" corresponds to multiple documents involving catalysis, and the node "reaction rate" corresponds to multiple documents involving rate measurement. The system further generates data edges based on relational attribute data, such as establishing a data edge between "catalyst" and "reaction rate" to represent their association in the literature. The system then obtains a conceptual relationship data set containing multiple data nodes and data edges, providing a clear and structured input for subsequent generation of knowledge graph data.

[0110] Sub-step 2032: Generate knowledge graph data based on concept relationship data.

[0111] In some embodiments of this application, after determining the conceptual relationship data, it is necessary to further transform it into knowledge graph data to form a graph representation that can intuitively represent the semantic structure of the document dataset. At this time, knowledge graph data can be generated based on the conceptual relationship data. Knowledge graph data refers to a graph structure composed of multiple conceptual nodes and the edges between them. Each node corresponds to a conceptual name and its associated documents, and each edge corresponds to the association between two conceptual names. In this way, the concepts and relationships in the document dataset can be uniformly expressed in graph form, enabling the system to have the capabilities of semantic retrieval, relational reasoning, and visualization.

[0112] In a specific example, the system has generated conceptual relationship data, including nodes "catalyst," "reaction rate," and "activation energy," as well as edges "catalyst-reaction rate" and "activation energy-reaction rate." The system can then organize these nodes and edges into a knowledge graph data set, forming a graph structure: the node "catalyst" corresponds to multiple documents related to catalysis, the node "reaction rate" corresponds to documents on rate measurement, the node "activation energy" corresponds to documents related to energy thresholds, and the edges visually represent the semantic relationships between them. This ultimately yields a complete knowledge graph data set, providing structured support for the subsequent establishment of a retrieval data model.

[0113] Step 204: Establish a retrieval data model based on the relationships between the various concept names represented in the knowledge graph data and the data documents corresponding to each concept name.

[0114] The retrieval data model is used to determine the target concept name based on the input retrieval text, and to determine the target data document based on the determined target concept name.

[0115] In some embodiments of this application, a structured model supporting retrieval needs to be established based on knowledge graph data. This enables the system to quickly locate the target concept name and trace back to the corresponding data document based on the input retrieval text. At this point, a retrieval data model can be established based on the relationships between various concept names represented in the knowledge graph data, and the data documents corresponding to each concept name. The retrieval data model is used to determine the target concept name based on the input retrieval text, and to determine the target data document based on the determined target concept name. The retrieval data model refers to the retrieval structure formed by structurally modeling the knowledge graph data. It can combine concept nodes, concept relationships, and document mapping information to achieve matching from text input to target documents. This allows the system to quickly complete concept identification and document location upon receiving the retrieval text, thereby improving the accuracy and efficiency of the retrieval.

[0116] In a specific example, when researchers are processing a dataset of literature in the field of physics, the system has already generated knowledge graph data containing concepts such as "quantum state," "superposition principle," and "measurement collapse," along with their corresponding literature. A retrieval data model can then be built based on the relationships between these concept names and their mapping to literature. When a user inputs the search text "measurement process of quantum state," the system identifies the target concept name as "measurement collapse" through the retrieval data model and further determines the target data literature as relevant literature containing this concept. Ultimately, a retrieval data model is obtained that supports text input and literature location, providing clear model support for subsequent retrieval matching.

[0117] Step 205: In response to the search text of the target data document, a match is performed in the search data model to determine the target concept name for the search text.

[0118] In some embodiments of this application, semantic matching of the input search text is required in the retrieval data model to determine the conceptual object corresponding to the data document used for retrieval and location. Specifically, in response to the search text of the target data document, matching is performed in the retrieval data model to determine the target concept name of the search text. The target concept name refers to the retrieval object selected after matching the search text with concept nodes in the knowledge graph. It is used to characterize the core concept pointed to by the search intent and serves as the basis for subsequent location of data documents. By establishing a clear correspondence between the search text and concept nodes in the knowledge graph, it is possible to support the subsequent generation of document lists and the determination of target data documents.

[0119] In a specific example, the user inputs the search text "thermal stability of polymer materials." In this step, the system calls the search data model for matching, identifies the target concept name as "thermal stability," and further determines that this concept name has a spectral association with "polymer materials." The system then obtains the target concept name "thermal stability," which will serve as the core input for subsequently generating a literature list and filtering target data literature.

[0120] Step 206: Generate a list of literatures for the target data based on the search text and the target concept name, and identify the target data literatures from the list of literatures.

[0121] In some embodiments of this application, after identifying the target concept name, it is necessary to further generate a list of documents related to that concept name in order to filter out the target data documents that best match the semantics of the search text from multiple candidate documents. At this time, a list of target data documents can be generated based on the search text and the target concept name, and the target data documents can be determined from the list. This will enable the rapid location of the target data documents that best meet the search requirements from the candidate document set, thereby improving the accuracy and efficiency of the search.

[0122] In a specific example, a user enters the search text "conductivity properties of nanomaterials" in a scientific research search scenario. The system has identified the target concept as "conductivity properties" and generates a literature list based on this concept in this step, which includes multiple data documents related to "conductivity properties." The system further filters the literature list by combining the semantics of "nanomaterials" in the search text, ultimately determining that the target data documents are several research papers on "conductivity property testing of nanomaterials." Then, the system sorts them according to preset conditions (such as publication time), resulting in the final literature list for the user to select from. This ensures that the literature not only matches the semantics of the target concept name but also remains consistent with the research object of the search text.

[0123] like Figure 3 The diagram illustrates the workflow of a data storage and retrieval system based on an embodiment of this application. It aims to process complex, multimodal raw data and transform it into structured knowledge, ultimately supporting intelligent retrieval and semantic reasoning for users. The entire process is divided into two stages: a knowledge generation stage and a knowledge retrieval stage.

[0124] The raw data is first input into the system (S1) and then enters the data processing stage. During this process, the system parses and preprocesses the complex multimodal data and generates semantic labels (S2). These labels are uniformly managed by the labeling system to ensure consistency of data from different sources in subsequent processing. The labeling results then enter multimodal processing and storage, where semantic fusion and context enhancement are completed with the assistance of a large model. The data is then written to the multimodal database through associative storage (S4) to store non-textual information such as images and tables. Simultaneously, the results of data processing are also provided for knowledge graph reading and construction. With the assistance of a large model, concept nodes and relation edges are organized to generate a knowledge graph (S3), which is then stored in the knowledge graph database to intuitively represent the logical connections between concepts.

[0125] When a user needs to perform a search: The user initiates a query (T1), which is received and parsed by the intelligent agent. Subsequently, in the tool invocation (T2) stage, knowledge retrieval and retrieval enhancement begin. The system can retrieve structured semantic relationships from a knowledge graph database (T3) or obtain contextual information such as images or tables from a multimodal database (T4), and integrate this data into a verifiable context. This context is input into a large model for semantic understanding and response generation (T5), and finally, the large model returns the results to the user (T6), completing a full intelligent retrieval and knowledge response process.

[0126] In summary, in this embodiment, by performing structured parsing on the literature dataset, the concept names, concept explanations, and their relationships in the literature are extracted as an initial concept dataset. This allows knowledge units scattered across different documents to be stored and managed in a unified structured form, ensuring data organization and processability. While ensuring traceability in data retrieval, this provides a foundation for subsequent concept disambiguation and knowledge graph generation. Furthermore, by comparing the concept explanations, different concept names representing the same concept are identified and merged, reducing duplicate or incorrect matching problems caused by concept ambiguity. This lowers the risk of mismatches common in semantic retrieval, making the concept dataset more unified and accurate, providing semantically consistent input for knowledge graph construction. Finally, knowledge graph data is generated based on the disambiguated concept dataset, enabling each concept name to not only establish clear relationships with other concepts but also to establish a clear mapping with the corresponding literature data. This allows searchers to quickly locate target concepts in the graph and directly trace back to relevant literature fragments, thereby improving search efficiency and the reliability of results. Therefore, the method based on the embodiments of this application, through the core idea of ​​joint storage-concept disambiguation, realizes the layer-by-layer transformation from original documents to structured knowledge, thereby effectively improving the problems of semantic retrieval mismatch and lack of precise document correspondence in knowledge graphs, thus improving the accuracy and efficiency of retrieval and providing researchers with higher quality knowledge retrieval support.

[0127] refer to Figure 4 It illustrates a knowledge graph and digital object hybrid device 30 for scientific discovery provided in an embodiment of this application, comprising:

[0128] The parsing storage module 301 is used to parse the literature dataset to obtain the initial concept dataset of the literature dataset; the literature dataset contains at least one data document; the concept dataset contains multiple concept data; each concept data is used to represent the corresponding concept name and concept explanation data in the literature dataset, as well as the relationship between the concept names;

[0129] The ambiguity elimination module 302 is used to compare the concept explanation data of each concept data in the initial concept dataset, so as to merge the concept data corresponding to the concept names that represent the same concept connotation, so as to obtain the concept dataset after ambiguity elimination.

[0130] The knowledge graph creation module 303 is used to generate knowledge graph data of the literature dataset based on the concept dataset after the ambiguity is eliminated; the knowledge graph data is used to represent the relationship between the various concept names appearing in the literature dataset, as well as the data literature corresponding to each concept name.

[0131] Optionally, the parsing storage module 301 includes:

[0132] The parsing submodule is used to perform structured parsing on each data document in the literature dataset to obtain multiple sets of corresponding concept names and concept explanation data, and generate relational attribute data corresponding to each concept name; the relational attribute data is used to describe at least one concept name that has an association with the corresponding concept name;

[0133] The organization submodule is used to organize the corresponding concept names, concept explanation data, and relational attribute data to form three-dimensional data groups, and to identify each three-dimensional data group as a concept data in the initial concept dataset.

[0134] Optionally, the parsing submodule includes:

[0135] The parsing unit is used to perform hierarchical parsing of the data documents according to the data types contained in each data document, so as to obtain the concept names in each data document, and generate corresponding concept explanation data based on the contextual semantics of each concept name in the corresponding data document.

[0136] The relation unit is used to determine the semantic connections and / or citation relationships of each concept name in the corresponding data documents, and to generate corresponding relation attribute data based on the determined semantic connections and / or citation relationships of each concept name in the corresponding data documents.

[0137] Optionally, the parsing unit includes:

[0138] The first parsing subunit is used to perform a first extraction process on the first structured data recorded as text data in the case of the data document, to obtain multiple first concept names, and to generate concept explanation data for each first concept name based on the first analysis; the first extraction process includes sentence parsing extraction and / or semantic segmentation extraction; the first analysis includes at least one of lexical semantic analysis, syntactic structure analysis, and contextual logical relationship analysis;

[0139] The second parsing subunit is used to perform a second extraction process on the second structured data recorded as chart data in the case where the data document contains second structured data, to obtain multiple second concept names, and to generate concept explanation data for each second concept name based on the second analysis; the second extraction process includes table unit parsing extraction and / or graphic element identification extraction; the second analysis includes at least one of numerical attribute analysis, graphic annotation analysis, and table logical relationship analysis;

[0140] The third parsing subunit is used to perform third extraction processing on unstructured data in the case of unstructured data in the data document to obtain multiple third concept names, and generate concept explanation data for each third concept name based on the third analysis; the third extraction processing includes feature extraction and / or pattern recognition extraction; the third analysis includes at least one of semantic clustering analysis, context matching analysis and statistical feature analysis.

[0141] Optionally, the ambiguity resolution module 302 includes:

[0142] The inductive submodule is used to determine the similarity evaluation value between the various concept explanation data in order to identify concept names that are represented as having the same semantic connotation;

[0143] The association submodule is used to associate concept names with the same semantic connotations with term entries in a pre-defined domain term dataset;

[0144] The disambiguation module is used to merge the concept data corresponding to all concept names associated with the same term in the initial concept dataset to obtain a disambiguated concept dataset.

[0145] Optionally, the disambiguation submodule includes:

[0146] The merging unit is used to merge the concept explanation data and concept name of various concept data corresponding to the same term entry to obtain the explanation merged data;

[0147] The integration unit is used to integrate the relational attribute data of various conceptual data corresponding to the same term entry to generate attribute integration data;

[0148] The disambiguation unit is used to organize the corresponding term entries, explanation merged data, and attribute integration data to form at least one concept update data. All the concept update data are added to the concept dataset as new concept data, and all target concept data in the concept dataset are deleted to obtain the disambiguated concept dataset. The target concept data is the concept data corresponding to each concept data corresponding to the term entries in the added concept update data.

[0149] Optionally, the map building module 303 includes:

[0150] The relational data submodule is used to determine the conceptual relational data used to characterize the literature dataset based on all conceptual data in the disambiguated conceptual dataset. The conceptual relational data contains multiple data nodes and data edges used to connect the multiple data nodes. Each data node corresponds to a concept name and all the data documents corresponding to the corresponding concept name. Each data edge corresponds to two concept names that have an association relationship.

[0151] The knowledge graph data submodule is used to generate knowledge graph data based on concept relationship data.

[0152] Optionally, the knowledge graph and digital object hybrid device 30 for scientific discovery also includes:

[0153] The model building module is used to build a retrieval data model based on the relationships between the various concept names represented in the knowledge graph data and the data documents corresponding to each concept name. The retrieval data model is used to determine the target concept name based on the input retrieval text and to determine the target data documents based on the determined target concept name.

[0154] Optionally, the knowledge graph and digital object hybrid device 30 for scientific discovery also includes:

[0155] The matching module is used to perform matching in the retrieval data model in response to the retrieval text of the target data document in order to determine the target concept name for the retrieval text;

[0156] The output module is used to generate a list of literatures for the target data based on the search text and the target concept name, and to identify the target data literatures from the list of literatures.

[0157] In summary, in this embodiment, by performing structured parsing on the literature dataset, the concept names, concept explanations, and their relationships in the literature are extracted as an initial concept dataset. This allows knowledge units scattered across different documents to be stored and managed in a unified structured form, ensuring data organization and processability. While ensuring traceability in data retrieval, this provides a foundation for subsequent concept disambiguation and knowledge graph generation. Furthermore, by comparing the concept explanations, different concept names representing the same concept are identified and merged, reducing duplicate or incorrect matching problems caused by concept ambiguity. This lowers the risk of mismatches common in semantic retrieval, making the concept dataset more unified and accurate, providing semantically consistent input for knowledge graph construction. Finally, knowledge graph data is generated based on the disambiguated concept dataset, enabling each concept name to not only establish clear relationships with other concepts but also to establish a clear mapping with the corresponding literature data. This allows searchers to quickly locate target concepts in the graph and directly trace back to relevant literature fragments, thereby improving search efficiency and the reliability of results. Therefore, the method based on the embodiments of this application, through the core idea of ​​joint storage-concept disambiguation, realizes the layer-by-layer transformation from original documents to structured knowledge, thereby effectively improving the problems of semantic retrieval mismatch and lack of precise document correspondence in knowledge graphs, thus improving the accuracy and efficiency of retrieval and providing researchers with higher quality knowledge retrieval support.

[0158] Reference Figure 5 The electronic device 500 may include one or more of the following components: processing component 502, memory 504, power supply component 506, multimedia component 508, audio component 510, input / output (I / O) interface 512, sensor component 514, and communication component 516.

[0159] Processing component 502 typically controls the overall operation of electronic device 500, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 502 may include one or more processors 520 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 502 may include one or more modules to facilitate interaction between processing component 502 and other components. For example, processing component 502 may include a multimedia module to facilitate interaction between multimedia component 508 and processing component 502.

[0160] Memory 504 is used to store various types of data to support the operation of electronic device 500. Examples of this data include instructions for any application or method operating on electronic device 500, contact data, phonebook data, messages, pictures, multimedia, etc. Memory 504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0161] Power supply component 506 provides power to various components of electronic device 500. Power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 500.

[0162] Multimedia component 508 includes an interface that provides an output interface between electronic device 500 and user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). If the interface includes a touch panel, the interface may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 508 includes a front-facing camera and / or a rear-facing camera. When electronic device 500 is in an operating mode, such as shooting mode or multimedia mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0163] Audio component 510 is used to output and / or input audio signals. For example, audio component 510 includes a microphone (MIC) used to receive external audio signals when electronic device 500 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 504 or transmitted via communication component 516. In some embodiments, audio component 510 also includes a speaker for outputting audio signals.

[0164] Input / output (I / O) interface 512 provides an interface between processing component 502 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0165] Sensor assembly 514 includes one or more sensors for providing state assessments of various aspects of electronic device 500. For example, sensor assembly 514 may detect the on / off state of electronic device 500, the relative positioning of components such as the display and keypad of electronic device 500, changes in position of electronic device 500 or a component of electronic device 500, the presence or absence of user contact with electronic device 500, orientation or acceleration / deceleration of electronic device 500, and temperature changes of electronic device 500. Sensor assembly 514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 514 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 514 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0166] Communication component 516 facilitates wired or wireless communication between electronic device 500 and other devices. Electronic device 500 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 516 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0167] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement the methods provided in the embodiments of this application.

[0168] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, which can be executed by a processor 520 of an electronic device 500 to perform the above-described method. For example, the non-transitory storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0169] In an exemplary embodiment, the electronic device 500 may also be provided as a server, including a processing component 502, which further includes one or more processors, and memory resources represented by memory 504 for storing instructions, such as applications, that can be executed by the processing component 502. The applications stored in memory 504 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 502 is configured to execute instructions to perform the methods provided in the embodiments of this application.

[0170] Electronic device 500 may also include a power supply component 506 configured to perform power management of electronic device 500, a wired or wireless communication component 516 configured to connect electronic device 500 to a network, and an input / output (I / O) interface 512. Electronic device 500 may operate on an operating system stored in memory 504, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0171] It should be noted that, for the sake of simplicity, the method embodiments of this application are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of this application.

[0172] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.

[0173] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the claimed rights.

Claims

1. A method for integrating knowledge graphs and digital objects for scientific discovery, characterized in that, include: The literature dataset is parsed to obtain the initial conceptual dataset of the literature dataset; The document dataset contains at least one data document; The concept dataset contains multiple concept data; each concept data is used to characterize the corresponding concept name and concept explanation data in the literature dataset, as well as the relationship between the concept names; By comparing the concept explanation data of each concept data in the initial concept dataset, the concept data corresponding to concept names that are characterized as having the same conceptual connotation are merged to obtain the concept dataset after disambiguation; Based on the unambiguous concept dataset, a knowledge graph of the literature dataset is generated; the knowledge graph is used to characterize the relationships between the various concept names appearing in the literature dataset, as well as the data literature corresponding to each concept name. The process of parsing the literature dataset to obtain the initial conceptual dataset of the literature dataset includes: The document dataset is structured and parsed to obtain multiple sets of corresponding concept names and concept explanations, and relational attribute data corresponding to each concept name is generated; the relational attribute data is used to describe at least one concept name that is associated with the corresponding concept name. The corresponding concept names, concept explanation data, and relation attribute data are organized to form three-dimensional data groups, and each of the three-dimensional data groups is determined as a concept data in the initial concept dataset. The comparison involves taking the concept explanation data for each concept data in the initial concept dataset, and fusing the concept data corresponding to concept names that represent the same conceptual connotation, to obtain the disambiguated concept dataset, including: Determine the similarity evaluation value between each of the concept explanation data to identify concept names that are represented as having the same semantic connotation; Associate concept names with the same semantic connotations with term entries in a pre-defined domain term dataset; The concept data corresponding to all concept names associated with the same term in the initial concept dataset are merged to obtain the concept dataset after disambiguation; The step of generating the knowledge graph data of the literature dataset based on the disambiguated concept dataset includes: Based on all concept data in the concept dataset after deambiguation, concept relationship data is determined to characterize the literature dataset; the concept relationship data includes multiple data nodes and data edges for connecting the multiple data nodes; each data node corresponds to a concept name and all data documents corresponding to the corresponding concept name, and each data edge corresponds to two concept names that have an association relationship; The knowledge graph data is generated based on the conceptual relationship data.

2. The method for hybridizing knowledge graphs and digital objects for scientific discovery as described in claim 1, characterized in that, The step involves performing structured parsing on each data document in the document dataset to obtain multiple sets of corresponding concept names and concept explanations, and generating relational attribute data corresponding to each concept name, including: The data documents are hierarchically parsed according to the data types contained in each data document to obtain the concept names in each data document, and corresponding concept explanation data is generated based on the contextual semantics of each concept name in the corresponding data document. Determine the semantic connections and / or citation relationships of each concept name in the corresponding data document, and generate the corresponding relationship attribute data based on the determined semantic connections and / or citation relationships of each concept name in the corresponding data document.

3. The method for hybridizing knowledge graphs and digital objects for scientific discovery as described in claim 1, characterized in that, The method for blending knowledge graphs and digital objects for scientific discovery also includes: Based on the relationships between the concept names represented in the knowledge graph data and the data documents corresponding to each concept name, a retrieval data model is established; the retrieval data model is used to determine the target concept name based on the input retrieval text, and to determine the target data document based on the determined target concept name.

4. The method for hybridizing knowledge graphs and digital objects for scientific discovery as described in claim 3, characterized in that, The method for blending knowledge graphs and digital objects for scientific discovery also includes: In response to the search text of the target data document, a match is performed in the search data model to determine the target concept name for the search text; A list of documents for the target data is generated based on the search text and the target concept name, and the target data documents are identified from the list of documents.

5. A hybrid device for knowledge graphs and digital objects used in scientific discovery, characterized in that, include: The parsing storage module is used to parse the literature dataset to obtain the initial conceptual dataset of the literature dataset; The document dataset contains at least one data document; The concept dataset contains multiple concept data; each concept data is used to characterize the corresponding concept name and concept explanation data in the literature dataset, as well as the relationship between the concept names; The ambiguity resolution module is used to compare the concept explanation data of each concept data in the initial concept dataset, so as to merge the concept data corresponding to the concept names that represent the same concept connotation, so as to obtain the concept dataset after ambiguity resolution. The knowledge graph building module is used to generate knowledge graph data of the literature dataset based on the concept dataset after deambiguation; the knowledge graph data is used to represent the relationship between various concept names appearing in the literature dataset, and the data literature corresponding to each concept name; The parsing and storage module includes: The parsing submodule is used to perform structured parsing on each data document in the document dataset to obtain multiple sets of corresponding concept names and concept explanation data, and generate relational attribute data corresponding to each concept name; the relational attribute data is used to describe at least one concept name that has an association with the corresponding concept name; The organization submodule is used to organize the corresponding concept name, concept explanation data, and relation attribute data to form a three-data group, and to determine each three-data group as a concept data in the initial concept dataset; The ambiguity resolution module includes: The inductive submodule is used to determine the similarity evaluation value between the various concept explanation data in order to identify concept names that are represented as having the same semantic connotation; The association submodule is used to associate concept names with the same semantic connotations with term entries in a pre-defined domain term dataset; The disambiguation submodule is used to merge the concept data corresponding to all concept names associated with the same term in the initial concept dataset to obtain the disambiguated concept dataset. The map creation module includes: The relational data submodule is used to determine the conceptual relational data used to characterize the document dataset based on all conceptual data in the conceptual dataset after disambiguation; the conceptual relational data includes multiple data nodes and data edges for connecting the multiple data nodes; each data node corresponds to a concept name and all data documents corresponding to the corresponding concept name, and each data edge corresponds to two concept names that have an association relationship; The knowledge graph data submodule is used to generate the knowledge graph data based on the concept relationship data.

6. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the knowledge graph and digital object hybrid method for scientific discovery as described in any one of claims 1 to 4.

7. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the knowledge graph and digital object hybrid method for scientific discovery as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Knowledge concept construction method and device

    CN113268608A

  • Information retrieval query method for legal data service platform

    CN118981512A