Multi-field scientific knowledge base automatic construction method, system, equipment and medium

By using an automated method based on identifiers and large language models to parse and standardize scientific literature, combined with manual review, the problem of low efficiency and high error rate in the construction of scientific knowledge bases in existing technologies has been solved, and cross-domain adaptability and data quality have been improved.

CN120930742APending Publication Date: 2025-11-11TIANJIN INST OF IND BIOTECH CHINESE ACADEMY OF SCI
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511028741.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing methods for constructing scientific knowledge bases are inefficient, rely on manual reading of literature, lack automation, have poor cross-domain adaptability, pose a high risk of error propagation, and are insufficient in standardization, making it difficult to cope with massive amounts of literature.

Method used

The system retrieves full-text documents based on a list of input identifiers or domain search terms, parses and converts them into structured text using a large language model, extracts key information using domain prompts, filters and standardizes the text using automated scripts, inserts it into a knowledge base and verifies the association logic, and human experts participate in key nodes to ensure quality.

Benefits of technology

It enables rapid cross-domain adaptation, improves knowledge extraction speed and data quality, reduces error rate, and supports efficient automated construction in fields such as biomedicine, materials science, and synthetic biology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930742A_ABST
    Figure CN120930742A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of natural language processing and knowledge engineering, and discloses a method, a system, equipment and a medium for automatically constructing a multi-field scientific knowledge base, and the method comprises the following steps: based on an input identifier list or a field search word, retrieving a full text of a literature, analyzing and converting the full text into a structured text; combining a pre-configured large language model with a field cue word, extracting key information in the structured text, and outputting structured data; the automatic script filters irrelevant, repeated or incomplete structured data according to a pre-configured filtering rule; performing standardization processing on the filtered structured data so as to realize the consistency and comparability of the data; and inserting the standardized structured data into an interactive knowledge base or database, establishing data association, and verifying association logic. The method supports cross-field rapid adaptation, solves the problems of low efficiency and high error rate of a traditional method, and can be widely applied to the fields of biomedicine, material science, synthetic biology and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of natural language processing and knowledge engineering, and in particular to a method, system, device, and medium for the automated construction of multi-domain scientific knowledge bases. Background Technology

[0002] In current technologies, the construction of scientific knowledge bases mainly relies on manual reading of literature and manual data entry, which is inefficient and struggles to handle massive amounts of literature. Some automated methods (such as keyword-based text mining) have the following drawbacks:

[0003] 1. Relying on abstract data while ignoring details in the full text (such as tables and experimental conditions);

[0004] 2. Poor cross-domain adaptability, requiring rule redesign;

[0005] 3. Lack of manual closed-loop verification leads to a high risk of error propagation;

[0006] 4. Insufficient standardization makes data difficult to reuse.

[0007] Therefore, how to provide an automated construction method, system, device, and medium for multi-domain scientific knowledge bases is an urgent problem to be solved. Summary of the Invention

[0008] This invention provides a method, system, device, and medium for the automated construction of a multi-domain scientific knowledge base to solve the aforementioned technical problems in the prior art.

[0009] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or to describe the scope of protection of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.

[0010] According to a first aspect of the present invention, an automated method for constructing a multi-domain scientific knowledge base is provided.

[0011] In one embodiment, the automated construction method for the multi-domain scientific knowledge base includes:

[0012] Based on the input list of identifiers or domain search terms, retrieve the full text of the literature, parse it and convert it into structured text;

[0013] A pre-configured large language model, combined with domain cue words, extracts key information from the structured text and outputs structured data.

[0014] The automated script filters irrelevant, duplicate, or incomplete structured data according to pre-configured filtering rules;

[0015] The filtered structured data is then standardized to ensure data consistency and comparability.

[0016] The standardized structured data is inserted into an interactive knowledge base or database to establish data relationships and verify the relationship logic.

[0017] In one embodiment, retrieving the full text of a document based on a list of input identifiers or domain search terms, parsing and converting it into structured text includes:

[0018] Based on a user-provided list of identifiers or domain search terms, retrieve identifier information and metadata of relevant literature from academic databases;

[0019] Full text is downloaded based on identifiers, the full text is parsed using a large language model, and the metadata of the document is completed.

[0020] Use parsing tools to convert documents into structured text while preserving table and chart formats.

[0021] In one embodiment, the step of retrieving the full text of a document based on a list of input identifiers or domain search terms, parsing and converting it into structured text, further includes:

[0022] Verify the parsing results and check whether the text has been completely extracted.

[0023] In one embodiment, the pre-configured large language model, combined with domain cue words, extracts key information from the structured text and outputs structured data including:

[0024] Design structured domain-specific keywords;

[0025] Break the task down into subtasks;

[0026] Check the extraction results, correct the model, and verify the consistency between the model output and the original text.

[0027] In one embodiment, the filtering rules include: experiment name data, experiment condition data, and performance index data.

[0028] In one embodiment, the standardization process for the filtered structured data to achieve data consistency and comparability includes:

[0029] Automated scripts are used to unify the data format of filtered structured data with that of external databases.

[0030] In one embodiment, inserting the standardized structured data into an interactive knowledge base or database, establishing data associations, and verifying the association logic includes:

[0031] Use automated scripts to import structured data into a knowledge base or database;

[0032] Establish connections between data and form an interconnected knowledge framework.

[0033] According to a second aspect of the present invention, an automated construction system for a multi-domain scientific knowledge base is provided.

[0034] In one embodiment, the automated construction system for the multi-domain scientific knowledge base includes:

[0035] The document parsing module is used to retrieve the full text of documents based on a list of input identifiers or domain search terms, and then parse and convert it into structured text.

[0036] The data extraction module is used to extract key information from the structured text by combining a pre-configured large language model with domain cue words, and output structured data.

[0037] The data filtering module is used to automate scripts to filter irrelevant, duplicate, or incomplete structured data according to pre-configured filtering rules.

[0038] The data standardization module is used to standardize the filtered structured data to achieve data consistency and comparability.

[0039] The data insertion module is used to insert standardized structured data into an interactive knowledge base or database, establish data relationships, and verify the relationship logic.

[0040] According to a third aspect of the present invention, a computer device is provided.

[0041] In some embodiments, the computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described above.

[0042] According to a fourth aspect of the present invention, a computer-readable storage medium is provided.

[0043] In one embodiment, a computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the above method.

[0044] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0045] 1) This invention extracts full-text data through a document parsing module, generates structured information using LLM semantic understanding, and imports it into a knowledge base after rule filtering and standardization. Human experts define rules and perform quality checks at key nodes, forming a closed loop of "automated processing - human optimization." This system supports rapid cross-domain adaptation, solving the problems of low efficiency and high error rate of traditional methods, and can be widely applied in fields such as biomedicine, materials science, and synthetic biology.

[0046] 2) Flexible adaptation: By adjusting search terms, suggestion words, filtering rules and standardization criteria, this process can be adapted to any scientific field.

[0047] 3) High-efficiency collaboration: LLM processes massive amounts of text, with program code implementing an automated pipeline, while human staff focus on high-level review and knowledge integration, significantly improving the speed of knowledge extraction.

[0048] 4) Continuous Optimization (Human-in-the-Loop, HITL): Human involvement is present throughout the entire process, from search term optimization to data validation, ensuring the reliability and usability of the results. The large language model provides initial extraction capabilities, the program code automates the processing, and human correction is the ultimate guarantee of quality; all three are indispensable.

[0049] 5) Closed-loop mechanism: Collect user query logs and review feedback to iteratively optimize LLM prompts and filtering rules. Regularly update the standardized rule base (e.g., add materials science terminology).

[0050] 6) Human role: Domain experts lead rule iteration to ensure that the process adapts to the forefront of the discipline.

[0051] 7) Accurate and reliable: Manual review is conducted throughout the entire process to ensure data quality; LLM is linked with authoritative databases to reduce the spread of errors.

[0052] 8) Scalability: The modular design supports seamless integration with new tools (such as new PDF parsing libraries) and adapts to technological evolution.

[0053] 9) Comprehensiveness: Based on full-text extraction rather than being limited to the summary, it captures richer data details.

[0054] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0055] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0056] Figure 1This is a flowchart illustrating an automated construction method for a multi-domain scientific knowledge base according to an exemplary embodiment;

[0057] Figure 2 This is a structural block diagram illustrating an automated construction system for a multi-domain scientific knowledge base, according to an exemplary embodiment.

[0058] Figure 3 This is a schematic diagram of the structure of a computer device according to an exemplary embodiment. Detailed Implementation

[0059] The following description and accompanying drawings fully illustrate specific embodiments described herein to enable those skilled in the art to practice them. Some embodiments may include or substitute parts and features of other embodiments. The scope of the embodiments herein encompasses the entire scope of the claims and all available equivalents thereof. Throughout this document, the terms “first,” “second,” etc., are used only to distinguish one element from another without requiring or implying any actual relationship or order between the elements. Indeed, a first element can also be referred to as a second element, and vice versa. Furthermore, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a structure, apparatus, or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a structure, apparatus, or device. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the structure, apparatus, or device that includes said element. The various embodiments described herein are presented in a progressive manner, with each embodiment focusing on its differences from other embodiments; similar or identical parts between embodiments can be referred to interchangeably.

[0060] The terms "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer" used in this document to indicate orientations or positional relationships are based on the orientations or positional relationships shown in the accompanying drawings. They are used solely for the convenience of describing the document and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In the description herein, unless otherwise specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly. For example, they can refer to mechanical or electrical connections, or internal connections between two elements; they can be direct connections or indirect connections through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.

[0061] In this document, unless otherwise stated, the term "multiple" means two or more.

[0062] In this article, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.

[0063] In this article, the term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0064] It should be understood that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order constraint on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the diagram may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0065] The modules in the apparatus or system of this application can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0066] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0067] Figure 1 An embodiment of an automated construction method for a multi-domain scientific knowledge base according to the present invention is shown.

[0068] In this optional embodiment, the method for automatically constructing a multi-domain scientific knowledge base includes:

[0069] Step S101: Based on the input identifier list or domain search terms, retrieve the full text of the document, parse it and convert it into structured text;

[0070] The step of retrieving full-text documents based on a list of input identifiers or domain search terms, parsing and converting them into structured text includes:

[0071] Based on a user-provided list of identifiers or domain search terms, retrieve identifier information and metadata of relevant literature from academic databases;

[0072] Full text is downloaded based on identifiers, the full text is parsed using a large language model, and the metadata of the document is completed.

[0073] Use parsing tools to convert documents into structured text while preserving table and chart formats;

[0074] Verify the parsing results and check whether the text has been completely extracted.

[0075] Specifically, the literature analysis process includes:

[0076] Input: A list of DOIs (identifiers) provided by the user or domain-specific search terms (such as "metabolic engineering" in biology, "clinical trials" in medicine, and "materials synthesis" in materials science).

[0077] Process: Based on the DOI or search terms, retrieve the DOI information and metadata (i.e., metadata, which some documents do not have) of relevant documents from academic databases (such as Web of Science and PubMed); download the full text based on the DOI, and use a large model to parse the full text and complete the metadata of the document; use a PDF parsing tool (such as Docling) to convert the PDF file of the document into structured text, retaining the format of tables, charts, etc., to ensure the integrity of the content.

[0078] Using a large model to parse the full text, the supplementary metadata of the document includes:

[0079] Text parsing: The PDF file is parsed using the Docling tool to extract structured text. Docling efficiently preserves text formatting and accurately extracts table and chart content. The parsed text is output in Markdown format to ensure seamless processing by subsequent data extraction modules.

[0080] Metadata completion: First, use efetch to query the abstract; if it's not found there, then use semantics. Use EntrezAPI to retrieve country, keywords, and institution information. Other information can be obtained from the Web of Science starter API.

[0081] Tool selection: Choose the appropriate parsing tool based on the needs of your domain. For example, Docling excels in preserving formatting and extracting tables, making it suitable for scenarios requiring precise data.

[0082] Human involvement: Experts optimize search terms based on the characteristics of the field to ensure that the retrieved literature is highly relevant to the research objectives. The parsing results are verified, and the completeness of the text extraction is checked, especially for tables and key data sections.

[0083] Output: Structured full-text data, laying the foundation for subsequent extraction.

[0084] In this embodiment, literature parsing aims to obtain raw text data related to metabolic engineering from scientific literature, laying the foundation for subsequent processing, including:

[0085] Literature search: Literature was searched through the Web of Science database using the search terms "((TS=("Escherichia coli")OR TS=("E.coli"))AND(TS=("metabolic engineering")OR TS=("genetic modification")))" to target metabolic engineering research related to target microorganisms such as Escherichia coli.

[0086] Text parsing: The PDF file is parsed using the Docling tool to extract structured text. Docling efficiently preserves text formatting and accurately extracts table and chart content. The parsed text is output in Markdown format to ensure seamless processing by subsequent data extraction modules.

[0087] Step S102: The pre-configured large language model, combined with domain cue words, extracts key information from the structured text and outputs structured data;

[0088] The pre-configured large language model, combined with domain-specific cue words, extracts key information from the structured text and outputs structured data, including:

[0089] Design structured domain-specific keywords;

[0090] Break the task down into subtasks;

[0091] Check the extraction results, correct the model, and verify the consistency between the model output and the original text.

[0092] Specifically, the data extraction process includes:

[0093] Input: The parsed document text.

[0094] Workflow: Use large language models (such as GPT-4) combined with domain-specific prompts to extract key information. For example: biology (extracting gene function, protein-protein interactions), medicine (extracting drug effects, clinical trial data), materials science (extracting material properties, experimental conditions); break down complex tasks into sub-tasks (such as extracting experimental conditions and results separately) to improve extraction accuracy; focus on key parts of the literature (such as methods, results, and tables) to reduce interference from irrelevant content.

[0095] Optimization strategy: Design structured prompts and output in a format (such as JSON) that is easy for the program to process; retain original text references for easy subsequent verification.

[0096] Human intervention: Experts design and adjust prompts based on domain needs to ensure that the extracted information is accurate and meets expectations; check the extraction results and correct any omissions or misunderstandings that may occur in the model; verify the consistency between the LLM output and the original text (e.g., compare the original table with the extraction results); and correct LLM errors (e.g., misinterpreted abbreviations, ambiguous terms).

[0097] Output: Structured raw extracted data containing domain-specific key information.

[0098] In this embodiment, data extraction utilizes a large language model (LLM) to extract key information on metabolic engineering modifications from the parsed literature text, including:

[0099] Model configuration: GPT-4 and other large language models were selected, and domain-specific prompts were designed to extract three types of core information: gene modification (such as gene name and modification type), fermentation conditions (such as culture medium and fermentation mode), and performance indicators (such as titer, yield, and rate).

[0100] Extraction Strategy: To improve efficiency, only the Methods, Results, and tables sections of the literature are processed, excluding non-critical content such as the Introduction and Discussion sections. A dedicated extraction process is designed for tabular data to ensure the capture of its structured information.

[0101] The designed extraction process includes: for tabular data, extracting the content before and after the table, such as titles, notes, and paragraph descriptions, to supplement the semantic information of the tabular data. Using the table and its context as text input, extracting prompt words using the gene modification and performance indicators designed above, and extracting structured information.

[0102] Implementation details: For example, when extracting gene modification information, prompts require the model to identify the gene name (e.g., ilvB), modification type (e.g., overexpression), and specific method (e.g., CRISPR editing) to generate detailed structured data.

[0103] Step S103: The automated script filters irrelevant, duplicate, or incomplete structured data according to the pre-configured filtering rules.

[0104] The filtering rules include: experiment name data, experiment condition data, and performance index data.

[0105] Specifically, data filtering includes:

[0106] Input: The extracted raw data.

[0107] Process: Use automated scripts to filter irrelevant, duplicate, or incomplete data based on domain rules, such as biology (removing records lacking genetic identifiers), medicine (removing entries without key clinical indicators), and materials science (filtering out experimental data without specific attributes).

[0108] The automated scripts efficiently process data based on domain rules, including irrelevant data removal, duplicate data deduplication, and incomplete data filtering. By configuring keyword attributes and constraint rules, it verifies data integrity and validity; for data that does not meet requirements, it can choose to mark, filter, or repair it. The scripts adopt a modular design, supporting dynamic configuration and high-performance processing, ensuring the accuracy and traceability of data cleaning, while outputting standardized data and anomaly reports. For example, in gene modification data, if the keywords "gene name" and "gene modification type" fields are empty or None, they are filtered; if all keyword combinations uniquely identify duplicates, they are also filtered.

[0109] Human involvement: Experts define filtering rules to ensure that the retained data is consistent with the research objectives; the filtering results are checked regularly and the rules are adjusted to improve accuracy.

[0110] Output: A cleaned, high-quality dataset.

[0111] In this embodiment, data filtering cleans the extracted data using automated scripts to ensure its quality and relevance, including filtering rules:

[0112] Genetic modification data (i.e., experiment name data): Remove records that are missing both the "Gene" and "Type" fields, and merge duplicate entries with the same "Gene" and "Type".

[0113] Fermentation condition data (i.e. experimental condition data): Delete records that are missing key fields such as "Media" or "Fermentation_Mode", and merge duplicate entries with identical fields.

[0114] Performance metrics data: Records with no specific values ​​for "Titer", "Yield", and "Rate" were excluded.

[0115] Step S104: Standardize the filtered structured data to achieve data consistency and comparability;

[0116] The standardization process for the filtered structured data to achieve data consistency and comparability includes:

[0117] Automated scripts are used to unify the data format of filtered structured data with that of external databases.

[0118] Specifically, data standardization includes:

[0119] Input: Filtered data.

[0120] Process: Standardize data according to domain conventions to ensure consistency and comparability, such as biology (unifying gene names and calibrating using UniProtAPI), medicine (standardizing drug names and units, such as dosage units mg), and materials science (unifying units of physical properties, such as Pa and nm); combine LLMs and external databases (such as PubChem) to handle synonyms and aliases.

[0121] Human intervention: Experts develop standardized rules to resolve complex or ambiguous naming issues; they manually review standardized results, especially those involving ambiguity or outliers; they review entries that cannot be automatically mapped (such as newly discovered compounds) and manually add standardized rules; and they verify the accuracy of cross-database associations (such as gene-disease relationships).

[0122] Output: A standardized dataset that facilitates comparison and analysis across studies.

[0123] In this embodiment, data standardization unifies the data format through scripts and external databases, ensuring consistency and comparability of data from different literature sources, including:

[0124] Gene normalization: Decompose complex gene sets (such as ilvBNCD) into individual genes (such as ilvB, ilvN, ilvC, ilvD); call UniProtAPI to map gene names to standard UniProt identifiers.

[0125] Compound standardization: Construct a compound alias mapping library, for example, unify "Acetic acid" and "Acetate" as standard names; use PubChemAPI to obtain the IUPAC name of the compound and its synonyms to further ensure consistency.

[0126] Performance metrics standardization: Titer is standardized to g / L, yield to g / g, and rate to g / L / h or mmol / g / h. For example, titer data in mg / L is converted to g / L.

[0127] Step S105: Insert the standardized structured data into an interactive knowledge base or database, establish data associations, and verify the association logic.

[0128] The step of inserting standardized structured data into an interactive knowledge base or database, establishing data associations, and verifying the association logic includes:

[0129] Use automated scripts to import structured data into a knowledge base or database;

[0130] Establish connections between data and form an interconnected knowledge framework.

[0131] Specifically, data insertion includes:

[0132] Input: Standardized data.

[0133] Process: Use automated scripts to import data into a domain-specific knowledge base or database; establish connections between data (such as links between experimental conditions and results) to form an interconnected knowledge framework; develop user-friendly interfaces (such as web platforms) to facilitate querying and exploration.

[0134] Human involvement includes: defining the knowledge base architecture (e.g., categorizing it by "synthesis method-performance-application" in materials science); optimizing the search logic and data display format based on user feedback; having experts verify the accuracy and relevance of imported data to ensure the scientific rigor of the knowledge base; and continuously maintaining and updating the knowledge base in response to user feedback.

[0135] Output: A fully functional, searchable domain knowledge base.

[0136] In this embodiment, the data insertion module integrates the standardized data into the SEEK platform and provides user interaction functions, including:

[0137] Data Linking: Automated scripts link genetic modification, fermentation conditions, and performance metrics data according to predefined rules. For example, genetic modification information of the same strain can be linked with its corresponding fermentation conditions and performance metrics to form a complete data chain.

[0138] Key attributes for performance index data include: culture medium name, modified gene name, product concentration, etc. Key attributes for fermentation condition data include: culture medium name, culture medium composition, etc. Key attributes for gene modification data include: gene name, mutation type, mutation site, etc. More precise data links are established through joint matching of multiple keywords. For example, the gene name in performance index data can be linked to the gene name in gene modification data, and the culture medium name in performance index data can be associated with the culture medium name in fermentation condition data, thus achieving a link between fermentation conditions, performance indicators, and gene modification.

[0139] Knowledge base construction: Import the associated data into the database to build an interconnected knowledge framework.

[0140] User Interface: Develop a web interface that allows users to search by product name or gene name to find details of gene modification, fermentation conditions, historical yield data, and relevant literature citations.

[0141] Figure 2 An embodiment of an automated construction system for a multi-domain scientific knowledge base according to the present invention is shown.

[0142] In this optional embodiment, the automated construction system for the multi-domain scientific knowledge base includes:

[0143] The document parsing module 201 is used to retrieve the full text of documents based on a list of input identifiers or domain search terms, and then parse and convert it into structured text.

[0144] The data extraction module 202 is used to extract key information from the structured text by combining a pre-configured large language model with domain cue words, and output structured data.

[0145] The data filtering module 203 is used to automate the filtering of irrelevant, duplicate, or incomplete structured data by scripts according to pre-configured filtering rules.

[0146] The data standardization module 204 is used to standardize the filtered structured data to achieve data consistency and comparability.

[0147] The data insertion module 205 is used to insert standardized structured data into an interactive knowledge base or database, establish data associations, and verify the association logic.

[0148] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores static and dynamic information data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0149] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0150] In addition, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0151] In addition, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0152] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0153] This invention is not limited to the structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this invention is limited only by the appended claims.

Claims

1. A method for automatically constructing a multi-domain scientific knowledge base, characterized in that, include: Based on the input list of identifiers or domain search terms, retrieve the full text of the literature, parse it and convert it into structured text; A pre-configured large language model, combined with domain cue words, extracts key information from the structured text and outputs structured data. The automated script filters irrelevant, duplicate, or incomplete structured data according to pre-configured filtering rules; The filtered structured data is then standardized to ensure data consistency and comparability. The standardized structured data is inserted into an interactive knowledge base or database to establish data relationships and verify the relationship logic.

2. The method for automatically constructing a multi-domain scientific knowledge base according to claim 1, characterized in that, The process of retrieving full-text documents based on a list of input identifiers or domain search terms, parsing and converting them into structured text includes: Based on a user-provided list of identifiers or domain search terms, retrieve identifier information and metadata of relevant literature from academic databases; Full text is downloaded based on identifiers, the full text is parsed using a large language model, and the metadata of the document is completed. Use parsing tools to convert documents into structured text while preserving table and chart formats.

3. The method for automatically constructing a multi-domain scientific knowledge base according to claim 1, characterized in that, The process of retrieving full-text documents based on a list of input identifiers or domain search terms, parsing and converting them into structured text, also includes: Verify the parsing results and check whether the text has been completely extracted.

4. The method for automatically constructing a multi-domain scientific knowledge base according to claim 1, characterized in that, The pre-configured large language model, combined with domain-specific cue words, extracts key information from the structured text and outputs structured data including: Design structured domain-specific keywords; Break the task down into subtasks; Check the extraction results, correct the model, and verify the consistency between the model output and the original text.

5. The method for automatically constructing a multi-domain scientific knowledge base according to claim 1, characterized in that, The filtering rules include: experiment name data, experiment condition data, and performance index data.

6. The method for automatically constructing a multi-domain scientific knowledge base according to claim 1, characterized in that, The standardization process for the filtered structured data to achieve data consistency and comparability includes: Automated scripts are used to unify the data format of filtered structured data with that of external databases.

7. The method for automatically constructing a multi-domain scientific knowledge base according to claim 1, characterized in that, The process of inserting standardized structured data into an interactive knowledge base or database, establishing data relationships, and verifying the relationship logic includes: Use automated scripts to import structured data into a knowledge base or database; Establish connections between data and form an interconnected knowledge framework.

8. An automated construction system for a multi-domain scientific knowledge base, characterized in that, include: The document parsing module is used to retrieve the full text of documents based on a list of input identifiers or domain search terms, and then parse and convert it into structured text. The data extraction module is used to extract key information from the structured text by combining a pre-configured large language model with domain cue words, and output structured data. The data filtering module is used to automate scripts to filter irrelevant, duplicate, or incomplete structured data according to pre-configured filtering rules. The data standardization module is used to standardize the filtered structured data to achieve data consistency and comparability. The data insertion module is used to insert standardized structured data into an interactive knowledge base or database, establish data relationships, and verify the relationship logic.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • High-quality metal material process data set construction method based on large language model

    CN121789818A

  • High-quality metal material process dataset construction method based on large language model

    CN121789818B

  • Biomedical literature scientific problem extraction system and extraction method

    CN122019654A