An evidence-based method and system for knowledge discovery from experimental data in battery materials
By performing structural deconstruction, cross-documentary evidence alignment, and explicit-implicit consistency analysis on battery material experimental data, the problem of the lack of traceability and alignment of battery material experimental data was solved, and reliable comparison and knowledge discovery of battery material experimental data were achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI
- Filing Date
- 2026-01-16
- Publication Date
- 2026-07-17
AI Technical Summary
The lack of traceable and aligned evidence organization mechanisms in experimental data on battery materials makes it difficult to reliably compare experimental conclusions across different literatures.
By acquiring a collection of experimental papers on battery materials, performing structural deconstruction processing to generate a collection of papers in Markdown format, constructing a battery material experimental library, and performing cross-document evidence alignment with materials as the alignment core, combined with explicit and implicit consistency evidence analysis, a multi-dimensional abstract mechanism is introduced to perform dual-channel source tracing abstract parsing.
This approach enables reliable comparison and knowledge discovery of experimental data for battery materials while maintaining the integrity of original literature evidence, thus solving the problem of reliable comparison of experimental conclusions across different literatures.
Smart Images

Figure CN121960490B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to an evidence-based method and system for discovering knowledge from experimental data on battery materials. Background Technology
[0002] In the field of battery materials research, a large amount of experimental data and performance conclusions are scattered across different research papers and experimental reports. This information is often presented in unstructured text or simple tables. Experimental results are often given only as conclusive descriptions, lacking a clear connection to specific experimental conditions, procedures, and fragments of original literature, making it difficult to effectively trace the source of experimental data. Furthermore, experimental conclusions for the same battery material in different publications vary in testing conditions, characterization methods, and data presentation, lacking a unified alignment and organization centered on the material. This makes it difficult to accurately correlate and compare experimental conclusions, affecting the reliability of our understanding of battery material performance. Summary of the Invention
[0003] This application provides an evidence-based knowledge discovery method and system for battery material experimental data, which addresses the technical problem that the lack of traceable and alignable evidence organization mechanisms in existing battery material experimental data makes it difficult to reliably compare experimental conclusions across different literatures.
[0004] In view of the above problems, this application provides an evidence-based method and system for discovering knowledge from experimental data of battery materials.
[0005] The first aspect of this application provides an evidence-based method for knowledge discovery from experimental data on battery materials, the method comprising:
[0006] A collection of experimental papers on battery materials is obtained, and structural deconstruction processing is performed to obtain a collection of papers in Markdown format. Semantic parsing is then performed on this Markdown collection to construct a battery material experimental library. Using materials as the alignment core, cross-document evidence alignment is performed on the battery material experimental library to obtain multiple experimental pieces of evidence. Explicit and implicit consistency evidence analysis is then performed on these multiple experimental pieces of evidence to obtain multiple consistent experimental pieces of evidence and multiple differing experimental pieces of evidence. A multi-dimensional summarization mechanism is introduced to perform dual-channel source tracing and summarization on the multiple consistent experimental pieces of evidence and the multiple differing experimental pieces of evidence, obtaining multiple rapid semantic summaries of experiments and multiple fine-grained evidence summaries. These multiple rapid semantic summaries of experiments and multiple fine-grained evidence summaries are used as knowledge discovery results.
[0007] A second aspect of this application provides an evidence-based knowledge discovery system for experimental data on battery materials, the system comprising:
[0008] The system comprises the following modules: a deconstruction module, which acquires a collection of experimental papers on battery materials, performs structural deconstruction processing, and obtains a collection of papers in Markdown format; a semantic parsing module, which traverses the Markdown format paper collection to perform semantic parsing and construct a battery material experimental library; an evidence alignment module, which uses materials as the alignment core and traverses the battery material experimental library to perform cross-document evidence alignment and obtain multiple experimental pieces of evidence; a consistency evidence analysis module, which traverses the multiple experimental pieces of evidence to perform explicit and implicit consistency evidence analysis and obtain multiple consistent experimental pieces of evidence and multiple differing experimental pieces of evidence; and a summary parsing module, which introduces a multi-dimensional summarization mechanism to perform dual-channel source-tracing summary parsing on the multiple consistent experimental pieces of evidence and the multiple differing experimental pieces of evidence, obtaining multiple rapid semantic summaries of experiments and multiple fine-grained evidence summaries, and using these multiple rapid semantic summaries of experiments and multiple fine-grained evidence summaries as knowledge discovery results.
[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0010] This application obtains a collection of experimental papers on battery materials, performs structured deconstruction processing to obtain a collection of papers in Markdown format; it then traverses the Markdown collection for semantic parsing to construct a battery material experimental library; using materials as the alignment core, it traverses the battery material experimental library for cross-document evidence alignment to obtain multiple experimental pieces of evidence; it then traverses these multiple pieces of experimental evidence for explicit and implicit consistency evidence analysis to obtain multiple consistent experimental pieces of evidence and multiple differing experimental pieces of evidence; finally, it introduces a multi-dimensional summarization mechanism to perform dual-channel source-tracing summary parsing on the multiple consistent experimental pieces of evidence and the multiple differing experimental pieces of evidence to obtain multiple rapid semantic summaries of experiments and multiple fine-grained evidence summaries, which are then used as knowledge discovery results. This invention addresses the technical problem in existing technologies where battery material experimental data lacks a traceable and alignable evidence organization mechanism, making reliable comparison of cross-document experimental conclusions difficult. By using material-centric cross-document evidence alignment combined with explicit and implicit consistency evidence analysis, it achieves the technical effect of reliable comparison and knowledge discovery of battery material experimental data while maintaining the integrity of the original documentary evidence. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1A schematic flowchart of an evidence-based battery material experimental data knowledge discovery method provided in this application embodiment;
[0013] Figure 2 This is a schematic diagram of an evidence-based battery material experimental data knowledge discovery system provided in an embodiment of this application.
[0014] Figure labeling: Deconstruction module 11, Semantic parsing module 12, Evidence alignment module 13, Consistency evidence analysis module 14, Summary parsing module 15. Detailed Implementation
[0015] This application provides an evidence-based method and system for knowledge discovery from experimental data of battery materials. It addresses the technical problem in the prior art that the lack of traceable and alignable evidence organization mechanisms for experimental data of battery materials makes it difficult to reliably compare experimental conclusions across different documents. By aligning cross-documentary evidence with materials as the core and combining explicit and implicit consistency evidence analysis, it achieves the technical effect of reliable comparison and knowledge discovery of experimental data of battery materials while maintaining the integrity of the original documentary evidence.
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0017] It should be noted that any variation of the terms "comprising" and "having" is intended to cover non-exclusive inclusion, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such processes, methods, products, or devices.
[0018] Example 1, as Figure 1 As shown, this application provides an evidence-based method for knowledge discovery from experimental data of battery materials, the method comprising:
[0019] Step S100: Obtain the collection of experimental papers on battery materials, perform structural deconstruction processing, and obtain the collection of papers in Markdown format.
[0020] In this embodiment of the application, a collection of experimental papers on battery materials is first obtained through literature retrieval and data acquisition. This collection includes publicly published research papers related to the preparation, characterization, and performance testing of battery materials, all in PDF format as the original input medium. The sources of these papers may include academic journal databases, open access research platforms, or internal research document storage systems.
[0021] After obtaining the collection of experimental papers on battery materials, the PDF research papers underwent structured deconstruction processing, which involved parsing and reconstructing the original papers' layout structure and semantic hierarchy. Specifically, MinerU was first used to analyze the layout and break down the content into blocks, identifying chapter titles, main text paragraphs, figure and table areas, and the structure of references. Then, Grobid was used to analyze the academic structure of the split text, extracting the chapter hierarchy, paragraph order, and bibliographic metadata, and then standardizing and reorganizing the text content. Through this structured deconstruction process, the implicit chapter logic and semantic order in the original PDF papers were explicitly expressed, generating clearly structured, structurally stable, and semantically continuous Markdown format papers, thus obtaining a collection of Markdown format papers.
[0022] Furthermore, the method provided in the application embodiments also includes:
[0023] MinerU and Grobid were used to deconstruct each battery material experimental paper in the collection to obtain a collection of papers in Markdown format.
[0024] In this embodiment of the application, when using MinerU and Grobid to deconstruct each battery material experimental paper in the battery material experimental paper collection, MinerU is first used to analyze the layout structure of the PDF format paper. By identifying the page layout, text block position, paragraph boundary and figure area, the title, body text, figure caption and reference in the paper are initially separated.
[0025] Subsequently, Grobid was used to perform academic structure analysis on the text content processed by MinerU, standardizing and reconstructing the chapter hierarchy, title numbering, paragraph order, and bibliographic metadata of the papers, thus explicitly expressing the implicit logical structure in the original PDF. Through the collaborative deconstruction processing of MinerU and Grobid, the PDF format papers with complex layouts and weak semantic continuity were transformed into Markdown format text with clear hierarchy, stable structure, and semantic continuity, resulting in a collection of Markdown format papers.
[0026] Step S200: Traverse the collection of Markdown format papers, perform semantic parsing, and construct a battery material experimental library.
[0027] In this embodiment, when traversing the collection of Markdown format papers for semantic parsing to construct a battery material experimental library, the interface of the deepseek-reasoner large language model is first invoked under the constraint of preset prompt words to automatically identify and extract experiment-related content from the Markdown format papers, resulting in an information extraction result set containing material information, experimental procedures, test conditions, and performance data. Subsequently, according to predefined data structure specifications, the information extraction result set is stored in Elasticsearch in a unified structured form, thereby completing the construction of a battery material experimental library with unified semantic representation and searchability.
[0028] Furthermore, the method provided in the application embodiments, which involves traversing the collection of Markdown format papers for semantic parsing to construct a battery material experimental library, also includes:
[0029] The interface of the large language model deepseek-reasoner is called to automatically extract experimental information from the collection of Markdown format papers under the constraint of preset prompt words, and obtain an information extraction result set; according to the preset data structure, the information extraction result set is stored in Elasticsearch in a structured form to build the battery material experimental library.
[0030] Furthermore, the method provided in the application embodiments also includes:
[0031] The preset data structure includes entities and relationships between entities; among them, entities include materials, properties, experimental procedures, test conditions, experimental steps, and parameters, and the four types of entities, materials, properties, experimental procedures, and experimental steps, are associated with original text fragment information for traceability and evidence-based purposes; the relationships between entities include eight types of relationships, including the material to which the material belongs, the experimental procedure for preparing the material, a list of all tested properties, test conditions, experimental steps that make up the procedure, experimental parameters, input materials, and output materials.
[0032] In this embodiment, when performing semantic parsing on a collection of Markdown format papers, the interface of the large language model deepseek-reasoner is first invoked under the constraint of preset prompt words to automatically extract the experiment-related content from each Markdown format paper. The preset prompt words are used to limit the target and boundary of information extraction, so that deepseek-reasoner only parses the experimental facts objectively described in the paper and analyzes the content directly related to the battery material experiment in the main text of the paper paragraph by paragraph. In this way, information such as material name and chemical composition, experimental preparation and characterization process, experimental step sequence, performance testing process and test results are identified and extracted, forming an information extraction result set based on text-level semantic units, while preserving the contextual boundaries of each experimental information in the original Markdown text during the extraction process.
[0033] After obtaining the information extraction result set, the set is structured according to a preset data structure. The preset data structure includes entities and relationships between entities, with entities categorized into six types: materials, performance, experimental procedures, test conditions, experimental steps, and parameters. During the structuring process, the content describing material composition and names in the information extraction results is organized into material entities; the content describing performance index values, units, and test ranges is organized into performance entities; the content describing the complete experimental preparation or characterization process is organized into experimental procedure entities, which are further broken down into sequentially arranged experimental step entities; environmental conditions directly related to performance testing are organized into test condition entities; and process variables and control quantities involved in the experiment are organized into parameter entities. The four types of entities—materials, performance, experimental procedures, and experimental steps—are synchronously linked to their original text fragments in the Markdown document during generation, establishing a one-to-one correspondence between experimental data and the original paper content, thereby supporting subsequent traceability and evidence collection.
[0034] While completing the entity organization, relationships between entities are constructed based on the explicit description of the relationships between experimental facts in the paper. These relationships include the material relationships corresponding to materials, the experimental procedures for preparing materials, the list of all tested properties of materials, the test conditions between properties and their testing environment, the experimental steps that make up the experimental procedures, the experimental parameters between experimental procedures and their parameters, and the input and output material relationships involved in the experimental steps. This ensures that each experimental fact remains consistent with the original description in the paper within the structured representation.
[0035] After completing the above information extraction, entity construction, and determination of relationships between entities, the resulting entities and relationships between entities are stored in Elasticsearch in a unified structured format, thereby constructing a battery material experimental library with a unified data structure, supporting efficient retrieval, and enabling traceability and evidence-based verification.
[0036] Step S300: Using materials as the alignment core, traverse the battery material experimental library to perform cross-document evidence alignment and obtain multiple experimental evidences.
[0037] In this embodiment, the battery material experimental database stored in Elasticsearch is traversed record by record, with materials serving as the core object for cross-document alignment during the traversal. Specifically, when traversing each experimental record, the associated material entity information is read, and experimental records from different papers that point to the same material are identified and grouped based on attributes such as material name, standard name, and chemical formula. After completing the material-level grouping, the structured information corresponding to the material, such as experimental procedures, performance, test conditions, experimental steps, and parameters, is organized accordingly. This aligns the experimental facts of the same material in different documents under a unified data structure, while retaining the source paper information and original text fragments for each experimental record.
[0038] Through the above-mentioned process of traversing the battery material experimental library, with materials as the alignment core, experimental records scattered in different literatures are organized into multiple sources that can be distinguished, whose content can be compared, and which have traceability capabilities.
[0039] Step S400: Traverse the multiple experimental evidences to perform explicit-latent consistency evidence analysis, and obtain multiple consistent experimental evidences and multiple differing experimental evidences.
[0040] In this embodiment, when performing explicit and implicit consistency evidence analysis on multiple experimental pieces of evidence, explicit consistency evidence analysis is first performed on the multiple experimental pieces of evidence. Based on the performance indicators, test conditions, and experimental procedures of the structured representation in the experimental pieces of evidence, comparable content between different experimental pieces of evidence is compared, thereby obtaining multiple explicit consistent experimental pieces of evidence and multiple explicit differing experimental pieces of evidence. Then, source tracing and evidence verification are performed on the multiple explicit consistent experimental pieces of evidence to extract the corresponding set of related original text fragments. Next, implicit consistency evidence analysis is performed based on the set of related original text fragments to identify experimental pieces of evidence that are consistent or differing at the semantic level of the original text, obtaining multiple consistent experimental pieces of evidence and multiple implicit differing experimental pieces of evidence. Finally, the multiple explicit differing experimental pieces of evidence and the multiple implicit differing experimental pieces of evidence are summarized to form multiple differing experimental pieces of evidence.
[0041] Furthermore, the method provided in the application embodiments, which involves traversing the multiple experimental pieces of evidence to perform explicit-implicit consistency evidence analysis to obtain multiple consistent experimental pieces of evidence and multiple differing experimental pieces of evidence, also includes:
[0042] Explicit consistency evidence analysis is performed on the multiple experimental evidences to obtain multiple explicit consistent experimental evidences and multiple explicit dissimilar experimental evidences; source tracing and evidence-based analysis are performed on the multiple explicit consistent experimental evidences to obtain multiple sets of related original text fragment information; implicit consistency evidence analysis is performed on the multiple sets of related original text fragment information to obtain multiple consistent experimental evidences and multiple implicit dissimilar experimental evidences; the multiple explicit dissimilar experimental evidences and the multiple implicit dissimilar experimental evidences are summarized to obtain multiple dissimilar experimental evidences.
[0043] In this embodiment, when performing explicit consistency evidence analysis on multiple experimental pieces of evidence, the overall similarity between each piece of experimental evidence and the remaining experimental evidence is first calculated, and the experimental evidence corresponding to the maximum overall similarity is selected as the explicit consistency center. Then, using a preset minimum similarity as a constraint and according to a preset diffusion similarity bandwidth, a neighborhood is constructed around the explicit consistency center in the multiple experimental pieces of evidence, and progressive neighborhood edge diffusion is performed to obtain the neighborhood of the target explicit consistency center. Next, the experimental evidence in the neighborhood of the target explicit consistency center is determined as multiple explicit consistent experimental pieces of evidence. Finally, the experimental evidence other than the multiple explicit consistent experimental pieces of evidence is determined as multiple explicit difference experimental pieces of evidence.
[0044] Next, source tracing and evidence-based processing was performed on multiple pieces of explicit consistent experimental evidence. Specifically, for each piece of explicit consistent experimental evidence, the original text fragments associated with it in the battery material experimental library were retrieved, the original description position of the experimental evidence in the corresponding paper was located, and the original text content related to materials, performance, experimental procedures and experimental steps was extracted, thereby forming a set of associated original text fragments that correspond one-to-one with multiple pieces of explicit consistent experimental evidence.
[0045] Subsequently, implicit consistency evidence analysis was performed based on multiple sets of related original text fragments. This process first traversed the sets of related original text fragments, extracting semantic features for each fragment to obtain corresponding semantic features. Then, implicit consistency evidence analysis was performed on multiple semantic features, identifying features with consistent semantic expression as implicitly consistent semantic features and features with semantic differences as implicitly differing semantic features. Finally, the explicit consistency experimental evidence corresponding to multiple implicitly consistent semantic features was determined as multiple consistency experimental evidence, and the explicit consistency experimental evidence corresponding to multiple implicitly differing semantic features was determined as multiple implicitly differing experimental evidence.
[0046] Finally, multiple pieces of evidence from dominant and recessive experimental differences were combined to form a unified set of evidence from differential experiments.
[0047] Furthermore, the method provided in the application embodiments, which performs explicit consistency evidence analysis on the plurality of experimental evidences to obtain plurality of explicit consistency experimental evidences and plurality of explicit difference experimental evidences, further includes:
[0048] The experimental evidence corresponding to the maximum overall similarity with the remaining experimental evidence is extracted from the plurality of experimental evidence and used as the explicit consistency center. With a preset minimum similarity as a constraint, and according to a preset diffusion similarity bandwidth, a neighborhood is constructed and the neighborhood edge is gradually diffused among the plurality of experimental evidence to obtain the neighborhood of the target explicit consistency center. The experimental evidence in the neighborhood of the target explicit consistency center is used as a plurality of explicit consistent experimental evidence. The experimental evidence other than the plurality of explicit consistent experimental evidence is used as a plurality of explicit difference experimental evidence.
[0049] In this embodiment, when determining the explicit center of consistency from multiple experimental evidences, each piece of experimental evidence is traversed one by one, and the overall similarity is calculated using cosine similarity. Specifically, the directly comparable structured fields in each piece of experimental evidence are concatenated into a feature vector of the same dimension. The feature vector at least includes performance values and their unit conversions, test ratios, voltage windows, temperature, and other fields, and missing fields are filled with preset placeholder values. Subsequently, the cosine similarity is calculated for the feature vectors of any two pieces of experimental evidence to obtain a similarity set between that piece of experimental evidence and the other experimental evidences, and the maximum value is taken from the similarity set. After repeating the above process for all experimental evidences, the largest of the maximum values is selected, and the corresponding experimental evidence is taken as the experimental evidence in the experiment, and this experimental evidence in the experiment is determined as the explicit center of consistency.
[0050] Next, neighborhood construction and progressive neighborhood edge diffusion are performed on the dominant consistency center under the constraint of a preset minimum similarity and according to a preset diffusion similarity bandwidth. In this process, firstly, a neighborhood of the dominant consistency center is constructed based on multiple experimental evidences. Then, under the preset minimum similarity constraint, the neighborhood edges of the dominant consistency center neighborhood are diffused according to the preset diffusion similarity bandwidth to obtain the diffused dominant consistency center neighborhood. During this process, the neighborhood size of the diffused dominant consistency center neighborhood is judged. When its neighborhood size is greater than or equal to the neighborhood size of the dominant consistency center neighborhood, neighborhood edge diffusion continues until the preset number of diffusions is met and / or the constraint conditions are no longer met. Finally, the diffused dominant consistency center neighborhood obtained from the last diffusion is determined as the target dominant consistency center neighborhood.
[0051] Then, the experimental evidence in the neighborhood of the target explicit consistency center is taken as multiple explicit consistent experimental evidences. That is, all experimental evidence contained in the neighborhood of the target explicit consistency center is directly output and uniformly labeled as multiple explicit consistent experimental evidences.
[0052] Finally, the experimental evidence other than the multiple dominant consistency experimental evidences is regarded as multiple dominant difference experimental evidences. That is, the multiple experimental evidences are subjected to set difference processing, and all the remaining experimental evidences that do not belong to the neighborhood of the target dominant consistency center are extracted and marked as multiple dominant difference experimental evidences.
[0053] Furthermore, in the method provided in the application embodiments, constrained by a preset minimum similarity, and according to a preset diffusion similarity bandwidth, the method constructs a neighborhood and gradually diffuses the neighborhood edges of the explicit consistency center in the plurality of experimental evidence to obtain the neighborhood of the target explicit consistency center, and further includes:
[0054] Constrained by a preset minimum similarity, and according to a preset diffusion similarity bandwidth, a neighborhood of the dominant consistency center is constructed from the multiple experimental evidences. Again, constrained by the preset minimum similarity, the neighborhood edges of the neighborhood of the dominant consistency center are diffused according to the preset diffusion similarity bandwidth to obtain a diffused neighborhood of the dominant consistency center. It is determined whether the number of neighborhoods within the diffused neighborhood of the dominant consistency center is greater than or equal to the number of neighborhoods of the dominant consistency center. If so, the neighborhood edges of the diffused neighborhood of the dominant consistency center are continued to be diffused according to the preset minimum similarity and the preset diffusion similarity bandwidth until a preset number of diffusions is satisfied and / or the constraints are not satisfied. The diffused neighborhood of the dominant consistency center obtained from the last diffusion is taken as the target neighborhood of the dominant consistency center.
[0055] In this embodiment, when constructing the neighborhood of the dominant consistency center among multiple experimental evidences, each piece of experimental evidence is traversed one by one. Cosine similarity is used as the overall similarity calculation method. After concatenating the structured fields of each piece of experimental evidence into a feature vector, the cosine similarity between each piece of experimental evidence and the dominant consistency center is calculated. A preset minimum similarity is used as a screening threshold for preliminary filtering. At the same time, a preset diffusion similarity bandwidth is introduced as a constraint parameter for the similarity variation range. The preset diffusion similarity bandwidth is used to limit the similarity deviation range between the experimental evidence that can be included in the neighborhood and the dominant consistency center, thereby avoiding experimental evidence with excessively large similarity differences from being directly included in the neighborhood. Under the condition that the preset minimum similarity is met and the similarity falls within the range limited by the preset diffusion similarity bandwidth, the corresponding experimental evidence is added to the set to obtain the neighborhood of the dominant consistency center.
[0056] Next, when diffusing the neighborhood edges of the neighborhood of the dominant consistency center, the experimental evidence newly added to the neighborhood of the dominant consistency center in the previous round is identified as the neighborhood edges. The cosine similarity between the neighborhood edges and the experimental evidence not yet included in the neighborhood of the dominant consistency center is calculated one by one. The preset minimum similarity is used as the basic screening constraint again, and the preset diffusion similarity bandwidth is used as the similarity offset control condition. Experimental evidence that meets the preset diffusion similarity bandwidth range in terms of similarity value with the experimental evidence of the neighborhood edges is gradually added to the neighborhood, thus obtaining the diffused neighborhood of the dominant consistency center.
[0057] Subsequently, when judging the neighborhood quantity within the neighborhood of the diffusion dominant consistency center and the neighborhood quantity of the dominant consistency center, the quantity of experimental evidence within the neighborhood of the diffusion dominant consistency center is counted as the neighborhood quantity within the neighborhood of the diffusion dominant consistency center, and the quantity of experimental evidence within the neighborhood of the dominant consistency center is counted as the neighborhood quantity of the dominant consistency center. The neighborhood quantity judgment result is formed by comparing the size relationship between the two.
[0058] Finally, the diffusion continues on the neighborhood edges of the neighborhood of the diffusion dominant consistency center until termination, and the target dominant consistency center neighborhood is determined. During this process, if the neighborhood size determination result is greater than or equal to a certain value, the experimental evidence newly added to the diffusion dominant consistency center neighborhood in this round is updated as new neighborhood edges. The diffusion process based on a preset minimum similarity constraint and a preset diffusion similarity bandwidth is repeated, while the number of diffusions is counted. Diffusion terminates when the preset number of diffusions is reached and / or when there is no longer any new experimental evidence satisfying the preset minimum similarity and preset diffusion similarity bandwidth constraints. The diffusion dominant consistency center neighborhood obtained from the last diffusion is determined as the target dominant consistency center neighborhood.
[0059] Furthermore, the method provided in the application embodiments, which performs implicit consistency evidence analysis based on the multiple sets of related original text fragment information to obtain multiple consistency experimental evidence and multiple implicit difference experimental evidence, also includes:
[0060] Semantic features are extracted by traversing the multiple sets of related original text fragments to obtain multiple semantic features; implicit consistency evidence analysis is performed on the multiple semantic features to obtain multiple implicit consistent semantic features and multiple implicit dissimilar semantic features; explicit consistency experimental evidence corresponding to the multiple implicit consistent semantic features is used as multiple consistency experimental evidence; explicit consistency experimental evidence corresponding to the multiple implicit dissimilar semantic features is used as multiple implicit dissimilar experimental evidence.
[0061] In this embodiment, when extracting semantic features by traversing multiple sets of related original text fragments, the information is read one by one, and semantic features are generated using a keyword window extraction method. Specifically, in each original text fragment, anchor keywords are used, such as material name, performance name, experimental procedure name, and core actions of experimental steps. A preset length of context text is extracted before and after the anchor keywords to form a context window. Test condition limiting words, conditional descriptive words such as multiplier / voltage window / temperature, pre-processing or post-processing descriptive words, comparison words, and negation words appearing within the window are merged into the semantic features of that original text fragment. Thus, each original text fragment obtains a semantic feature that can be used for contextual comparison, thereby obtaining multiple semantic features.
[0062] Next, when performing implicit consistency evidence analysis on multiple semantic features, a term overlap rate calculation method is used for discrimination. Specifically, the semantic elements in multiple semantic features are deduplicated, and the semantic elements of different semantic features are compared pairwise. The overlap ratio of the semantic elements is calculated, and a preset overlap rate threshold is used as the discrimination criterion for implicit consistency. Semantic features with an overlap ratio not lower than the threshold are judged as having consistent semantic expression at the contextual level and are classified as implicitly consistent semantic features. Semantic features with an overlap ratio lower than the threshold are judged as having different semantic expression at the contextual level and are classified as implicitly differing semantic features. This yields multiple implicitly consistent semantic features and multiple implicitly differing semantic features, which are used to supplement the contextual loss problem caused by entity abstraction in explicit consistency analysis.
[0063] Subsequently, when multiple implicitly consistent semantic features are combined with explicit consistent experimental evidence as multiple consistent experimental evidences, an identifier-based back-reference method is used. Specifically, during the semantic feature generation process, a unique identifier of the explicit consistent experimental evidence from which each semantic feature originates is retained. After the implicit consistent evidence analysis is completed, the implicitly consistent semantic features are back-referenced one by one to the corresponding explicit consistent experimental evidence based on the unique identifier. The experimental evidence obtained from the back-reference is then deduplicated and organized to obtain multiple consistent experimental evidences.
[0064] When treating multiple implicit difference semantic features corresponding to explicit consistency experimental evidence as multiple implicit difference experimental evidence, the same label back-reference method is used. Specifically, based on the source labels retained by the implicit difference semantic features, the corresponding explicit consistency experimental evidence is back-referenced and organized one by one, so that experimental evidence that is consistent at the explicit structural level but differs at the original text context level can be distinguished, thus obtaining multiple implicit difference experimental evidence.
[0065] Step S500: Introduce a multi-dimensional summarization mechanism to perform dual-channel source tracing and summarization parsing on the multiple consistent experimental evidence and the multiple differing experimental evidence to obtain multiple experimental semantic fast summaries and multiple fine-grained evidence summaries. Use the multiple experimental semantic fast summaries and multiple fine-grained evidence summaries as knowledge discovery results.
[0066] In this embodiment, after introducing a multi-dimensional summarization mechanism, multiple consistent experimental evidences and multiple differing experimental evidences are processed through corresponding dual-channel source-tracing summary parsing processes. One channel is for consistent experimental evidences, which performs rapid summary analysis on multiple consistent experimental evidences to extract stable and consistent experimental semantic information across documents and generate multiple rapid experimental semantic summaries. The other channel is for differing experimental evidences, which extracts traceable original text fragments associated with each differing experimental evidence and performs source-tracing summary parsing to generate multiple fine-grained evidence summaries that accurately reflect the source of experimental differences. Finally, the multiple rapid experimental semantic summaries and multiple fine-grained evidence summaries are unified as the knowledge discovery result output.
[0067] Furthermore, the method provided in the application embodiment introduces a multi-dimensional summarization mechanism to perform dual-channel source tracing and summarization parsing on the multiple consistent experimental evidence and the multiple differing experimental evidence, obtaining multiple rapid semantic summaries of experiments and multiple fine-grained evidence summaries. The multiple rapid semantic summaries of experiments and the multiple fine-grained evidence summaries are used as knowledge discovery results, and the method further includes:
[0068] Based on the multiple consistent experimental evidence, a rapid summary analysis is performed to obtain multiple experimental semantic rapid summaries; the traceable related original text fragment information of each of the multiple differing experimental evidences is extracted to obtain multiple sets of differing related original text fragment information; based on the multiple sets of differing related original text fragment information, a traceable summary parsing is performed to obtain multiple fine-grained evidence summaries.
[0069] In this embodiment, when performing rapid summary analysis based on multiple consistent experimental evidences, a semantic fitting method is used to generate common experimental semantics from the consistent experimental evidences. Specifically, firstly, multiple consistent experimental evidences are traversed one by one, reading structured fields such as materials, performance, experimental procedures, and test conditions, while simultaneously reading the corresponding original text fragments. Then, using materials as a unified object, performance indicators with the same name, test conditions, and key points of the same experimental procedures in different experimental evidences are aligned. Numerical fields are converted and normalized according to a unified unit, and synonymous descriptions in textual fields are normalized and mapped. After completing the field-level alignment, the experimental semantic content that appears repeatedly in multiple consistent experimental evidences is merged, and modifying descriptions that appear only in individual documents but do not affect the overall conclusion are weakened. Thus, through semantic integration and compression, a summary expression that can represent the stable experimental conclusion of the material is formed. Finally, the fitted experimental semantic content is output as a rapid summary of multiple experimental semantics.
[0070] Next, when extracting the traceable original text fragments for each piece of differential experimental evidence, a text location extraction method is used. Specifically, for each piece of differential experimental evidence, based on the source paper identifier and original text fragment location identifier associated with it in the battery material experimental library, the original text content in the corresponding paper is directly located, and original text segments related to performance values, test condition limitations, experimental procedure descriptions, and conclusion statements are extracted to form original text fragment records that correspond one-to-one with each piece of differential experimental evidence, thereby obtaining multiple differentially associated original text fragment information.
[0071] Finally, when performing source tracing and summary analysis based on multiple discrepancies in the original text fragments, a source-preserving inductive approach is used to generate the summary. Specifically, each extracted original text fragment is read and categorized according to material name, performance name, and test conditions. Experimental conclusions that differ across different documents are presented in parallel, and the corresponding source paper identifiers and original text fragment content are clearly marked in the summary content. This ensures that each discrepancy description can be directly traced back to the original document, ultimately forming multiple fine-grained evidence summaries, which, together with the rapid experimental semantic summary, serve as the knowledge discovery result output.
[0072] In summary, the embodiments of this application have at least the following technical effects:
[0073] This application obtains a collection of experimental papers on battery materials, performs structured deconstruction processing to obtain a collection of papers in Markdown format; it then traverses the Markdown collection for semantic parsing to construct a battery material experimental library; using materials as the alignment core, it traverses the battery material experimental library for cross-document evidence alignment to obtain multiple experimental pieces of evidence; it then traverses these multiple pieces of experimental evidence for explicit and implicit consistency evidence analysis to obtain multiple consistent experimental pieces of evidence and multiple differing experimental pieces of evidence; finally, it introduces a multi-dimensional summarization mechanism to perform dual-channel source-tracing summary parsing on the multiple consistent experimental pieces of evidence and the multiple differing experimental pieces of evidence to obtain multiple rapid semantic summaries of experiments and multiple fine-grained evidence summaries, which are then used as knowledge discovery results. This invention addresses the technical problem in existing technologies where battery material experimental data lacks a traceable and alignable evidence organization mechanism, making reliable comparison of cross-document experimental conclusions difficult. By using material-centric cross-document evidence alignment combined with explicit and implicit consistency evidence analysis, it achieves the technical effect of reliable comparison and knowledge discovery of battery material experimental data while maintaining the integrity of the original documentary evidence.
[0074] Example 2, based on the same inventive concept as the evidence-based battery material experimental data knowledge discovery method in the foregoing examples, such as... Figure 2 As shown, this application provides an evidence-based knowledge discovery system for battery materials experimental data. The system and method embodiments in this application are based on the same inventive concept. The system includes:
[0075] The deconstruction module 11 is used to obtain a set of experimental papers on battery materials, perform structural deconstruction processing, and obtain a set of papers in Markdown format; the semantic parsing module 12 is used to traverse the set of papers in Markdown format for semantic parsing and construct a battery material experimental library; the evidence alignment module 13 is used to traverse the battery material experimental library with materials as the alignment core for cross-document evidence alignment and obtain multiple experimental evidences; the consistency evidence analysis module 14 is used to traverse the multiple experimental evidences for explicit and implicit consistency evidence analysis and obtain multiple consistent experimental evidences and multiple differing experimental evidences; the abstract parsing module 15 is used to introduce a multi-dimensional abstracting mechanism to perform dual-channel source tracing abstract parsing on the multiple consistent experimental evidences and multiple differing experimental evidences, obtain multiple experimental semantic fast abstracts and multiple fine-grained evidence abstracts, and use the multiple experimental semantic fast abstracts and multiple fine-grained evidence abstracts as knowledge discovery results.
[0076] Furthermore, the system is also used to implement the following functions:
[0077] The interface of the large language model deepseek-reasoner is called to automatically extract experimental information from the collection of Markdown format papers under the constraint of preset prompt words, and obtain an information extraction result set; according to the preset data structure, the information extraction result set is stored in Elasticsearch in a structured form to build the battery material experimental library.
[0078] Furthermore, the system is also used to implement the following functions:
[0079] The preset data structure includes entities and relationships between entities; among them, entities include materials, properties, experimental procedures, test conditions, experimental steps, and parameters, and the four types of entities, materials, properties, experimental procedures, and experimental steps, are associated with original text fragment information for traceability and evidence-based purposes; the relationships between entities include eight types of relationships, including the material to which the material belongs, the experimental procedure for preparing the material, a list of all tested properties, test conditions, experimental steps that make up the procedure, experimental parameters, input materials, and output materials.
[0080] Furthermore, the system is also used to implement the following functions:
[0081] Explicit consistency evidence analysis is performed on the multiple experimental evidences to obtain multiple explicit consistent experimental evidences and multiple explicit dissimilar experimental evidences; source tracing and evidence-based analysis are performed on the multiple explicit consistent experimental evidences to obtain multiple sets of related original text fragment information; implicit consistency evidence analysis is performed on the multiple sets of related original text fragment information to obtain multiple consistent experimental evidences and multiple implicit dissimilar experimental evidences; the multiple explicit dissimilar experimental evidences and the multiple implicit dissimilar experimental evidences are summarized to obtain multiple dissimilar experimental evidences.
[0082] Furthermore, the system is also used to implement the following functions:
[0083] The experimental evidence corresponding to the maximum overall similarity with the remaining experimental evidence is extracted from the plurality of experimental evidence and used as the explicit consistency center. With a preset minimum similarity as a constraint, and according to a preset diffusion similarity bandwidth, a neighborhood is constructed and the neighborhood edge is gradually diffused among the plurality of experimental evidence to obtain the neighborhood of the target explicit consistency center. The experimental evidence in the neighborhood of the target explicit consistency center is used as a plurality of explicit consistent experimental evidence. The experimental evidence other than the plurality of explicit consistent experimental evidence is used as a plurality of explicit difference experimental evidence.
[0084] Furthermore, the system is also used to implement the following functions:
[0085] Constrained by a preset minimum similarity, and according to a preset diffusion similarity bandwidth, a neighborhood of the dominant consistency center is constructed from the multiple experimental evidences. Again, constrained by the preset minimum similarity, the neighborhood edges of the neighborhood of the dominant consistency center are diffused according to the preset diffusion similarity bandwidth to obtain a diffused neighborhood of the dominant consistency center. It is determined whether the number of neighborhoods within the diffused neighborhood of the dominant consistency center is greater than or equal to the number of neighborhoods of the dominant consistency center. If so, the neighborhood edges of the diffused neighborhood of the dominant consistency center are continued to be diffused according to the preset minimum similarity and the preset diffusion similarity bandwidth until a preset number of diffusions is satisfied and / or the constraints are not satisfied. The diffused neighborhood of the dominant consistency center obtained from the last diffusion is taken as the target neighborhood of the dominant consistency center.
[0086] Furthermore, the system is also used to implement the following functions:
[0087] Semantic features are extracted by traversing the multiple sets of related original text fragments to obtain multiple semantic features; implicit consistency evidence analysis is performed on the multiple semantic features to obtain multiple implicit consistent semantic features and multiple implicit dissimilar semantic features; explicit consistency experimental evidence corresponding to the multiple implicit consistent semantic features is used as multiple consistency experimental evidence; explicit consistency experimental evidence corresponding to the multiple implicit dissimilar semantic features is used as multiple implicit dissimilar experimental evidence.
[0088] Furthermore, the system is also used to implement the following functions:
[0089] Based on the multiple consistent experimental evidence, a rapid summary analysis is performed to obtain multiple experimental semantic rapid summaries; the traceable related original text fragment information of each of the multiple differing experimental evidences is extracted to obtain multiple sets of differing related original text fragment information; based on the multiple sets of differing related original text fragment information, a traceable summary parsing is performed to obtain multiple fine-grained evidence summaries.
[0090] Furthermore, the system is also used to implement the following functions:
[0091] MinerU and Grobid were used to deconstruct each battery material experimental paper in the collection to obtain a collection of papers in Markdown format.
[0092] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0093] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. An evidence-based method for knowledge discovery from experimental data on battery materials, characterized in that, The method includes: Obtain a collection of experimental papers on battery materials, perform structural deconstruction processing, and obtain a collection of papers in Markdown format. The collection of Markdown format papers is traversed for semantic parsing to construct a battery material experimental library. Using materials as the core of alignment, cross-documentary evidence alignment was performed by traversing the battery material experimental library to obtain multiple experimental evidences. By traversing the multiple experimental pieces of evidence and performing explicit-latent consistency analysis, multiple consistent experimental pieces of evidence and multiple differing experimental pieces of evidence were obtained, including: A dominant consistency evidence analysis was performed on the multiple experimental evidences to obtain multiple dominant consistent experimental evidences and multiple dominant difference experimental evidences; By tracing the source of the multiple explicit and consistent experimental evidences, multiple sets of related original text fragment information were obtained; Implicit consistency evidence analysis was performed based on the aforementioned set of multiple related original text fragments to obtain multiple experimental evidences of consistency and multiple experimental evidences of implicit differences. The multiple evidences of dominant and recessive differences are combined to obtain multiple pieces of evidence of difference. Specifically, a dominant consistency evidence analysis was performed on the multiple experimental pieces of evidence to obtain multiple pieces of dominant consistent experimental evidence and multiple pieces of dominant difference experimental evidence, including: The experimental evidence corresponding to the maximum overall similarity with the remaining experimental evidence is extracted from the multiple experimental evidence and used as the explicit consistency center; With a preset minimum similarity as a constraint and according to a preset diffusion similarity bandwidth, the neighborhood of the explicit consistency center is constructed and the neighborhood edge is gradually diffused in the multiple experimental evidences to obtain the neighborhood of the target explicit consistency center. Experimental evidence in the neighborhood of the target explicit consistency center is used as multiple explicit consistency experimental evidences; Experimental evidence excluding the multiple explicit consistent experimental evidence is considered as multiple explicit differing experimental evidence. A multi-dimensional summarization mechanism is introduced to perform dual-channel source tracing and summarization parsing on the multiple consistent and multiple differing experimental evidence, obtaining multiple rapid semantic summaries and multiple fine-grained evidence summaries. These multiple rapid semantic summaries and multiple fine-grained evidence summaries are used as knowledge discovery results, including: Based on the aforementioned multiple consistent experimental evidence, a rapid summary analysis is performed to obtain a rapid summary of the semantics of multiple experiments. Extract the traceable related original text fragment information for each of the multiple differential experimental evidences to obtain multiple sets of differential related original text fragment information; Based on the multiple sets of differentially related original text fragment information, source tracing and summary parsing are performed to obtain multiple fine-grained evidence summaries.
2. The evidence-based knowledge discovery method for battery materials experimental data as described in claim 1, characterized in that, The collection of Markdown-formatted papers is traversed for semantic parsing to construct a battery material experimental library, including: By calling the interface of the large language model deepseek-reasoner, experimental information is automatically extracted from a collection of Markdown-formatted papers under the constraint of preset prompt words, and the information extraction result set is obtained. According to the preset data structure, the information extraction result set is stored in Elasticsearch in a structured form to build the battery material experimental library.
3. The evidence-based knowledge discovery method for battery materials experimental data as described in claim 2, characterized in that, The preset data structure includes entities and relationships between entities; The entities include materials, performance, experimental procedures, test conditions, experimental steps and parameters. The four types of entities, namely materials, performance, experimental procedures and experimental steps, are associated with original text fragments for traceability and evidence-based purposes. The relationships between entities include eight categories: the material to which it belongs, the experimental procedure for preparing the material, a list of all the properties tested, the test conditions, the experimental steps that make up the procedure, the experimental parameters, the input materials, and the output materials.
4. The evidence-based knowledge discovery method for battery materials experimental data as described in claim 1, characterized in that, Constrained by a preset minimum similarity, and according to a preset diffusion similarity bandwidth, the neighborhood of the target explicit consistency center is constructed and progressively diffused along its edges within the multiple experimental evidences to obtain the neighborhood of the target explicit consistency center, including: With a preset minimum similarity as a constraint, and according to a preset diffusion similarity bandwidth, a neighborhood of the explicit consistency center is constructed in the multiple experimental evidences. Again, constrained by a preset minimum similarity, the neighborhood edges of the neighborhood of the dominant consistency center are diffused according to the preset diffusion similarity bandwidth to obtain the diffused dominant consistency center neighborhood. Determine whether the number of neighborhoods within the neighborhood of the dominant consistency center is greater than or equal to the number of neighborhoods of the dominant consistency center. If so, continue to diffuse the neighborhood edges of the neighborhood of the dominant consistency center according to the preset diffusion similarity bandwidth, with the preset minimum similarity as a constraint, until the preset number of diffusions is met and / or the constraint is not met. The neighborhood of the dominant consistency center obtained by the last diffusion is taken as the target dominant consistency center neighborhood.
5. The evidence-based knowledge discovery method for battery materials experimental data as described in claim 1, characterized in that, Implicit consistency evidence analysis was performed based on the aforementioned sets of information from multiple related original text fragments, resulting in multiple experimental evidences of consistency and multiple experimental evidences of implicit differences, including: Semantic features are extracted by traversing the multiple sets of related original text fragments; Implicit consistency evidence analysis is performed on the multiple semantic features to obtain multiple implicit consistent semantic features and multiple implicit dissimilar semantic features; The explicit consistency experimental evidence corresponding to the multiple implicit consistency semantic features is used as multiple consistency experimental evidence. The explicit consistency experimental evidence corresponding to the multiple latent difference semantic features is used as multiple latent difference experimental evidence.
6. The evidence-based knowledge discovery method for battery materials experimental data as described in claim 1, characterized in that, MinerU and Grobid were used to deconstruct each battery material experimental paper in the collection to obtain a collection of papers in Markdown format.
7. An evidence-based knowledge discovery system for battery materials experimental data, characterized in that, The system is used to execute an evidence-based knowledge discovery method for battery materials based on experimental data as described in any one of claims 1-6, the system comprising: The deconstruction module is used to obtain a collection of experimental papers on battery materials, perform structural deconstruction processing, and obtain a collection of papers in Markdown format. The semantic parsing module is used to traverse the collection of Markdown format papers, perform semantic parsing, and build a battery material experimental library. The evidence alignment module is used to perform cross-document evidence alignment by traversing the battery material experimental library with materials as the alignment core, and to obtain multiple experimental evidences. The consistency evidence analysis module is used to traverse the multiple experimental evidences to perform explicit and implicit consistency evidence analysis, and obtain multiple consistent experimental evidences and multiple differing experimental evidences. The abstract parsing module is used to introduce a multi-dimensional abstracting mechanism to perform dual-channel source-tracing abstract parsing on the multiple consistent experimental evidence and the multiple differing experimental evidence, thereby obtaining multiple experimental semantic fast summaries and multiple fine-grained evidence summaries. The multiple experimental semantic fast summaries and multiple fine-grained evidence summaries are used as knowledge discovery results.