A large model semantic arrangement multi-agent material data mining method and system

By employing a large-scale model semantic orchestration and multi-agent collaboration approach, the problem of reliance on manual labor and insufficient automation in materials science literature data mining has been solved. This approach enables efficient and accurate acquisition and structured storage of materials data, adapts to literature scenarios of different scales, and improves the consistency and usability of results.

CN122494073APending Publication Date: 2026-07-31BEIJING TECH & BUSINESS UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING TECH & BUSINESS UNIV
Filing Date
2026-05-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, data mining of materials science literature relies heavily on manual processing, which is inefficient and prone to errors. Automated methods are difficult to adapt to complex professional semantics, the information extraction is incomplete, the processing cost of long documents is high, and there is a lack of a unified collaborative architecture for various types of data mining tasks. The consistency and usability of the results are poor, and they are difficult to use directly for database entry and subsequent analysis.

Method used

We employ a large-scale model semantic orchestration and multi-agent collaboration approach. Through literature parsing and unified document representation, we differentiate between long and short texts, assign multiple agents to perform material data mining tasks in parallel, and improve the consistency and reliability of results through an audit feedback mechanism, ultimately generating structured output.

Benefits of technology

It enables automated, efficient, and accurate mining of materials science literature data, improves data acquisition efficiency, enhances result consistency and usability, adapts to literature scenarios of different scales, and supports joint mining and structured storage of multiple types of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122494073A_ABST
    Figure CN122494073A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-agent material data mining method and system with large-scale model semantic orchestration, relating to the fields of scientific literature data mining and materials informatics. The invention first parses scientific literature such as journal articles and dissertations, constructing a unified document representation and employing differentiated processing paths for long and short texts based on document size. A large-scale model is then used to complete semantic understanding and task planning, distributing mining tasks such as material composition, magnetocaloric properties, charts, and metadata to multiple agents for collaborative execution. The output results of the multiple agents are audited, verified, analyzed for consistency, and corrected based on feedback, identifying low-reliability and conflicting data, triggering supplementary processing or redistribution, and ultimately generating structured material data that can be stored, retrieved, and used for scientific analysis. This invention significantly improves the automation, accuracy, and consistency of material data mining, and is applicable to literature mining of functional materials such as magnetocaloric, magnetic, optoelectronic, and semiconductor materials, exhibiting good versatility and scalability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of scientific literature data mining and materials informatics, specifically to a multi-agent materials data mining method and system with large-scale model semantic orchestration. Background Technology

[0002] As materials science research deepens, a vast amount of data on material composition, structure, preparation processes, performance parameters, and mechanisms of action has accumulated in scientific literature. This data forms the core foundation for materials database construction, material screening, performance prediction, and intelligent scientific research analysis. In the research fields of functional materials such as magnetic materials, optoelectronic materials, and semiconductors, key data is widely distributed in journal articles, dissertations, conference papers, book chapters, and their figures, tables, captions, and supplementary materials. This data is characterized by its dispersed sources, complex expression formats, and strong contextual dependence. Therefore, efficient and accurate extraction of materials-related data has become a key requirement in materials informatics and intelligent scientific research. Currently, traditional manual processing, rule matching, general NLP extraction, and single large model processing are the mainstream existing technologies for materials literature data mining.

[0003] Current methods for organizing materials data heavily rely on manual reading, extraction, and entry. Researchers must retrieve target information from each document and organize it into usable records, a model with significant drawbacks. Manual methods are extremely labor-intensive and inefficient, and are susceptible to variations in researcher experience, misunderstandings, and inconsistent recording habits, making it difficult to guarantee the accuracy and consistency of data extraction. This approach cannot meet the practical needs of large-scale materials data accumulation and continuous updating. Furthermore, traditional automated technologies such as rule matching, keyword retrieval, and fixed template extraction are ill-suited to the complex semantic scenarios of materials literature. When faced with multiple material objects, intertwined performance data, and correlations between experimental conditions and performance results, they commonly suffer from incomplete information extraction, ambiguous field binding relationships, and high error rates, failing to support the accurate extraction of complex materials data.

[0004] The approach of directly processing scientific literature using a single large model also has significant limitations. Long dissertations, reviews, and books have lengthy chapters, high information density, and a lot of noisy content; processing them uniformly as a whole would significantly increase computational costs and easily lead to problems such as loss of key information, omission of local evidence, and insufficient result stability. Furthermore, this approach lacks differentiated processing mechanisms for texts of varying lengths, making it difficult to balance processing efficiency and data mining accuracy. In addition, existing technologies lack a unified task organization and collaborative processing architecture. Data mining tasks of different types, such as material composition, structure, performance, charts, and metadata, are fragmented, failing to form a scalable and reusable collaborative processing system. Simultaneously, the lack of result auditing and closed-loop feedback mechanisms results in poor consistency and low reliability of automatically extracted data, making it unsuitable for direct database entry, model training, and subsequent scientific analysis, hindering efficient, accurate, and standardized data mining from materials and literature.

[0005] Therefore, a multi-agent material data mining method and system with large-scale model semantic orchestration is proposed to improve the automation level, result consistency and practical usability of scientific literature material data acquisition. Summary of the Invention

[0006] 1. The technical problem to be solved by the present invention

[0007] The purpose of this invention is to propose a multi-agent material data mining method and system with large-scale model semantic orchestration to solve the following problems existing in the prior art: (1) This technical solution mainly addresses the problems of traditional material research literature data mining being highly dependent on manual labor, inefficient, and prone to errors. At the same time, it overcomes the shortcomings of existing automated methods, such as reliance on rules and fixed templates, difficulty in adapting to complex professional semantics, incomplete information extraction, and weak cross-paragraph association processing capabilities.

[0008] (2) It solves the technical problems of high cost, lack of focus, and unstable results when large models directly process long texts, as well as the lack of a unified collaborative architecture, no verification and feedback mechanism for output results, poor consistency and availability, and difficulty in direct use for database entry and subsequent analysis in multi-type material data mining tasks.

[0009] 2. Technical Solution To address the problems in existing materials science literature data processing, such as reliance on manual processing, low automation, incomplete information extraction, low efficiency in processing long documents, insufficient coordination of information on different types of materials, and inadequate consistency and usability of results, this invention provides a materials data mining method and system based on large-model semantic arrangement and multi-agent collaboration. The aim is to automatically mine, organize, verify, and output structured materials-related information from research literature. The invention provides the following technical solution: A multi-agent materials data mining method based on large-model semantic arrangement, comprising the following steps: S1, Literature Analysis and Unified Document Representation Construction: Full-text analysis of input scientific literature materials, including text, tables, pictures, figure captions and metadata, is performed and organized into a unified structured document representation, which serves as the basic input for subsequent semantic task planning and collaborative processing; S2, Differentiated path division for short and long texts: Based on the document size, information density and content distribution characteristics, full-text semantic planning is performed on short texts, and candidate content is first screened and then semantic planning is performed on long texts; S3, Large Model Semantic Task Orchestration: The large model performs semantic understanding of documents or candidate content, identifies material objects, information types and processing needs, and generates multi-agent task distribution plans. S4, Multi-agent Parallel Collaborative Mining: Establish multiple types of mining tasks, including material composition, structure, magnetic / magnetic-thermal properties, charts, tables, and metadata, and distribute them to corresponding domain agents for parallel execution to obtain preliminary results; S5, Audit Feedback and Closed-Loop Correction: Perform confidence assessment, conflict detection, integrity verification and object binding verification on the preliminary results, identify low-confidence, conflict, missing and unclear information, and trigger supplementary processing, redistribution or result correction. S6, Structured Output: Organizes the audited and corrected results into standardized material data that can be stored, retrieved, and analyzed.

[0010] As a preferred embodiment, the differentiated path division described in S2 specifically includes: short texts directly entering full-text semantic planning; long texts first filtering candidate segments related to the target task through keyword matching and semantic similarity, and then arranging the candidate segments for tasks.

[0011] As a preferred embodiment, the large model semantic task orchestration described in S3 includes dynamic evidence pool construction, multi-round task planning, uncertainty management and reading budget control, to achieve stable parsing of complex scientific research semantics and cross-paragraph related information.

[0012] Preferably, the multiple agents in S4 include a material information agent, a magnetic data agent, a magnetocaloric effect agent, a chart parsing agent, a table extraction agent, and a metadata agent. Each agent executes in parallel and interacts through shared files.

[0013] Preferably, the audit feedback described in S5 includes result fusion, conflict decision-making, quality audit, and generation of corrective suggestions, forming a closed-loop mechanism of "planning—execution—fusion—audit—feedback—reprocessing".

[0014] Preferably, the material data includes the composition, structure, preparation, performance parameters, and experimental conditions of magnetocaloric materials, magnetic materials, optoelectronic materials, and semiconductor materials, and is preferentially applicable to the mining of magnetocaloric performance data such as magnetic entropy change.

[0015] Furthermore, a multi-agent materials data mining system with large-scale model semantic orchestration is proposed, comprising: Literature parsing and document representation module: used to parse materials science literature and construct a unified structured document representation; Semantic orchestration module: Based on a large model, it realizes differentiated processing of long and short texts, semantic understanding, task planning and agent scheduling; Multi-agent collaborative processing module: contains multiple domain-specific intelligent agents that perform parallel extraction and mining of material composition, structure, performance, charts, tables, and metadata; Audit feedback module: used for result fusion, confidence assessment, conflict resolution, integrity verification, and feedback correction; Structured output module: Used to generate standardized, database-ready material data results.

[0016] Preferably, the semantic orchestration module includes an OrchestratorAgent, a dynamic evidence pool, a multi-round planning unit, an uncertainty management unit, and a parallel scheduling unit, supporting adaptive processing of long and short texts.

[0017] Preferably, the audit feedback module includes a MergerAgent, a shared file building unit, a conflict decision engine, a QA auditor, and a feedback correction unit, forming a closed-loop quality control.

[0018] Preferably, the system supports the addition of new intelligent agents to adapt to new material systems and new performance types, while maintaining the overall process of unified document representation, semantic orchestration, collaborative execution, audit feedback, and structured output.

[0019] Compared with existing technologies, the multi-agent material data mining method and system with large-scale model semantic orchestration provided by this invention has the following beneficial effects: 1. This invention can automatically mine relevant data from scientific research literature, reducing reliance on manual reading, manual extraction, and manual organization, and improving the efficiency of data acquisition.

[0020] 2. By adopting differentiated processing paths for documents of different sizes, this invention improves the data mining adaptability in both long and short document scenarios, which is beneficial to balancing processing efficiency and result quality.

[0021] 3. This invention improves the ability to process complex scientific semantics, multi-type information, and cross-paragraph information relationships by using large models for semantic task planning, and overcomes the limitations of traditional rule-based methods in complex scenarios.

[0022] 4. This invention improves the ability to jointly mine material composition, structure, properties and related descriptive information by having multiple processing units work together to execute different types of material data processing tasks.

[0023] 5. This invention improves the consistency, reliability, and usability of the output results through auditing, verification, and feedback mechanisms, making the generated results more suitable for structured storage, retrieval, and subsequent research and analysis.

[0024] 6. This invention has good scalability. In addition to being preferably applied to data mining related to magnetocaloric materials, it can also be extended to data mining scenarios of other material systems and other material properties. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the overall architecture of a material data mining system based on large model semantic orchestration and multi-agent collaboration according to the present invention. Figure 2 This is a schematic diagram of the overall process of a material data mining method based on large model semantic orchestration and multi-agent collaboration according to the present invention; Figure 3 This is a schematic diagram of the multi-agent collaborative processing and audit feedback closed loop of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0027] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0028] Example 1, please refer to Figures 1 to 3 As shown: To address the problems mentioned in the technical solutions, this application provides a material data mining method based on large-scale model semantic orchestration and multi-agent collaboration. This method is designed for scenarios involving the acquisition of material and performance information from scientific literature. It parses and processes the input literature to construct a unified document representation, and then implements differentiated processing strategies based on the characteristics of the literature content. On this basis, a large-scale model is used to complete the semantic depth understanding and task planning of the literature content, and the relevant processing tasks of different types of material data are allocated to multiple processing units to achieve multi-unit collaborative execution. Subsequently, the output results of each processing unit are audited, verified, and adjusted as necessary to ultimately generate standardized structured material data.

[0029] Specifically, the steps include the following: The first step is to comprehensively analyze the input research literature. Research literature can include journal articles, dissertations, conference papers, book chapters, and other literature carriers containing research content. By systematically sorting and organizing the main text, chapter structure, table data, images, figure captions, and metadata of the literature, a unified document representation result is constructed, which serves as the basic input for subsequent semantic task planning and collaborative processing.

[0030] The second step involves dividing the processing path for the document content based on a unified document representation. For shorter documents with a concentrated structure suitable for overall semantic understanding, a full-text semantic planning model is adopted; for longer documents with complex content and dispersed information, a step-by-step processing model of "candidate content screening - semantic planning" is used. This differentiated processing strategy effectively improves the efficiency and accuracy of data mining from research documents of different sizes.

[0031] The third step involves using a large-scale model to perform semantic understanding and task planning of the document content. Based on the input full text of the document or the filtered candidate content, the large-scale model accurately identifies the target objects, target information types, and corresponding processing requirements in the material data mining task, and generates standardized task distribution results, providing a basis for subsequent multiple processing units to perform data mining tasks in a targeted manner.

[0032] Furthermore, based on the task distribution results, multiple processing units are invoked to collaboratively execute material data mining tasks. Each processing unit is responsible for processing one or more of the following: material composition information, structural information, performance information, tabular information, graphic and textual descriptions, and literature-related information. Each processing unit can execute tasks independently or construct collaborative processing relationships according to task requirements, achieving joint mining and efficient extraction of multiple types of material information from scientific literature.

[0033] After each processing unit completes data processing, the output results undergo comprehensive auditing, verification, and consistency analysis. The auditing process focuses on identifying low-reliability information, conflicting information, missing information, and ambiguous object correspondences in the results, and based on this, determines whether supplementary processing, task redistribution, data reprocessing, or result correction is necessary. This closed-loop audit feedback mechanism significantly improves the consistency and reliability of the final data results.

[0034] Finally, the audited and verified processing results are organized into structured material data and output. The structured results can be directly used for tasks such as file export, database storage, information retrieval, and subsequent material research and analysis, forming a complete solution for scientific literature material data mining.

[0035] This invention also proposes a materials data mining system based on large-model semantic orchestration and multi-agent collaboration. The system corresponds to the aforementioned method and includes a document parsing module, a document representation construction module, a task orchestration module, a collaborative processing module, an audit feedback module, and a result output module; wherein: The literature analysis module comprehensively analyzes scientific research literature and extracts various core information. The document representation construction module integrates the analyzed information to form a unified document representation. The task orchestration module realizes semantic understanding and task planning based on a large model and generates task distribution results. The collaborative processing module calls multiple processing units to complete the collaborative mining of multi-type material information. The audit feedback module audits, verifies, and analyzes the processing results to ensure the reliability of the results. The result output module organizes the audited results into structured material data and completes the output.

[0036] Specifically, the present invention differs from the prior art in that: 1. Design of the extraction protocol for the top-level field registry driver; This invention establishes a unified field registry at the system top level, uniformly defining the field path, field alias, field type, domain affiliation, unit type, evidence requirements, and output format of the fields to be extracted. This design ensures that each intensive reading agent does not generate results arbitrarily, but rather performs field-level extraction under a unified protocol.

[0037] This feature addresses the issues of inconsistent field naming, difficulty in merging synonymous fields, and difficulty in merging outputs from different extraction modules in existing technologies, and provides a unified extension entry point for subsequent additions of fields, material domains, and sub-Agents.

[0038] 2. Protocol-based management of field aliases and search targets; For the same material field, this invention can maintain multiple professional expressions, abbreviations, synonyms, and different documentary notations in the field registry, such as various expressions for fields like magnetocaloric, magnetic, structural, and process. Both the recall layer and the agent task can generate search targets based on this field alias protocol.

[0039] This feature addresses the issue of inconsistencies in terminology used in literature, which can lead to missed keyword retrievals. It also reduces the cost of manually maintaining prompts and field rules for each agent.

[0040] 3. Field-level resource budgeting and task granularity control; Because fields are registered uniformly, the system can perform field-level control over token consumption and reading budget based on the number of fields, field importance, field domain, and estimated context length.

[0041] This feature addresses the issues of high cost and uncontrollable resource consumption in indiscriminate full-text reading of long text materials, enabling large model calls to shift from "document-level coarse-grained processing" to "field-level budgetable processing".

[0042] 4. Deterministic preprocessing for the top-level unit normalization table; This invention establishes a unified unit standardization table to standardize the conversion of units in the materials science field, such as temperature, magnetic field, magnetization intensity, magnetic entropy change, cooling capacity, length, angle, and energy. Unit standardization is preferably completed through deterministic rules before the audit agent is audited, rather than being determined ad hoc by a large model.

[0043] This feature is designed to address the problems in existing technologies, such as diverse unit representations, incomparable numerical values, reliance on large models for conversion that is prone to errors and cannot be reproduced.

[0044] 5. The results of unit unification are used for initial conflict screening before the audit; After unit unification, the system can perform consistency checks by using the unified unit string and the unified unit value; if there are still units that cannot be identified, units that are not ununified, or unit conflicts in the fields, the system will proceed to the audit agent for further processing.

[0045] This feature is used to address the difficulty of directly comparing the output results of multiple agents at the unit level, and to reduce the token waste caused by the audit agent processing a large number of deterministic unit conversions.

[0046] 6. Audit feedback-based supplementation and registration of unit rules; For newly emerging unit notation or uncovered conversion relationships, the audit agent can generate supplementary normalization rules or prompt for the addition of new conversion methods, and register them in the unit normalization table for reuse in subsequent documents.

[0047] This feature addresses the issue of constantly changing unit notation in the materials field and high manual maintenance costs, enabling the system to form iteratively enhanceable unit knowledge assets rather than a one-time static rule base.

[0048] 7. Domain keyword scoring mechanism for the long text recall layer; This invention establishes a keyword lookup table in the materials science field during long text processing, including core terms such as coercivity, magneticocaloric, entropy change, curie temperature, ΔS, and RCP. When the corresponding keyword appears in an evidence block, the system assigns a score to the block based on field importance, contextual position, chapter type, and signal strength, and prioritizes its submission to the Orchestrator.

[0049] This feature addresses the challenges of large text size, scattered target information, and high cost of feeding the full text into a large model in long papers, dissertations, and reviews.

[0050] 8. A fusion mechanism between keyword recall and Embedding / RAG semantic recall; As an optional implementation, the system can vectorize document blocks and vectorize field aliases, task descriptions, and professional query statements in the field registry. It can then use RAG similarity retrieval to discover high-value blocks that are semantically related to the target field. Finally, it can integrate these with keyword scoring results to obtain a comprehensive score. When the score reaches a preset threshold, Top-K, or adaptive threshold, it can be submitted to the Orchestrator.

[0051] This feature addresses the issue of insufficient recall for synonyms, implicit expressions, and cross-sentence expressions when relying solely on keywords. This approach is more rigorous.

[0052] 9. Recall layer self-reflection and keyword registry self-evolution mechanism; In the process of reading long texts in multiple rounds, this invention does not rely solely on a pre-fixed keyword list or a one-time RAG recall result. Instead, after each round of processing by the Orchestrator, it conducts a reflective analysis on "high-value evidence blocks that were discovered but failed to be hit by the initial recall layer." Specifically, when the Orchestrator or audit agent discovers certain blocks with high-value information through expanded reading, sub-agent feedback, or evidence chain verification, but these blocks are not effectively identified by the initial recall layer, the system extracts domain keywords, synonyms, unit expressions, field context patterns, or material system terms from these blocks and calls the registration tool to add them to the recall layer keyword registry. At the same time, it adjusts the scoring weights or recall rules of the corresponding fields.

[0053] This feature addresses the problems of static recall rules and poor adaptability to new terminology, material systems, and non-standard expressions in existing technologies. By feeding back the reasons for missed recalls to the recall layer, this invention enables the system to continuously accumulate new high-value keywords and scoring rules during document processing, thereby reducing invalid reading and repeated calls to similar documents, improving long text reading efficiency, recall accuracy, and system boundary expansion capabilities. This mechanism transforms the recall layer from a static rule base into a dynamic recall system that can be progressively enhanced with audit and compilation feedback.

[0054] 10. Standardize the EvidenceBlock evidence carrier; This invention organizes text paragraphs, tables, images, captions, and other information into a unified EvidenceBlock or equivalent evidence structure with block_id, page number, chapter, evidence type, signal, and source. Each agent's output must be bound to evidence, rather than simply providing extracted values.

[0055] This feature is used to address the problems of existing material data mining results being untraceable, unverifiable, and difficult to audit.

[0056] 11. Material records should be distinguished by their "data production environment," rather than simply by their chemical formula; This invention recognizes that materials with the same chemical formula may produce different performance data, such as magnetic entropy change, Curie temperature, and magnetization, depending on the preparation process, test magnetic field, temperature range, experimental method, calculation method, or sample condition. Therefore, the system not only identifies the chemical formula but also the corresponding production environment of the data, and determines whether to record and merge or separate the data accordingly.

[0057] This feature is used to solve the problem in existing technologies where data from different experimental conditions are mistakenly merged and material property records are distorted due to polymerization based solely on chemical formulas.

[0058] 12. A unified schema output mechanism for detailed reading of sub-agents; Each domain-specific sub-agent outputs field values, units, confidence levels, notes, evidence citations, and data production environment information according to a unified schema. The sub-agent's responsibility extends beyond simply finding numerical values; it also includes determining which material object, experimental / calculation condition, and whether the evidence is sufficient.

[0059] This feature is used to solve the problems of difficulty in comparing, merging, and storing free text outputs from multiple agents.

[0060] 13. Mechanism for determining the merging and separation of audit agent's material records; The audit agent uses chemical formulas as initial clues, but further examines the notes, evidence chains, and data production environments of each sub-agent to determine whether data with the same chemical formula comes from the same material record; if the experimental conditions, preparation methods, or calculation methods are different, they are kept separate; if the evidence is consistent and the environment is the same, they are merged.

[0061] This feature is used to solve the dual problems of "mistaken merging of materials with the same name" and "duplicate recording of the same material" in multi-agent systems.

[0062] 14. Audit feedback-driven targeted rereading mechanism; When the audit agent determines that the evidence for a certain field is insufficient, the confidence level is low, the unit is abnormal, or the material attribution is unclear, it does not directly output the final result. Instead, it generates audit notes and uncertainty information and feeds them back to the Orchestrator. The Orchestrator then expands the relevant blocks in a targeted manner and reassigns them to the corresponding sub-agents for in-depth reading.

[0063] This feature is used to solve the problems of one-time extraction being unable to be corrected and erroneous results being unable to be closed-loop processed.

[0064] 15. Stateful multi-round Orchestrator and reading budget stop strategy; The Orchestrator of this invention maintains rounds, coverage, reading budget, accessed evidence, pending uncertainties, and downstream extraction readiness, and iteratively executes the process of "discovering material objects, assigning in-depth reading tasks, auditing feedback, generating uncertainty, building extension requests, expanding the evidence pool, and updating the budget / stopping." Extension methods include nearest-neighbor block extension, same-chapter extension, captionfamily extension, table / figure related block extension, and field-guided extensions such as process / property / structure.

[0065] This feature addresses the issues of blindly reading the entire text, repeated reading, and diminishing marginal returns in long text processing, enabling the system to stop when there is sufficient evidence and continue targeted reading where there is insufficient evidence.

[0066] 16. A hierarchical memory mechanism where the Orchestrator and Audit Agent share memory, while sub-agents have locally unshared memory; This invention employs a hierarchical memory-sharing mechanism, preferably using caching components such as Redis to store the global working memory between the Orchestrator and the Audit Agent. This memory includes the current document's list of material objects, accessed evidence blocks, audit conclusions, unresolved conflicts, uncertainty records, reading budget status, and historical task allocation results. The Orchestrator maintains the overall workflow based on this shared memory, while the Audit Agent uses this shared memory to determine which material records meet the evidence requirements, which fields need to be read again, and which objects should not be processed repeatedly.

[0067] In contrast, each in-depth reading sub-agent does not force the sharing of global memory. Instead, it focuses on the local evidence package of its respective domain according to the field registry and task allocation results. The reason for this design is that different sub-agents are targeting different extraction domains, such as material composition, magnetic properties, magnetocaloric properties, table parsing, and figure annotation parsing. If the global context is shared indiscriminately with all sub-agents, it will introduce irrelevant noise, increase token consumption, and distract the model's attention, thereby reducing the accuracy of extracting fixed domain fields.

[0068] This feature addresses the issues of excessive shared context, noise accumulation, unclear task boundaries, and token waste in existing multi-agent systems. Through a hierarchical design where "global memory is shared only between the Orchestrator and the auditing agent, while sub-agents maintain local domain-specific in-depth reading," this invention ensures both the continuity and traceability of the overall process, as well as the domain focus and scalability of the in-depth reading agent. When a sub-agent discovers insufficient evidence or missing blocks during local in-depth reading, it can generate a missing information or expansion request and send it back to the Orchestrator. The Orchestrator then uses the shared memory to uniformly decide whether to perform nearest-neighbor expansion, same-chapter expansion, table / graph-related expansion, or field-oriented expansion, and reassigns the in-depth reading task.

[0069] Example 2: Based on Embodiment 1, but with some differences, the following description, in conjunction with specific examples and accompanying drawings, illustrates a multi-agent material data mining method and system for large-model semantic orchestration proposed in this invention. The specific details are as follows: First, the overall implementation method of the system; In one implementation, the system includes a document parsing module, a document representation construction module, a task orchestration module, a multi-agent collaborative processing module, an audit feedback module, a result output module, and a data storage module.

[0070] The system comprises the following modules: a literature parsing module for receiving input research literature and performing basic parsing; a document representation construction module for organizing the literature parsing results into document representations that can be used for subsequent processing; a task orchestration module for semantic understanding and task planning of the literature content based on a large model; a multi-agent collaborative processing module for performing division of labor processing around different types of information; an audit feedback module for verifying the processing results, performing consistency analysis, and providing necessary feedback; a result output module for generating structured results; and a data storage module for saving, managing, or using the results for subsequent research and analysis.

[0071] Second, implementation methods for literature analysis and document representation construction; In one implementation, the system first receives the research literature to be processed and parses and organizes the text, chapters, tables, images, figure captions, and related metadata within the literature. The parsed content is then uniformly represented as a document representation result that can be used for subsequent processing.

[0072] The document representation serves to convey the main organizational relationships within the literature, enabling subsequent modules to perform semantic understanding, task planning, collaborative processing, and result auditing based on a unified input. This approach avoids issues of inconsistent input formats and information organization that arise from research literature of different sources and structures during subsequent processing.

[0073] Third, the implementation method of dividing the document processing path; In one implementation, after obtaining the document representation result, the system performs processing path division on the input document. The processing path division can be determined based on document length, document size, content distribution characteristics, or information density.

[0074] For research literature that is relatively short, has relatively concentrated target information, and is suitable for overall understanding, the system adopts a full-text semantic processing path, with the task arrangement module performing semantic understanding and task planning based on the overall content.

[0075] For lengthy research documents with substantial content and dispersed information, the system employs a candidate content screening approach followed by semantic processing. This involves first selecting candidate content relevant to the target task from the lengthy text, and then performing semantic understanding and task planning based on this candidate content. This differentiated processing improves the system's adaptability to documents of varying sizes.

[0076] Fourth, implementation methods for semantic task planning; In one implementation, the task orchestration module performs semantic understanding of the document content based on a large model, identifies target objects, target fields, target information types and processing requirements related to material data mining, and generates task planning results.

[0077] The task planning results are used to determine the types of processing units to be invoked subsequently, the scope of processing objects, and the corresponding processing order. Compared with traditional fixed-rule routing methods, this approach is better suited to the characteristics of scientific literature, such as complex information expression, strong contextual dependencies, and obvious cross-paragraph connections, thereby improving the ability to identify and organize information from multiple types of materials.

[0078] Fifth, implementation method of multi-agent collaborative processing; In one implementation, the system assigns different types of data mining tasks to multiple processing units for collaborative execution based on task planning results. These multiple processing units are designed for different types of information and preferably include one or more of the following: a material basic information processing unit, a performance information processing unit, a tabular information processing unit, a graphic description processing unit, and a literature information processing unit.

[0079] Each processing unit can process material composition, structure, properties, tabular data, graphic descriptions, and literature information according to its assigned tasks, and generate corresponding intermediate results. In a preferred embodiment, multiple processing units can execute in parallel to improve overall processing efficiency; in other embodiments, they can also be executed sequentially according to a preset order, or a combination of partially parallel and partially serial execution.

[0080] Sixth, the implementation method of audit feedback; In one implementation, after multiple processing units complete preliminary processing, the system sends their output results to the audit feedback module. The audit feedback module is used to analyze the reliability, consistency, completeness, and object correspondence of the results.

[0081] In this invention, structured validation rules are set in the audit agent, which not only summarize the output of multiple agents, but also judge the field confidence, unit consistency, numerical conflict, evidence integrity and material object binding relationship, and decide whether to trigger reprocessing based on the validation results.

[0082] Specifically, the verification rules of this invention may include: The confidence threshold is verified as follows: When the confidence level of a certain field is lower than the preset threshold (<0.7), or when it is marked as low confidence / weak evidence in the system, the field is not directly used as the final result output. Instead, it is marked as a low confidence field, and the Orchestrator is triggered to expand the relevant EvidenceBlock by neighbor, same chapter, or field-oriented expansion, and then handed over to the corresponding sub-Agent for rereading.

[0083] Automatic field conflict marking When multiple agents or multiple rounds of extraction results provide different values, units, or textual conclusions for the same field within the same material object and data output environment, the audit agent automatically generates conflict markers, recording the conflicting field, conflict source, candidate value, evidence block, and related Agentnote. If the conflict cannot be resolved through unit normalization, evidence strength, or contextual judgment, feedback is sent to the Orchestrator to trigger a targeted reread.

[0084] Unit consistency verification; For numerical fields such as temperature, magnetic field, magnetic entropy change, magnetization intensity, and cooling capacity, they are first standardized using a unit normalization table. If the unit cannot be identified, is not successfully normalized, or is still inconsistent after normalization, it is marked as a unit anomaly and enters the audit feedback process.

[0085] Evidence integrity verification All found fields must be linked to evidence sources such as EvidenceBlocks, tables, figure captions, or image descriptions. If a field only contains a numerical value but lacks supporting evidence, or if the evidence cannot support the connection between the field and the material object / experimental conditions, it will be marked as insufficient evidence, triggering a supplementary reading or task reassignment.

[0086] Material object and data output environment verification For different results under the same chemical formula, the audit agent further verifies whether the data output environment, such as the preparation process, heat treatment conditions, test magnetic field, temperature range, and experimental / calculation methods, is consistent. If the environment is inconsistent, it is kept separate; if the missing environment makes it impossible to determine, it is marked as an unclear object binding and targeted extended reading is triggered.

[0087] Through the above verification rules, this invention forms a quality control chain of "confidence judgment - conflict marking - unit normalization - evidence integrity check - object binding verification - feedback reprocessing". This supplement is used to solve the problems in the prior art, such as the lack of clear verification standards for extraction results, direct output of low-confidence fields, inability to interpret field conflicts, and inability to automatically reprocess.

[0088] Specifically, this invention establishes a unified unit standardization table to standardize the conversion of units in the materials science field, such as temperature, magnetic field, magnetization, magnetic entropy change, cooling capacity, length, angle, and energy. Unit standardization is prioritized and completed through deterministic rules before the audit agent, rather than being determined ad hoc by the large model. For newly emerging unit notations or uncovered conversion relationships, the audit agent can generate supplementary standardization rules or suggest new conversion methods, registering them in the unit standardization table for reuse in subsequent literature.

[0089] This feature addresses the issue of constantly changing unit notation in the materials field and high manual maintenance costs, enabling the system to form iteratively enhanceable unit knowledge assets rather than a one-time static rule base.

[0090] Stateful multi-round Orchestrator and reading budget stop strategy The Orchestrator of this invention maintains rounds, coverage, reading budget, accessed evidence, pending uncertainties, and downstream extraction readiness, and iteratively executes the process of "discovering material objects, assigning in-depth reading tasks, auditing feedback, generating uncertainty, building extension requests, expanding the evidence pool, and updating the budget / stopping." Extension methods include nearest-neighbor block extension, same-chapter extension, captionfamily extension, table / figure related block extension, and field-guided extensions such as process / property / structure.

[0091] This feature addresses the issues of blindly reading the entire text, repeated reading, and diminishing marginal returns in long text processing, enabling the system to stop when there is sufficient evidence and continue targeted reading where there is insufficient evidence.

[0092] A hierarchical memory mechanism where the Orchestrator and Audit Agent share memory, while sub-agents have locally unshared memory. This invention employs a hierarchical memory-sharing mechanism, preferably using caching components such as Redis to store the global working memory between the Orchestrator and the Audit Agent. This memory includes the current document's list of material objects, accessed evidence blocks, audit conclusions, unresolved conflicts, uncertainty records, reading budget status, and historical task allocation results. The Orchestrator maintains the overall workflow based on this shared memory, while the Audit Agent uses this shared memory to determine which material records meet the evidence requirements, which fields need to be read again, and which objects should not be processed repeatedly.

[0093] In contrast, each in-depth reading sub-agent does not force the sharing of global memory. Instead, it focuses on the local evidence package of its respective domain according to the field registry and task allocation results. The reason for this design is that different sub-agents are targeting different extraction domains, such as material composition, magnetic properties, magnetocaloric properties, table parsing, and figure annotation parsing. If the global context is shared indiscriminately with all sub-agents, it will introduce irrelevant noise, increase token consumption, and distract the model's attention, thereby reducing the accuracy of extracting fixed domain fields.

[0094] This feature addresses the issues of excessive shared context, noise accumulation, unclear task boundaries, and token waste in existing multi-agent systems. Through a hierarchical design where "global memory is shared only between the Orchestrator and the auditing agent, while sub-agents maintain local domain-specific in-depth reading," this invention ensures both the continuity and traceability of the overall process, as well as the domain focus and scalability of the in-depth reading agent. When a sub-agent discovers insufficient evidence or missing blocks during local in-depth reading, it can generate a missing information or expansion request and send it back to the Orchestrator. The Orchestrator then uses the shared memory to uniformly decide whether to perform nearest-neighbor expansion, same-chapter expansion, table / graph-related expansion, or field-oriented expansion, and reassigns the in-depth reading task.

[0095] This feature is designed to address the problems in existing technologies, such as diverse unit representations, incomparable numerical values, reliance on large models for conversion that is prone to errors and cannot be reproduced.

[0096] Specifically, the audit feedback module can identify one or more of the following: 1. Some results have low reliability; 2. Conflicts exist between results from different processing units; 3. Some target information is missing; 4. Although some results are extracted, the object binding relationships are unclear; Seventh, some results are usable but do not meet the preset output specifications.

[0097] After identifying the above situations, the audit feedback module can generate feedback results to trigger supplementary processing, redistribution, reprocessing, result correction, or merging, thereby forming more consistent and reliable output results. Through this closed-loop mechanism, the system can not only complete one-time data mining but also further verify and correct the results, improving the actual usability of the structured results.

[0098] Eighth, implementation of structured output; In one implementation, the results after audit feedback processing are organized into structured material data results. The structured results may include one or more of the following: material composition information, structural information, performance information, literature information, and related explanatory information, and are output in a uniform format.

[0099] Structured results can be written to files, databases, or other data storage media for subsequent tasks such as retrieval, statistical analysis, database construction, model training, or research-assisted decision-making. Because they undergo collaborative processing and audit feedback before output, the resulting outputs offer advantages in consistency, readability, and subsequent usability.

[0100] Ninth, a preferred implementation method for magnetocaloric performance data; In a preferred embodiment, the present invention is applied to data mining scenarios related to literature on magnetocaloric materials. The system extracts information related to material composition, structure, magnetocaloric properties, and related experimental descriptions from scientific literature, and through the aforementioned literature analysis, task planning, multi-agent collaborative processing, audit feedback, and structured output process, generates structured results that can be used for materials database construction and subsequent research and analysis.

[0101] In this preferred embodiment, the present invention is particularly suitable for processing magnetocaloric performance data such as magnetic entropy change, but the present invention is not limited thereto. The same overall method and system framework can also be used for data mining of other magnetic properties, thermal properties, structural parameters, process conditions, and general material properties.

[0102] Tenth, extended implementation methods; In another implementation, the system can be expanded with new processing units to adapt to new data types or new material performance targets, based on the needs of different materials research tasks. This expansion does not change the overall processing flow of the invention; it still operates based on unified document representation, semantic task planning, multi-agent collaborative processing, audit feedback, and structured output.

[0103] Therefore, this invention is not only applicable to the currently preferred material data mining scenario, but can also be extended to more material systems and more scientific literature data organization tasks, and has good versatility and scalability.

[0104] Example 3, based on the specific implementation process of Example 2, is as follows: A journal article related to magnetocaloric materials is used as input literature. The system first parses the literature to form a unified document representation; then, based on the size of the literature, it determines which short text processing path to use. in: Short text processing path When the number of document blocks >= 500, pages >= 30, and the number of tables and figures are all at a low level, the system can directly submit the full text or main content to the Orchestrator for overall semantic planning.

[0105] Long text processing path When the number of document blocks is greater than or equal to 500 and the number of pages is greater than or equal to 30, especially for dissertations, reviews, book chapters, or long documents containing a large number of experimental results, the system does not directly perform full-text large model processing, but instead first enters the candidate content screening process.

[0106] Keyword filtering rules The system performs an initial screening of each EvidenceBlock based on a materials science keyword registry. For example, blocks containing keywords such as magneticocaloric, entropychange, ΔS, Curietemperature, Tc, coercivity, RCP, RC, saturationmagnetization, heat treatment, and synthesis receive higher recall scores. The system also weights the responses based on chapter titles, block location, table / figure citations, and field importance.

[0107] Semantic similarity filtering rules As an optional implementation, the system vectorizes the EvidenceBlock and also vectorizes the field aliases, task descriptions, and professional query statements in the field registry, calculating semantic similarity through Embedding / RAG. Blocks with a similarity score reaching a preset threshold of 0.75, or those located in the Top-K / Top-p candidate set, are identified as candidate high-value evidence blocks.

[0108] Fusion scoring rules The system can integrate keyword scores, semantic similarity scores, chapter position scores, evidence signal scores, and table / graph association scores. For example, if the overall score reaches a preset threshold of ≥0.8, the EvidenceBlock is submitted to the Orchestrator; content that does not reach the threshold but is adjacent to the selected block, in the same chapter, or related to tables / graphs can be retained as expansion candidates.

[0109] The semantic task planning module understands the full text and generates task distribution results. Then, multiple processing units collaboratively process material composition information, performance information, tabular information, and literature information. After obtaining preliminary results, the audit feedback module performs consistency analysis and verification, and triggers reprocessing for parts requiring supplementary processing. Finally, structured material data results are generated. This embodiment demonstrates that the present invention can automatically mine, organize, and output material-related information from scientific literature.

[0110] Example 4, differing from Example 3, is based on Example 2 and its specific implementation process is as follows: A lengthy dissertation is used as the input document. The system first parses the document and forms a unified document representation. Then, based on the document's size, it enters the long text processing path. Content relevant to the material data mining task is extracted through candidate content filtering. The semantic task planning module then generates task distribution results, which are then collaboratively processed by multiple processing units. After verification by the audit feedback module, a structured output result is generated. This example demonstrates that the present invention is applicable to data mining scenarios involving long scientific research documents.

[0111] Example 5 differs from Examples 3 and 4 in that, after multiple processing units process the same input document, the output results may exhibit differences in confidence levels or inconsistencies in some fields. The system analyzes the results through an audit feedback module, identifies low-confidence or conflicting results, and triggers supplementary processing or redistribution, ultimately resulting in a more consistent structured output. This example demonstrates that the present invention can not only perform material data mining but also verify and correct the results, thereby improving the reliability of the output.

[0112] Example 6: Multi-material mixed document processing scenario; In one embodiment, the input literature includes both magnetocaloric materials and optoelectronic materials. For example, in the same review or comprehensive research paper, one section describes the magnetic entropy change, Curie temperature, test magnetic field, and cooling capacity of magnetocaloric materials, while another section describes the band gap, absorption spectrum, photoelectric conversion efficiency, or carrier mobility of optoelectronic materials.

[0113] The system first constructs a unified EvidenceBlock evidence carrier through the literature parsing module. Then, the Orchestrator, based on the field registry and material object identification results, initially distinguishes different material objects and their related evidence blocks. For evidence blocks containing magnetocaloric keywords and magnetic / magnetocaloric fields, the system distributes them to the Magnetic Agent and Magnetocaloric Agent; for evidence blocks containing photoelectric performance keywords and photoelectric fields, they can be distributed to the extended photoelectric performance Agent. Each Agent outputs field values, units, confidence levels, notes, and evidence citations according to a unified schema.

[0114] The audit agent performs object binding validation on the outputs of different agents based on shared archives to determine whether each field belongs to the same material object, the same data production environment, and the same performance system. If it is found that the magnetocaloric field is mistakenly bound to the optoelectronic material object, or that the same chemical formula corresponds to different sample states under different research objectives, the conflict is automatically marked and reported to the Orchestrator for targeted rereading.

[0115] This embodiment is used to illustrate that the present invention can adapt to complex scenarios in which multiple material systems and multiple performance types coexist in the same document, verify the ability of multi-agent division of labor, field protocol constraints and object binding, and avoid data crosstalk between different material objects or different performance systems.

[0116] Example 7: Low-quality document processing scenario; In another embodiment, the input literature has problems such as blurry figures and tables, OCR recognition errors, non-standard expressions, inconsistent unit writing, scattered descriptions of experimental conditions, or incomplete figure captions.

[0117] The system first uses the document parsing module to structure the text, tables, figure captions, and image descriptions, organizing the identifiable content into EvidenceBlocks. If the confidence level of a field output by the figure parsing agent or table parsing agent is lower than a preset threshold, such as less than 0.7, or if the field lacks units, test conditions, or the source of evidence is incomplete, the audit agent will not directly output the final result, but will instead mark the field as low confidence or insufficient evidence.

[0118] Subsequently, the audit agent provides feedback to the orchestrator regarding the problematic field, related evidence blocks, missing evidence types, and suggested expansion directions. The orchestrator then performs expansions based on this feedback, including expansions of neighboring blocks, blocks within the same chapter, figure / table notes, related blocks in tables / figures, or field-oriented expansions, and reassigns the supplementary evidence to the corresponding sub-agent for closer review. If the output conditions still cannot be met after supplementing the evidence, the system marks the field as unconfirmable or requiring manual review, and retains the audit note and evidence chain.

[0119] This embodiment illustrates that the present invention does not simply output unreliable results in low-quality material literature, but improves the reliability of the results through confidence verification, evidence supplementation, audit feedback and multi-round reprocessing mechanisms, and verifies the closed-loop correction capability.

[0120] In summary, through Implementation 2 of this example, this solution is not a simple superposition of existing materials data mining, large-scale model applications, and multi-agent technologies, but rather a systematic innovation for materials science literature mining scenarios. Its core novelty and innovation are reflected in: pioneering a long-short text adaptive differential processing mechanism based on large-scale model semantic arrangement, combined with dynamic evidence pools, multi-round planning, and reading budget control to achieve accurate parsing of complex scientific research semantics; constructing a parallel collaborative mining system of dedicated agents in fields such as material composition, magnetic / magnetothermal properties, charts, and metadata, breaking through the information extraction limitations of traditional single-model and rule-based methods; establishing a closed-loop quality control mechanism of "planning-execution-fusion-audit-feedback-reprocessing" to complete confidence assessment, conflict detection, and result correction, ensuring data consistency and usability; and adopting a unified document representation and scalable agent architecture, taking into account the adaptability of multiple material systems such as magnetothermal, magnetic, optoelectronic, and semiconductor, comprehensively surpassing existing technologies in terms of processing efficiency, mining accuracy, result reliability, and scenario universality, forming a unique intelligent mining technology path for materials literature data.

[0121] Please refer to the above work process. Figures 1 to 3 .

[0122] It should be noted that the term "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0123] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multi-agent material data mining method for large model semantic orchestration, characterized in that, Includes the following steps: S1, Full-text parsing of materials science literature: Full-text parsing of input materials science literature, including text, tables, pictures, figure captions and metadata, organized into a unified structured document, which serves as the basic input for subsequent semantic task planning and collaborative processing; S2, based on document size and information density, adaptively divides the processing paths for long and short texts: based on document size, information density and content distribution characteristics, full-text semantic planning is performed on short texts, and candidate content is first screened and then semantic planning is performed on long texts. S3, Large Model Semantic Task Orchestration: The large model performs semantic understanding on the documents or candidate content, identifies the material objects, information types and processing needs, and generates a multi-agent task distribution plan; the semantic task orchestration further includes semantic anchoring based on the material knowledge graph, adaptive dynamic allocation of reading budget, and multi-agent evidence chain consensus planning. S4, Multi-agent Parallel Collaborative Mining: Establish multiple types of mining tasks, including material composition, structure, magnetic / magnetic-thermal properties, charts, tables, and metadata, and distribute them to corresponding domain agents for parallel execution to obtain preliminary results; S5, Audit Feedback and Closed-Loop Correction: Perform confidence assessment, conflict detection, integrity verification and object binding verification on the preliminary results, identify low-confidence, conflict, missing and unclear information, and trigger supplementary processing, redistribution or result correction. S6, Structured Output: Organizes the audited and corrected results into standardized material data that can be stored, retrieved, and analyzed.

2. The multi-agent material data mining method of large model semantic orchestration according to claim 1, characterized in that, The differentiated path division in S2 specifically includes: short texts directly enter full-text semantic planning; long texts first filter candidate segments related to the target task through keyword matching and semantic similarity, and then arrange the candidate segments for task.

3. The multi-agent material data mining method with large-scale model semantic orchestration according to claim 1, characterized in that, The S3 large model semantic task orchestration includes dynamic evidence pool construction, multi-round task planning, uncertainty management and reading budget control, to achieve stable parsing of complex scientific research semantics and cross-paragraph related information.

4. The multi-agent material data mining method with large-scale model semantic orchestration according to claim 1, characterized in that, The S4 contains multiple intelligent agents, including a material information intelligent agent, a magnetic data intelligent agent, a magnetocaloric effect intelligent agent, a chart parsing intelligent agent, a table extraction intelligent agent, and a metadata intelligent agent. These agents execute in parallel and interact through shared archives.

5. The multi-agent material data mining method with large-scale model semantic orchestration according to claim 1, characterized in that, The audit feedback in S5 includes result fusion, conflict decision-making, quality audit and correction suggestion generation, forming a closed-loop mechanism of "planning-execution-fusion-audit-feedback-reprocessing".

6. The multi-agent material data mining method with large-scale model semantic orchestration according to claim 1, characterized in that, The material data includes the composition, structure, preparation, performance parameters, and experimental conditions of magnetocaloric materials, magnetic materials, optoelectronic materials, and semiconductor materials, and is primarily applicable to the mining of magnetocaloric performance data such as magnetic entropy change.

7. A multi-agent material data mining system with large-scale model semantic orchestration, applicable to the multi-agent material data mining method with large-scale model semantic orchestration as described in any one of claims 1-6, characterized in that, include: Literature parsing and document representation module: used to parse materials science literature and construct a unified structured document representation; Semantic orchestration module: Based on a large model, it realizes differentiated processing of long and short texts, semantic understanding, task planning and agent scheduling; Multi-agent collaborative processing module: contains multiple domain-specific intelligent agents that perform parallel extraction and mining of material composition, structure, performance, charts, tables, and metadata; Audit feedback module: used for result fusion, confidence assessment, conflict resolution, integrity verification, and feedback correction; Structured output module: Used to generate standardized, database-ready material data results.

8. A multi-agent material data mining system with large-scale model semantic orchestration as described in claim 7, characterized in that, The semantic orchestration module includes an OrchestratorAgent, a dynamic evidence pool, a multi-round planning unit, an uncertainty management unit, and a parallel scheduling unit, supporting adaptive processing of both long and short texts.

9. A multi-agent material data mining system with large-scale model semantic orchestration according to claim 7, characterized in that, The audit feedback module includes MergerAgent, a shared archive building unit, a conflict decision engine, a QA auditor, and a feedback correction unit, forming a closed-loop quality control.

10. A multi-agent material data mining system with large-scale model semantic orchestration according to claim 7, characterized in that, The system supports the addition of new intelligent agents to adapt to new material systems and new performance types, while maintaining the overall process of unified document representation, semantic orchestration, collaborative execution, audit feedback, and structured output.