Large language model special for solid waste cementing material and innovative hypothesis generation method of large language model
By constructing a large language model dedicated to solid waste cementitious materials, the problems of long development cycle and lack of innovation in traditional methods have been solved, efficient and explainable material design and optimization have been achieved, and R&D efficiency and innovation capabilities have been improved.
Patent Information
- Application Number
- CN202510729131.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies have complexity and innovative demands in the development of solid waste cementitious materials. Traditional methods rely on empirical trial mixing, resulting in long development cycles and high costs. In addition, existing large-scale language models lack deep understanding and effective knowledge guidance, making it difficult to build a systematic performance prediction model.
Construct a large language model dedicated to solid waste cementitious materials, including a material knowledge graph module, a vector retrieval module, a hybrid retrieval enhancement generation module and a scientific reasoning module, to achieve knowledge representation, rapid retrieval and explainable innovation generation of multi-dimensional correlation relationships.
Through systematic knowledge representation and efficient data matching, the decision-making efficiency of material design and optimization has been significantly improved, the R&D cycle has been shortened, resource consumption has been reduced, and the explainability and feasibility of innovative solutions have been improved.
Smart Images

Figure CN120632038A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the intersection of artificial intelligence and materials science, and in particular to a large language model dedicated to solid waste cementitious materials and an innovative hypothesis generation method thereof. Background Art
[0002] As an important development direction in the field of green building materials, solid waste cementitious materials present multi-dimensional technical characteristics and challenges in their actual application and development. From the perspective of the raw material system, it involves a variety of industrial solid wastes such as fly ash, slag, and tailings, as well as construction waste. The mineral composition, particle size distribution, and active component content of different raw materials vary significantly, resulting in a diversity of raw material types that greatly increases the complexity of material design. In terms of the mechanism of action of stimulants, whether chemical stimulants (such as alkali metal compounds, sulfates, etc.) or physical stimulants (mechanical grinding, heat treatment) and composite stimulants, the mechanism of their influence on the activation path, hydration reaction kinetics, and product structure evolution of different solid waste active components is extremely complex, and the synergistic and antagonistic effects between the various stimulant factors have not yet been systematically understood.
[0003] At the process parameter level, there is a strong coupling between parameters such as grinding fineness, curing temperature, humidity, age, and mix design. Small adjustments to one parameter often trigger chain reactions in other parameters, making it difficult for traditional single-factor testing methods to accurately reveal the optimal parameter combination. Regarding performance evolution, performance indicators such as the mechanical properties (compressive strength, flexural strength), durability (impermeability, frost resistance), and volume stability of solid waste cementitious materials exhibit nonlinear evolution with changes in raw material composition and process conditions. The inherent action pathways of these performance indicators remain unclear due to the interweaving of multiple factors.
[0004] Current development methods for solid waste cementitious materials still rely heavily on empirical trial-and-error formulations, requiring technicians to conduct extensive repetitive experiments to screen formulations and verify performance. This model not only results in long product development cycles of months or even years, but also in high test material consumption and equipment operating costs. Furthermore, this overreliance on historical experience makes it difficult to break through the limitations of traditional formulation systems, resulting in a significant lack of innovation in the industry.
[0005] In the field of knowledge management and data analysis, traditional retrieval-based knowledge management systems can only achieve simple storage and query of literature and formula data, and cannot effectively explore the deep connections between material composition, excitation reaction processes, process conditions and mechanical properties. Isolated data analysis methods (such as univariate statistics and simple regression models) are unable to handle the complex coupling relationships between multiple variables, making it difficult to build systematic performance prediction models. When applied to the materials field, existing large-scale language models lack a deep understanding of professional knowledge of solid waste cementitious materials (such as mineral phase transformation laws and microstructural characterization of hydration products). When dealing with domain problems, they suffer from semantic parsing biases and logical reasoning breaks. Their black-box prediction results make it difficult to provide an explainable design basis, and the innovative design process lacks effective knowledge guidance and solution generation capabilities, which cannot meet the actual needs of the innovative development of solid waste cementitious materials.
[0006] In summary, existing technologies have obvious technical bottlenecks in dealing with the complexity and innovative needs of solid waste cementitious materials development. There is an urgent need to build a new system that can integrate domain knowledge and reasoning capabilities. Through multi-source data integration, complex relationship modeling and explainable reasoning, rapid design and innovative exploration of solid waste cementitious materials can be achieved, breaking through the dual constraints of efficiency and innovation of traditional development models. Summary of the Invention
[0007] The present invention provides a large language model for solid waste cementitious materials based on hybrid knowledge enhancement and scientific reasoning and an innovative hypothesis generation method thereof, which is used to solve the four core problems in the research and development of solid waste cementitious materials: knowledge structured expression, experimental data semantic retrieval, cross-modal reasoning support and explainable innovation generation.
[0008] The technical solution adopted by the present invention is: a large language model dedicated to solid waste cementitious materials, including the following collaborative modules:
[0009] The material knowledge graph module is used to construct and store a structured knowledge network in the field of solid waste cementitious materials. The knowledge network contains multi-dimensional correlations between raw material composition, activator ratio, curing conditions, microstructural characteristics, and mechanical performance indicators.
[0010] Vector retrieval module, used to embed material experimental ratio schemes, process parameters and performance test results into high-dimensional modeling, and can quickly retrieve historical experimental cases based on similarity matching;
[0011] A hybrid search enhancement generation module is used to integrate the logical reasoning results of the material knowledge graph module with the similar case matching results of the vector search module to dynamically generate highly relevant prompt information for the large language model to call;
[0012] The scientific reasoning module is used to generate material design rules, performance prediction models and innovative hypotheses through rule synthesis and data inference, and output visual reasoning chains and explanatory analysis reports.
[0013] Furthermore, the material knowledge graph module is constructed through the following steps:
[0014] Step 1-1, using natural language processing technology to automatically extract entity information from field literature, experimental data and technical standards, the entity including raw material type, chemical composition ratio, activator system, curing process parameters and performance indicators;
[0015] Steps 1-2: Through expert review and manual annotation correction, semantic consistency correction and terminology standardization are completed, ultimately constructing a multi-level heterogeneous knowledge network, in which each node covers material composition, excitation mechanism, and performance. The edge relationships between nodes are used to characterize the correlation between parameters and their evolution mechanism with changing conditions.
[0016] In steps 1-3, the number of triples extracted from the material knowledge graph is estimated using the following formula:
[0017]
[0018] Where: N is the number of extracted knowledge triples; E is the number of identified entity pairs; R is the number of possible relationship types between entities.
[0019] Furthermore, the vector retrieval module performs the following operations:
[0020] Step 2-1: Use the domain-tuned embedding model to convert the ratio parameters, process conditions, and performance data into high-dimensional vectors;
[0021] Step 2-2: Build a search index based on the approximate nearest neighbor algorithm and use the cosine similarity formula:
[0022]
[0023] Or Euclidean distance measures sample similarity;
[0024] Where: 、 A high-dimensional vector representation of the two experimental ratio samples; 、 is the modulus of the corresponding vector; is the vector angle; The closer it is to 1, the more similar the two experimental cases are.
[0025] Furthermore, the workflow of the hybrid search enhancement generation module includes:
[0026] Step 3-1: Screen candidate material systems and ratio combinations that meet the input target requirements based on material knowledge graph reasoning;
[0027] Step 3-2, further matching historical experimental cases in the screening set through vector retrieval method;
[0028] Step 3-3: Dynamically integrate the reasoning results and the search content through the prompt word optimization strategy to generate prompt information that meets the domain knowledge logic and context requirements.
[0029] Furthermore, the scientific reasoning module includes:
[0030] The rule synthesis submodule derives material design and performance prediction rules based on material science principles and knowledge graph facts;
[0031] The data inference submodule constructs a parameter-performance mapping relationship by summarizing experimental data:
[0032]
[0033] Where: P is the performance index; X is the ratio parameter vector; is a nonlinear mapping function; is the error term.
[0034] The present invention also provides an innovative hypothesis method for generating solid waste cementitious materials based on the above-mentioned large language model, comprising the steps of:
[0035] Step a, receiving material performance targets, application scenarios or optimization requirements input by the user;
[0036] Step b: Screen candidate material systems and stimulant solutions that meet the target conditions through material knowledge graph reasoning;
[0037] Step c: searching and matching similar historical cases and their matching parameters and performance data based on vectors;
[0038] Step d: Integrate the graph reasoning chain with the retrieval case features to generate possible matching innovation paths or new process hypotheses;
[0039] Step e: Output the innovation hypothesis text containing the recommended ratio, stimulant composition, expected performance, and attached reasoning chain and explainability analysis report.
[0040] The beneficial effects of the present invention are:
[0041] 1. A knowledge graph covering multi-dimensional elements such as the composition of covering materials, activator mechanism, process parameters, and mechanical properties has been constructed to systematically represent professional knowledge in the field of solid waste cementitious materials, support logical reasoning and causal chain analysis, and solve the problems of traditional data isolation and knowledge fragmentation;
[0042] 2. Through vectorized embedding and approximate nearest neighbor retrieval mechanisms, it is possible to quickly and accurately match highly relevant experimental cases in large-scale ratio-performance datasets, significantly reducing manual screening costs and improving decision-making efficiency in material system design and optimization;
[0043] 3. By integrating knowledge graph reasoning and screening with vector semantic refinement retrieval, it is possible to generate content with high contextual relevance and domain specificity, overcoming the problems of ambiguous content and logical discontinuities in traditional search-enhanced generation, and significantly improving the accuracy and professionalism of system responses.
[0044] 4. The scientific reasoning module automatically derives the causal relationship chain between material ratios and performance, generating reasonable new ratio schemes or process path hypotheses for specified target performance or application scenarios, and outputting the complete derivation process and key supporting evidence, thereby improving the traceability, explainability and engineering feasibility of innovative solutions.
[0045] 5. It effectively shortens the R&D cycle of solid waste cementitious materials from preliminary design to performance optimization, reduces the number of experiments and resource input, and is suitable for the fields of green building materials, sustainable building materials and high-value utilization of solid waste resources, with significant economic benefits and social value. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Schematic diagram of the system structure of the present invention;
[0047] Figure 2 This is the combined workflow of knowledge enhancement and scientific reasoning in the present invention. DETAILED DESCRIPTION
[0048] The present invention will be further described below with reference to the accompanying drawings.
[0049] like Figure 1 and Figure 2 As shown, the present invention is a large language model dedicated to solid waste cementitious materials, including a material knowledge graph module, a vector retrieval module, a hybrid retrieval enhancement generation module and a scientific reasoning module that work together.
[0050] The material knowledge graph module is used to construct and store a structured knowledge network in the field of solid waste cementitious materials. This knowledge network contains the multi-dimensional correlation between raw material composition, activator ratio, curing conditions, microstructure characteristics, and mechanical performance indicators, and is represented in the form of a structured graph. The construction process is as follows:
[0051] Step 1-1: Use natural language processing technology to automatically mine field literature, experimental data, and technical standards to extract information such as raw material types, chemical composition ratios, activator systems, curing process parameters, and performance results;
[0052] Steps 1-2: Through expert review and manual annotation correction, semantic consistency correction and terminology standardization are completed, ultimately constructing a multi-level heterogeneous knowledge network, in which each node covers material composition, excitation mechanism, and performance. The edge relationships between nodes are used to characterize the correlation between parameters and their evolution mechanism with changing conditions.
[0053] In steps 1-3, the number of triples extracted from the material knowledge graph is estimated using the following formula:
[0054]
[0055] Where: N is the number of extracted knowledge triples; E is the number of identified entity pairs; R is the number of possible relationship types between entities.
[0056] The vector retrieval module is used to embed material experimental ratio schemes, process parameters, and performance test results into high-dimensional models, and can quickly retrieve historical experimental cases based on similarity matching. The process includes the following steps:
[0057] Step 2-1: Use the domain-tuned embedding model to convert the ratio parameters, process conditions, and performance data into high-dimensional vectors;
[0058] Step 2-2: Build a search index based on the approximate nearest neighbor algorithm and use the cosine similarity formula:
[0059]
[0060] Or Euclidean distance measures sample similarity;
[0061] Where: 、 A high-dimensional vector representation of the two experimental ratio samples; 、 is the modulus of the corresponding vector; is the vector angle; The closer it is to 1, the more similar the two experimental cases are.
[0062] The hybrid retrieval enhancement generation module is used to integrate the logical reasoning results of the material knowledge graph module with the similar case matching results of the vector retrieval module, dynamically generating highly relevant prompt information for the large language model to call. Its workflow includes:
[0063] Step 3-1: Screen candidate material systems and ratio combinations that meet the input target requirements based on material knowledge graph reasoning;
[0064] Step 3-2, further matching historical experimental cases in the screening set through vector retrieval method;
[0065] Step 3-3: Dynamically integrate the reasoning results and the search content through the prompt word optimization strategy to generate prompt information that meets the domain knowledge logic and context requirements.
[0066] The scientific reasoning module is used to generate material design rules, performance prediction models, and innovative hypotheses through rule synthesis and data inference, and outputs a visual reasoning chain and explanatory analysis report. The scientific reasoning module includes a rule synthesis submodule and a data inference submodule; the rule synthesis submodule derives material design and performance prediction rules based on material science principles and knowledge graph facts; the data inference submodule constructs parameter-performance mapping relationships by summarizing experimental data:
[0067]
[0068] Where: P is the performance index; X is the ratio parameter vector; is a nonlinear mapping function; is the error term.
[0069] An innovative hypothesis method for generating solid waste cementitious materials based on the above-mentioned large language model includes the following steps:
[0070] Step a, receiving material performance targets, application scenarios or optimization requirements input by the user;
[0071] Step b: Screen candidate material systems and stimulant solutions that meet the target conditions through material knowledge graph reasoning;
[0072] Step c: searching and matching similar historical cases and their matching parameters and performance data based on vectors;
[0073] Step d: Integrate the graph reasoning chain with the retrieval case features to generate possible matching innovation paths or new process hypotheses;
[0074] Step e: Output the innovation hypothesis text containing the recommended ratio, stimulant composition, expected performance, and attached reasoning chain and explainability analysis report.
[0075] The present invention proposes automatic construction and reasoning technology of material knowledge graphs, covering the composition-mechanism-process-performance relationship; domain-specific vector retrieval and approximate neighborhood matching methods to improve retrieval accuracy; hybrid knowledge-data retrieval enhancement generation mechanism to optimize large-scale language model input; scientific reasoning and deduction chain tracking module to support ratio optimization and innovative hypothesis generation.
[0076] It solves the problems of systematic structured expression and reasonable modeling of domain knowledge; efficient and accurate semantic retrieval of experimental data; cross-knowledge-data hybrid reasoning to support material design optimization; and automatic generation of explainable innovation hypotheses. It significantly improves the R&D efficiency and innovation output capacity of solid waste cementitious materials.
[0077] The following is an embodiment of the present invention. In this embodiment, the dedicated large language model proposed by the present invention is used to predict the performance and optimize the mix ratio of a geopolymer concrete system reinforced with a mixture of nano-SiO2 and steel fiber-polyvinyl alcohol fiber (PVA).
[0078] The input goal is "Ratio optimization suggestions for improving both rheological properties and compressive strength". The specific operation process is as follows:
[0079] 1. Input parameters:
[0080] Raw material type: fly ash (FA)
[0081] Target performance: 28d compressive strength ≥45MPa, plastic viscosity ≤0.7Pa·s, yield stress ≤80Pa.
[0082] Usage restrictions: Nano-SiO2 content ≤ 2.0%, steel fiber ≤ 0.5%, PVA fiber ≤ 0.3%.
[0083] 2. Knowledge graph reasoning results:
[0084] According to the existing map structure, it is found that nano-SiO2 can significantly enhance the gel formation rate and improve the compressive strength when it is 0.5% to 1.5%;
[0085] The hybrid fiber system (0.25% steel fiber, 0.2% PVA) can improve crack control ability and strength development without significantly reducing fluidity.
[0086] 3. Vector search matching results:
[0087] The system's vector search module retrieved a historical experimental sample with a highly similar mix ratio to the recommended mix ratio (0.5% nano-SiO2, 0.25% steel fiber by volume, and 0.2% PVA fiber by volume). Performance testing of this sample demonstrated a 28-day compressive strength of 49.1 MPa, a yield stress of 75 Pa, and a plastic viscosity of 0.64 Pa·s, meeting the established strength and workability requirements and verifying the feasibility and rationality of the recommended mix ratio.
[0088] The nonlinear fitting model showed that a better balance could be achieved if the comprehensive value of RI (Reinforcing Index) was controlled within the range of 0.35-0.5.
[0089] 4. Model generation results:
[0090] Recommended ratio: nano-SiO2: 0.5%; steel fiber: 0.25%; PVA fiber: 0.2%;
[0091] Activator: water glass modulus is 3.2, NaOH concentration is 7.7M;
[0092] The ratio scheme is N0.5S0.25P0.2;
[0093] Predicted performance: 28d compressive strength is 48.6MPa, yield stress is 72Pa, and plastic viscosity is 0.67Pa·s.
[0094] 5. Explainable reasoning chain output:
[0095] Graph chain: "Nano-SiO2 addition → improve interface density → improve early strength";
[0096] “The steel fiber length / diameter ratio is 62 → delays crack growth → enhances toughness”;
[0097] Vector space similarity: 0.93, highly matching the best historical sample;
[0098] The error between the generated recommended ratio and the target performance is <5%.
Claims
1. A large language model dedicated to solid waste cementitious materials, characterized by: Includes the following modules that work together: The material knowledge graph module is used to construct and store a structured knowledge network in the field of solid waste cementitious materials. The knowledge network contains multi-dimensional correlations between raw material composition, activator ratio, curing conditions, microstructural characteristics, and mechanical performance indicators. Vector retrieval module, used to embed material experimental ratio schemes, process parameters and performance test results into high-dimensional modeling, and can quickly retrieve historical experimental cases based on similarity matching; A hybrid search enhancement generation module is used to integrate the logical reasoning results of the material knowledge graph module with the similar case matching results of the vector search module to dynamically generate highly relevant prompt information for the large language model to call; The scientific reasoning module is used to generate material design rules, performance prediction models and innovative hypotheses through rule synthesis and data inference, and output visual reasoning chains and explanatory analysis reports.
2. The large language model for solid waste cementitious materials according to claim 1, characterized in that: The material knowledge graph module is constructed through the following steps: Step 1-1, using natural language processing technology to automatically extract entity information from field literature, experimental data and technical standards, the entity including raw material type, chemical composition ratio, activator system, curing process parameters and performance indicators; Steps 1-2: Through expert review and manual annotation correction, semantic consistency correction and terminology standardization are completed, ultimately constructing a multi-level heterogeneous knowledge network, in which each node covers material composition, excitation mechanism, and performance. The edge relationships between nodes are used to characterize the correlation between parameters and their evolution mechanism with changing conditions. In steps 1-3, the number of triples extracted from the material knowledge graph is estimated using the following formula: Where: N is the number of extracted knowledge triples; E is the number of identified entity pairs; R is the number of possible relationship types between entities.
3. The large language model for solid waste cementitious materials according to claim 1, characterized in that: The vector retrieval module performs the following operations: Step 2-1: Use the domain-tuned embedding model to convert the ratio parameters, process conditions, and performance data into high-dimensional vectors; Step 2-2: Build a search index based on the approximate nearest neighbor algorithm and use the cosine similarity formula: Or Euclidean distance measures sample similarity; Where: 、 A high-dimensional vector representation of the two experimental ratio samples; 、 is the modulus of the corresponding vector; is the vector angle; The closer it is to 1, the more similar the two experimental cases are.
4. The large language model for solid waste cementitious materials according to claim 1, characterized in that: The workflow of the hybrid search enhancement generation module includes: Step 3-1: Screen candidate material systems and ratio combinations that meet the input target requirements based on material knowledge graph reasoning; Step 3-2, further matching historical experimental cases in the screening set through vector retrieval method; Step 3-3: Dynamically integrate the reasoning results and the search content through the prompt word optimization strategy to generate prompt information that meets the domain knowledge logic and context requirements.
5. The large language model for solid waste cementitious materials according to claim 1, characterized in that: The scientific reasoning module includes: The rule synthesis submodule derives material design and performance prediction rules based on material science principles and knowledge graph facts; The data inference submodule constructs a parameter-performance mapping relationship by summarizing experimental data: Where: P is the performance index; X is the ratio parameter vector; is a nonlinear mapping function; is the error term.
6. An innovative hypothesis method for generating solid waste cementitious materials using the model according to any one of claims 1 to 5, characterized in that: Including steps: Step a, receiving material performance targets, application scenarios or optimization requirements input by the user; Step b: Screen candidate material systems and stimulant solutions that meet the target conditions through material knowledge graph reasoning; Step c: searching and matching similar historical cases and their matching parameters and performance data based on vectors; Step d: Integrate the graph reasoning chain with the retrieval case features to generate possible matching innovation paths or new process hypotheses; Step e: Output the innovation hypothesis text containing the recommended ratio, stimulant composition, expected performance, and attached reasoning chain and explainability analysis report.
Citation Information
Cited By
Solid waste high-valued nuclear engineering concrete material knowledge base establishment method and system
CN121303299A
Method and system for establishing solid waste high-value nuclear engineering concrete material knowledge base
CN121303299B
Ultra-high performance concrete knowledge graph construction and graph retrieval enhanced generation-based mechanism interpretation method
CN121390245A
Physical simulation test similar material reverse design exploration and discovery method and system
CN121687298A
A method and system for reverse engineering and discovery of similar materials through physical simulation experiments
CN121687298B