Solid waste high-valued nuclear engineering concrete material knowledge base establishment method and system
By establishing a knowledge base for nuclear engineering concrete materials with high-value solid waste, and utilizing intelligent retrieval and machine learning models, the problem of insufficient data integration in existing technologies has been solved. This has enabled the efficient and accurate generation of nuclear engineering concrete material formulas, meeting the needs of rapid SMR construction and improving performance stability.
Patent Information
- Application Number
- CN202511854401.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-12-10
AI Technical Summary
The existing solid waste-based nuclear engineering concrete technology lacks systematic data integration and quantitative model guidance, resulting in long research and development cycles, difficulty in adapting to the rapid construction needs of SMR, and inability to accurately control the proportion of solid waste admixture and performance stability, affecting radiation shielding and strength reliability.
A knowledge base for nuclear engineering concrete materials with high-value solid waste was established. External data was collected through a standardized retrieval terminology database and an intelligent retrieval engine. Combined with machine learning models and gene coding models, the correlation between solid waste composition, proportion, process, and performance was explored. Optimization schemes were automatically generated and reverse-engineered, and finally, a knowledge graph was constructed to store the data.
It has achieved data-driven intelligent generation of nuclear engineering concrete material formulas, which has improved R&D efficiency and accuracy, adapted to the rapid construction needs of SMR, and ensured the stability of radiation shielding and strength.
Smart Images

Figure CN121303299A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of nuclear power engineering technology, and in particular to a method and system for establishing a knowledge base for nuclear engineering concrete materials with high-value solid waste. Background Technology
[0002] The global nuclear power industry is accelerating its transition to small modular reactors (SMRs). Nuclear engineering concrete, as a core structural material for nuclear facilities, accounts for over 30% of the total carbon emissions throughout the nuclear power plant's lifecycle, contradicting the environmentally friendly goals of SMRs. Using industrial solid wastes such as steel slag and fly ash in the preparation of nuclear engineering concrete can reduce pollution from solid waste storage and the consumption of natural raw materials, while also ensuring radiation shielding performance, making it a key pathway to promoting the greening of nuclear engineering.
[0003] Existing solid waste-based nuclear engineering concrete technologies mostly revolve around "solid waste replacing traditional components," such as replacing fine aggregates with steel slag, adding waste glass to barite concrete, or replacing coarse aggregates with iron tailings and using fly ash as a cementing material, in order to achieve solid waste resource utilization and optimize concrete performance.
[0004] The core limitation of existing technologies lies in the fact that the research and development process relies heavily on empirical fixed dosage trial and error, and lacks systematic data integration and quantitative model guidance on the relationship between solid waste composition, proportion, process and performance. This results in not only a long research and development cycle, making it difficult to adapt to the rapid construction needs of SMR, but also an inability to accurately control the proportion of solid waste and performance stability. It is difficult to break through the conventional upper limit of 25% solid waste dosage, and it is also easy to affect the radiation shielding and strength reliability of nuclear engineering concrete due to parameter matching deviations. Summary of the Invention
[0005] In order to achieve a leap from relying on experience-based trial and error to data-driven intelligent generation in nuclear engineering concrete material formulation, and to realize precise and efficient optimal design, this application provides a method and system for establishing a knowledge base of nuclear engineering concrete materials with high value from solid waste.
[0006] Firstly, this application provides a method for establishing a knowledge base for nuclear engineering concrete materials in the high-value utilization of solid waste, employing the following technical solution:
[0007] A method for establishing a knowledge base for nuclear engineering concrete materials through high-value utilization of solid waste includes:
[0008] External data is collected through a pre-set standardized search terminology database and intelligent search engine, while internal data is collected simultaneously to complete the integration and classification of dual-source heterogeneous data;
[0009] Based on the preset tree-structured four-level parameter architecture, the integrated and classified data is cleaned in a structured manner, classified by parameter category, subcategory and indicator level, outliers are removed and data units are unified to form a standardized dataset.
[0010] The standardized dataset is imported into a pre-defined database framework to build an initial database. A pre-defined machine learning model is used to learn the correlation between components, processes, and performance. Supplementary samples are generated to expand the data and obtain a complete database.
[0011] By using a pre-defined gene coding model to mine and improve the structure-activity relationship of solid waste components, ratios, processes, and performance in the database, a quantitative structure-activity relationship between performance and solid waste components, ratios, and processes is established.
[0012] Input the target performance requirements of nuclear engineering concrete, and based on the above-mentioned quantitative structure-property relationship and preset derivation constraints, automatically generate multiple sets of initial parameter combinations and perform reverse derivation to generate multiple initial optimization schemes;
[0013] Substitute each initial optimization scheme into the quantitative structure-property relationship to verify the performance compliance: if it meets the standard, it is included in the candidate optimization scheme set; if it does not meet the standard, the parameters are corrected according to the preset adjustment rules and the derivation is repeated until a preset number of candidate schemes are generated.
[0014] Based on a preset screening index system, a preset weighted scoring model is used to score the candidate solution set, and the solution with the highest score is selected as the optimal solution.
[0015] The database, structure-function relationship, candidate solution set, optimal solution and iteration process data are linked and stored according to the preset knowledge graph structure to build a knowledge base including data layer, model layer and solution layer.
[0016] Secondly, this application provides a knowledge base establishment system for nuclear engineering concrete materials in the high-value utilization of solid waste, which adopts the following technical solution:
[0017] A system for establishing a knowledge base of nuclear engineering concrete materials with high value utilization from solid waste includes a memory, a processor, and a program stored in the memory and executable on the processor. When the program is loaded and executed by the processor, it implements the method for establishing a knowledge base of nuclear engineering concrete materials with high value utilization from solid waste as described in the first aspect. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the overall process of establishing a knowledge base for nuclear engineering concrete materials for high-value utilization of solid waste, according to an embodiment of this application. Detailed Implementation
[0019] The present application will be further described in detail below with reference to the accompanying drawings.
[0020] Reference Figure 1 This application discloses a method for establishing a knowledge base for nuclear engineering concrete materials through the high-value utilization of solid waste, comprising:
[0021] Step S100: Collect external data through a pre-set standardized search terminology database and intelligent search engine, and simultaneously collect internal data to complete the integration and classification of dual-source heterogeneous data.
[0022] The database constructed in this invention and all its data processing procedures are based on publicly available academic literature, patent data, and internal material experiment and production data in this field. This data does not involve any personally identifiable information or sensitive personal information, and its collection, storage, and use strictly comply with relevant regulations, and will not violate social ethics or harm the public interest.
[0023] External data refers to data obtained through publicly available literature, patents, industry reports, etc., and is typically stored in formats such as PDF and Excel. These data sources are broad, covering research findings and practical experience within the industry. Internal data refers to data directly collected during experiments and production processes, and is typically stored in formats such as Origin and Word. This data includes solid waste property parameters (such as elemental abundance), process parameters (such as stirring rate), and characterization data (such as gamma-ray attenuation coefficient).
[0024] Standardized Search Terminology Database: A search terminology database built based on national standards and industry terminology, used for accurate collection of external data. For example, a keyword system is constructed with "GB / T5072-2022 Nuclear Power Concrete" as the root node, expanding downwards to specific parameters such as "steel slag - moisture content" and "fly ash - loss on ignition".
[0025] Intelligent search engine: A search tool that combines Boolean logic and semantic association, capable of automatically identifying and supplementing synonyms, improving search efficiency and accuracy. For example, using Python-Whoosh combined with a cosine similarity algorithm, it can achieve joint searches such as "steel slag OR converter slag".
[0026] The acquisition method and process are described below:
[0027] 1. External Data Acquisition: 1.1 Input Search Query: Construct search queries based on industry standards and specific needs, such as "("smallmodularreactor" OR "nuclear containment structure") AND ("solid waste" OR "steel slag" OR "fly ash") AND ("gamma attenuation coefficient" OR "shielding")". 1.2 Search Engine: Use Python-Whoosh combined with semantic association algorithms to automatically parse key parameters from documents, patents, and industry reports, outputting standardized Excel templates. For example, incrementally crawl data daily from databases such as CNKI, USPTO, and IAEA-TECDOC, parsing DOI, abstracts, and key parameters. 1.3 Data Formatting: Format the collected external data into uniform Excel spreadsheets, aligning fields with the four-level parameter architecture of the internal data to ensure data consistency and operability.
[0028] 2. Internal Data Acquisition: 2.1 Data Sources: Data is collected from experimental and production sites, including solid waste property parameters, process parameters, and characterization data. For example, stirring rate and temperature data are uploaded every 30 seconds via PLC, and the laboratory LIMS automatically transmits XRF elemental abundance and gamma-ray spectrometer attenuation values. 2.2 Data Storage: The collected data is stored in a local MongoDB database for easy subsequent processing and analysis.
[0029] 3. Dual-Source Data Alignment: 3.1 Primary Key Matching: Using "Solid Waste Category + Chemical Composition" as the joint primary key, fuzzy matching is performed using the Levenshtein distance algorithm to merge external and internal data into records with the same feature. For example, "Steel Slag-MnO 8.2%" from the external data is merged with "Steel Slag-MnO 8.0%" from the internal data. 3.2 Initial Quality Screening: Box plot rules are applied to the merged data to remove outliers. For example, if the γ decay coefficient exceeds Q3+1.5IQR, it is marked as an outlier and processed uniformly in subsequent steps.
[0030] Step S200: Based on the preset tree-structured four-level parameter architecture, the integrated and classified data is cleaned in a structured manner, classified by parameter category, subcategory and indicator level, outliers are removed and data units are unified to form a standardized dataset.
[0031] Tree-structured four-level parameter architecture: A hierarchical data classification method that divides data into major categories such as solid waste materials, physical properties, chemical properties, and process parameters. Each major category is further subdivided into subcategories and specific indicators. For example, solid waste materials are divided into subcategories such as steel slag and fly ash, and steel slag is further divided into indicators such as physical properties (e.g., moisture content) and chemical properties (e.g., MnO content). Structured cleaning: Standardizes the collected data, including unifying units, removing outliers, and filling missing values, to ensure the data conforms to preset format and quality requirements. Feature deconstruction: Extracts useful feature information for model building from the raw data, such as extracting key parameters affecting concrete performance from the physical and chemical properties of solid waste materials.
[0032] Explanation of acquisition methods and processes:
[0033] 1. Data Classification and Hierarchical Construction: 1.1 Classification Basis: Based on the type and properties of solid waste materials, and combined with the actual needs of nuclear power plant concrete, a tree-like four-level parameter architecture is constructed. For example, solid waste materials are classified into steel slag, fly ash, iron tailings, etc.; physical properties include moisture content, density, etc.; chemical properties include elemental abundance (such as MnO, SiO2 content), etc. 1.2 Hierarchical Construction: The data is divided into four levels: solid waste materials (major category), physical / chemical properties (sub-category), specific indicators (such as moisture content, MnO content), and data values. For example, steel slag (major category) → physical properties (sub-category) → moisture content (indicator) → 6.2% (data value).
[0034] 2. Structured Data Cleaning: 2.1 Unit Standardization: Standardize all data units to a uniform standard, such as density to g / cm³ and moisture content to percentage (%). For example, convert density data from different sources from kg / m³ to g / cm³. 2.2 Outlier Removal: Use box plots to detect and remove outliers. For example, for loss on ignition data, if the loss on ignition of a batch of steel slag exceeds 5%, mark it as an outlier and remove it. 2.3 Missing Value Handling: For missing data, use the mean or median to impute it. For example, if the loss on ignition data for a batch of fly ash is missing, use the median of the known data to impute it.
[0035] 3. Feature Deconstruction: 3.1 Key Feature Extraction: Extract key features that significantly affect concrete performance from the physical and chemical properties of solid waste materials. For example, extract features such as moisture content and MnO content from steel slag, and features such as loss on ignition and SiO2 content from fly ash. 3.2 Feature Correlation: Establish correlations between features, such as the correlation between moisture content and compressive strength, and the correlation between MnO content and gamma-ray attenuation coefficient. For example, correlation analysis revealed that for every 1% increase in moisture content of steel slag, the compressive strength decreases by 2%.
[0036] Step S300: Import the standardized dataset into the preset database framework to build the initial database, use the preset machine learning model to learn the correlation between components, processes and performance, generate supplementary samples to expand the data, and obtain a complete database.
[0037] The dataset comprises the following components: Initial Database: A preliminary database built upon a standardized dataset, containing information on the physical and chemical properties, and process parameters of solid waste materials. Machine Learning Model: Used to learn the correlation between the composition of solid waste materials, process parameters, and concrete performance, and to generate supplementary samples to expand the dataset. Data Expansion: New sample data is generated through the machine learning model to increase the diversity and coverage of the dataset, thereby improving the model's generalization ability. The specific processes described above can be found in steps S310 to S340, and will not be elaborated upon here.
[0038] Step S400: By using a preset gene coding model to mine and improve the structure-activity relationship of solid waste components-proportion-process-performance in the database, a quantitative structure-activity relationship between performance and solid waste components, proportion, and process is established.
[0039] Gene-encoded model: A mathematical model used to analyze the structure-property relationship between solid waste composition, mix proportions, process parameters, and concrete performance. It establishes a quantitative structure-property relationship between input parameters (solid waste composition, mix proportions, process) and output properties (such as compressive strength and gamma-ray attenuation coefficient) through quantitative analysis. Structure-property relationship: Describes the interaction and influence between solid waste composition, mix proportions, process parameters, and concrete performance. For example, how does the MnO content of steel slag affect the compressive strength of concrete, and how does the SiO2 content of fly ash affect the gamma-ray attenuation coefficient?
[0040] The specific process of the above steps can be referred to in steps S410 to S450, and will not be repeated here.
[0041] Step S500: Input the target performance requirements of nuclear engineering concrete. Based on the above-mentioned quantitative structure-property relationship and preset derivation constraints, automatically generate multiple sets of initial parameter combinations and perform reverse derivation to generate multiple initial optimization schemes.
[0042] The process includes: Target performance requirements: Based on the specific application scenarios of nuclear engineering concrete, defined concrete performance indicators, such as compressive strength, gamma-ray attenuation coefficient, and freeze-thaw cycle resistance. Reverse derivation: Based on the target performance requirements, using an established gene-coding model, the solid waste composition, mix proportions, and process parameters that meet these performance requirements are derived from the performance indicators. Initial optimization schemes: Preliminary schemes generated through reverse derivation; these schemes require further verification and optimization to meet practical application needs. The specific processes described above can be found in steps S510 to S550, and will not be elaborated here.
[0043] Step S600: Substitute each initial optimization scheme into the quantitative structure-property relationship to verify the performance compliance: if it meets the standard, it is included in the candidate optimization scheme set; if it does not meet the standard, the parameters are corrected according to the preset adjustment rules and the derivation is repeated until a preset number of candidate schemes are generated.
[0044] Performance verification involves verifying, through experiments or simulations, whether the initial optimized scheme meets the target performance requirements. Quantifying structure-property relationships (SPRs) is a mathematical model—using a pre-defined gene-coding model—to analyze and mine a comprehensive database, identifying quantifiable correlations between solid waste components, mix proportions, process parameters, and concrete performance indicators. It mathematically represents the interaction between parameters and performance, thus possessing computability. Technically, "quantified structure-property relationship verification" involves calculating the corresponding performance indicator estimates for any given scheme based on this quantified correlation.
[0045] The specific process is as follows:
[0046] 1. Performance prediction based on quantitative structure-property relationship: The specific parameters (solid waste composition, ratio, process parameters) of each initial optimization scheme are calculated according to the quantitative structure-property relationship to obtain the estimated value of its key performance indicators (such as compressive strength and gamma-ray attenuation coefficient).
[0047] 2. Empirical verification through experiments and / or simulations:
[0048] Experimental verification: For schemes predicted to meet standards, concrete specimens were prepared based on their parameters, and performance indicators were measured. For example, the compressive strength of a specimen from a certain scheme was 47 MPa, and the gamma-ray attenuation coefficient was 0.53 / cm.
[0049] Simulation verification: Numerical simulation methods such as finite element analysis can be used to predict and verify the performance of the scheme. For example, software simulation can yield a predicted compressive strength of 46 MPa and a predicted gamma-ray attenuation coefficient of 0.54 / cm for a certain scheme.
[0050] 3. Comprehensive Compliance Assessment: The performance estimates based on the quantified structure-property relationship, experimental test results, and / or simulation prediction results are comprehensively compared with the target performance requirements (e.g., compressive strength ≥ 45 MPa, gamma-ray attenuation coefficient ≥ 0.55 / cm). If the solution is verified through one or more of the above methods, and the results all meet or exceed the target thresholds, the solution is deemed to meet the requirements and is directly included in the candidate optimization solution set; otherwise, it is deemed not to meet the requirements and enters the parameter adjustment stage.
[0051] 4. Parameter adjustment and recalculation verification:
[0052] Application of adjustment rules: Based on preset adjustment rules for substandard performance indicators, specific parameters of the scheme are modified. For example, if the gamma-ray attenuation coefficient does not meet the standard, the amount of barite powder is increased according to the rules.
[0053] Adjustment methods: Optimization can be achieved by step-by-step adjustment (adjusting one parameter at a time) or comprehensive adjustment (adjusting multiple parameters simultaneously).
[0054] Recalculation and verification: The adjusted new parameter combination is recalculated based on the quantitative structure-property relationship to obtain a new performance estimate. The empirical verification and comprehensive judgment process in steps 2 to 3 is repeated until the performance meets the standard.
[0055] 5. Generate a candidate optimization scheme set: Summarize all validated compliant schemes (including those that initially meet the standards and those that have been adjusted to meet the standards) to form a candidate optimization scheme set. Completely record the solid waste composition, proportions, process parameters, and corresponding performance estimates, experimental and / or simulation validation data for each candidate scheme.
[0056] Step S700: Based on the preset screening index system, the candidate solution set is scored using a preset weighted scoring model, and the solution with the highest score is selected as the optimal solution.
[0057] The weighted scoring model is a mathematical model that calculates the comprehensive score of candidate solutions by assigning weights to each scoring indicator. The comprehensive score reflects the overall performance and applicability of the solution. The optimal solution is the one with the highest comprehensive score among all candidate solutions, which is considered to best meet the performance requirements of nuclear engineering concrete.
[0058] In step S700, we comprehensively evaluate the candidate optimization schemes based on a comprehensive scoring index system to determine the optimal scheme. The specific process is as follows: First, a comprehensive scoring index system is established based on the performance requirements and actual application conditions of nuclear engineering concrete. For example, compressive strength has a weight of 30%, gamma-ray attenuation coefficient has a weight of 25%, cost has a weight of 20%, durability has a weight of 15%, and construction convenience has a weight of 10%. These weights are determined based on actual engineering needs and expert experience, ensuring that the importance of each index is reasonably reflected in the overall evaluation. Next, the performance index data of each candidate scheme are standardized to eliminate the influence of different dimensions and orders of magnitude. For example, the formula: Standardized Score = (Actual Value - Minimum Value) ÷ (Maximum Value - Minimum Value) is used to convert indices such as compressive strength and gamma-ray attenuation coefficient into standardized scores. This process ensures the comparability between different indices and provides a basis for subsequent weighted scoring.
[0059] Then, based on preset weights, the standardized performance indicators of each candidate scheme are weighted and summed to calculate the comprehensive score. For example, if a scheme has a standardized score of 0.9 for compressive strength, 0.85 for gamma-ray attenuation coefficient, 0.7 for cost, 0.8 for durability, and 0.9 for ease of construction, its comprehensive score is 0.3×0.9+0.25×0.85+0.2×0.7+0.15×0.8+0.1×0.9=0.8225. This calculation process ensures that the overall performance of each scheme is comprehensively evaluated.
[0060] Step S800 involves storing the improved database, quantified structure-function relationship, candidate solution set, optimal solution, and iteration process data in a pre-defined knowledge graph structure to construct a knowledge base containing a data layer, a model layer, and a solution layer.
[0061] Among them, knowledge graph: a structured semantic knowledge base that uses a graph structure to describe the relationships between entities (such as solid waste materials and performance indicators), facilitating information retrieval and knowledge discovery. Association storage: organizing and storing different types of data and models according to a pre-defined knowledge graph structure to ensure data consistency and traceability.
[0062] The specific process is as follows: First, the validated and optimized database, structure-property relationships, candidate solution sets, and optimal solutions are systematically integrated. For example, the expanded database in step S300, the structure-property relationships established in step S400, the candidate solution set generated in step S600, and the optimal solution selected in step S700 are summarized to form a comprehensive dataset. This process ensures the integrity and consistency of all relevant data, providing a solid foundation for subsequent knowledge base construction. Next, these data are linked and stored according to the preset knowledge graph structure. For example, the physical properties, chemical properties, and process parameters of solid waste materials are linked with corresponding concrete performance indicators to form a structured knowledge graph. Through the knowledge graph, the relationship between the composition, proportion, process parameters, and concrete performance of solid waste materials can be clearly displayed, facilitating users to quickly retrieve and understand the intrinsic connections between the data. Finally, the constructed knowledge base and knowledge graph are stored and managed. For example, professional database management systems (such as MySQL or MongoDB) and knowledge graph tools (such as Neo4j) are used for storage and management to ensure data security and accessibility. At the same time, a data update and maintenance mechanism should be established to regularly update and optimize the knowledge base to adapt to new research progress and practical application needs.
[0063] The integration and classification of dual-source heterogeneous data includes:
[0064] Step S110: Standardize the format of the collected external data, use natural language processing algorithms to extract key entities from unstructured text, and unify the data representation format through regular expressions. Key entities include solid waste type, performance indicators, and process description.
[0065] The specific process is as follows: First, the collected external data is standardized. The PyMuPDF library in Python is used to extract text content from PDF documents, and the pandas library is used to convert it into a structured Excel spreadsheet. For example, performance descriptions of steel slag and fly ash are extracted from literature and converted into tabular form for subsequent processing and analysis. Next, natural language processing algorithms are used to extract key entities from unstructured text. Named Entity Recognition (NER) algorithms are used to identify key information such as solid waste type, performance indicators, and process descriptions from literature and patent texts. For example, "steel slag" is extracted as the solid waste type, "compressive strength 45MPa" as the performance indicator, and "stirring rate 100rpm" as the process description. Then, regular expressions are used to unify the data representation format. For example, regular expressions are used to unify different representations of performance indicators into a standard format, such as unifying "compressive strength is 45MPa" or "compressive strength of 45MPa" into "compressive strength: 45MPa". This process ensures the consistency of data representation, facilitating subsequent data processing and analysis. Finally, the extracted key entities are matched and stored with preset fields. For example, key entities such as "steel slag", "compressive strength: 45MPa", and "stirring rate: 100rpm" are stored in preset fields to form structured data records.
[0066] Step S120: The collected internal data is structured and transformed. The discrete data in the experimental records and production logs are reorganized according to the preset fields of solid waste name, detection date, physical property parameters, process parameters and characterization data. The Z-score method is used to preliminarily screen the reorganized data and remove invalid data that deviates from the preset normal distribution range.
[0067] The specific process is as follows: First, the collected internal data is structured. Discrete data from experimental records and production logs are reorganized according to preset fields such as solid waste name, testing date, physical property parameters, process parameters, and characterization data. For example, data such as steel slag moisture content and fly ash loss on ignition in experimental records are reorganized into a structured data table according to the format of solid waste name-test date-physical property parameters-process parameters-characterization data. This process ensures the uniformity of the internal data format, facilitating subsequent processing and analysis. Next, the reorganized data is initially screened using the Z-score method. By calculating the Z-score value of each data point, it is determined whether it deviates from the preset normal distribution range. For example, for steel slag moisture content data, its Z-score value is calculated; if the Z-score value of a data point exceeds 3 (usually considered the threshold for outliers), it is marked as invalid data and removed. This process ensures the accuracy and consistency of the data, providing a reliable data foundation for subsequent data processing and analysis. Finally, the data after initial screening is stored in a local database for easy querying and use. For example, storing the filtered data in a local MongoDB database ensures long-term availability and reliability. Simultaneously, establishing a data update and maintenance mechanism allows for regular data checks and updates, ensuring timeliness and accuracy.
[0068] Step S130: Use the cosine similarity algorithm to calculate the feature matching degree between the key entities extracted from the external data and the preset fields of the internal data, and establish a bidirectional correlation mapping between performance indicators mentioned in the literature and experimental test performance data, and between patent process description and production process parameters.
[0069] Among them, bidirectional association mapping: establishes a mapping relationship between key entities extracted from external data and preset fields of internal data to ensure that the information between the two can correspond and be associated with each other.
[0070] The specific process is as follows: First, the key entities of the external data extracted in step S110 and the internal data fields after structured transformation in step S120 are preprocessed. For example, the text cleaning and standardization processes are performed on "steel slag compressive strength 45MPa" in the external data and "steel slag-compressive strength-45MPa" in the internal data to ensure consistency in format. This process provides the foundation for subsequent feature matching. Next, the cosine similarity algorithm is used to calculate the similarity between the key entities of the external data and the internal data fields. For example, the cosine similarity between "steel slag compressive strength 45MPa" in the external data and "steel slag-compressive strength-45MPa" in the internal data is calculated. By converting the text into vector form and calculating the cosine value of the angle between these vectors, a similarity score is obtained. If the similarity score exceeds a preset threshold (e.g., 0.8), the two are considered to be successfully matched. Then, based on the matching results, a bidirectional correlation mapping is established between the performance indicators mentioned in the literature and the experimental test performance data, and between the patent process description and the production process parameters. For example, the "compressive strength of steel slag 45MPa" mentioned in the literature is correlated with "compressive strength of steel slag - 45MPa" in the experimental record, and the "stirring rate of 100rpm" in the patent is correlated with "stirring rate of 100rpm" in the production log.
[0071] Step S140: Based on the association mapping results, classify the dual-source data according to solid waste type and nuclear engineering concrete application scenario, remove redundant data with feature matching degree lower than the preset feature matching degree, and finally form a dual-source integrated dataset with clear classification and effective data.
[0072] Feature matching degree: Calculated using the cosine similarity algorithm, it represents the similarity between key entities extracted from external data and preset fields in internal data.
[0073] The specific process is as follows: First, based on the bidirectional association mapping results established in step S130, external and internal data are classified according to solid waste type. For example, all data related to "steel slag" is grouped into one category, and data related to "fly ash" is grouped into another. This process ensures clear data classification, facilitating subsequent analysis and application. Next, the classified data is further subdivided according to nuclear engineering concrete application scenarios. For example, steel slag data and fly ash data related to "nuclear containment building construction" are classified separately to ensure the relevance of the data to specific application scenarios. This process further refines the data classification, improving the data's relevance and practicality. Then, the feature matching degree of the classified data is evaluated. Based on a preset feature matching degree threshold (e.g., 0.8), redundant data with a feature matching degree lower than this threshold is removed. For example, if the feature matching degree of a certain data is 0.7, which is lower than the preset threshold, it is marked as redundant data and removed. This process ensures the validity and consistency of the data and reduces the interference of redundant data on subsequent analysis. Finally, the classified and redundant data-removed data are integrated to form a dual-source integrated dataset. For example, integrating all valid steel slag and fly ash data into a single dataset ensures data integrity and consistency.
[0074] The process of structuring and transforming the collected internal data also includes blockchain traceability processing, including:
[0075] Step S121: Reorganize the discrete data in the experimental records and production logs according to the preset fields of solid waste name - blockchain traceability code - detection date - physical property parameters - process parameters - characterization data, and add a blockchain traceability code on the basis of the original fields.
[0076] Among them, the blockchain traceability code is a unique identifier used to record and track solid waste data throughout the entire process from its origin to its use, ensuring the traceability and immutability of the data.
[0077] The specific process is as follows: First, the discrete data in the experimental records and production logs are reorganized according to preset fields such as solid waste name, blockchain traceability code, testing date, physical property parameters, process parameters, and characterization data. A blockchain traceability code field is added to the original fields to ensure that each data record is associated with a unique blockchain traceability code. For example, data such as the moisture content of steel slag and the loss on ignition of fly ash in the experimental records are reorganized according to the format of solid waste name-blockchain traceability code-testing date-physical property parameters-process parameters-characterization data to form a structured data table. Next, the blockchain traceability code is linked to the entire solid waste data chain, including origin identification, transportation process information, original testing data, and warehousing information. Smart contracts are used to encrypt and store the entire data chain, ensuring data security and immutability. For example, for a batch of steel slag, its blockchain traceability code is linked to the origin identifier (e.g., "steel slag mine A"), transportation process information (e.g., "transport vehicle B, transportation time C"), raw testing data (e.g., "moisture content 6.2%)", and warehousing information (e.g., "warehousing time D, warehousing location E"). This information is encrypted and stored using smart contracts to ensure the authenticity and integrity of the data. Finally, the reconstructed data is stored in a local database for easy subsequent querying and use. For example, the reconstructed data is stored in a local MongoDB database to ensure long-term availability and reliability.
[0078] Step S122: The blockchain traceability code is linked to the solid waste full-chain data, including the origin identification, transportation process information, original test data and warehousing information. The full-chain data is encrypted and stored through smart contracts to ensure that it cannot be tampered with.
[0079] The specific process is as follows: First, the blockchain traceability code is linked to the entire chain of solid waste data. This data includes the solid waste's origin identification, transportation process information, original testing data, and warehousing information. For example, for a batch of steel slag, its blockchain traceability code is linked to the origin identification (e.g., "steel slag mine A"), transportation process information (e.g., "transport vehicle B, transportation time C"), original testing data (e.g., "moisture content 6.2%)", and warehousing information (e.g., "warehousing time D, warehousing location E"). This information is recorded using blockchain technology to ensure that data at each stage can be traced and verified. Next, the entire chain of data is encrypted and stored using smart contracts. A smart contract is an automatically executed contract that uses blockchain technology to encrypt and store data, ensuring data security and immutability. For example, using the Ethereum smart contract platform, the entire chain of solid waste data is written into a smart contract; once written, the data cannot be tampered with. Finally, the encrypted and stored data is stored on the blockchain to ensure long-term availability and traceability. For example, the entire chain of solid waste data can be stored on the Ethereum blockchain, with each data record having a unique blockchain hash value, which can be used to quickly trace and verify the authenticity of the data.
[0080] In step S123, when using the Z-score method to perform preliminary screening of the recombined data, internal data without blockchain traceability codes or with incomplete traceability data are simultaneously removed, and only structured data with complete traceability information is retained.
[0081] The specific process is as follows: First, the reorganized internal data is initially screened using the Z-score method. The Z-score method is a statistical method used to detect outliers in the data. By calculating the Z-score value of each data point, it is determined whether it deviates from the preset normal distribution range. For example, for the moisture content data of steel slag, its Z-score value is calculated. If the Z-score value of a data point exceeds 3 (usually considered the threshold for outliers), it is marked as invalid data and removed. This process ensures the accuracy and consistency of the data, providing a reliable data foundation for subsequent data processing and analysis. Next, the integrity of the blockchain traceability information is verified simultaneously. During the initial screening process, it is checked whether each data record contains a valid blockchain traceability code and to ensure the integrity of the traceability data. For example, it is checked whether the blockchain traceability code is associated with complete solid waste end-to-end data, including origin identification, transportation process information, original testing data, and warehousing information. If a data record lacks a blockchain traceability code or the traceability data is incomplete, it is marked as invalid data and removed. Finally, only the structured data with complete traceability information is retained. Data that has undergone initial screening and traceability integrity verification is stored in a local database for easy subsequent querying and use. For example, the filtered data can be stored in a local MongoDB database to ensure long-term availability and reliability.
[0082] The standardized dataset is imported into a pre-defined database framework to construct an initial database. A pre-defined machine learning model is used to learn the correlation between composition, process, and performance. Supplementary samples are generated to expand the data, resulting in a complete database including:
[0083] Step S310: The standardized dataset is classified into four levels of parameters: solid waste attributes, production process, material properties, and service environment, and imported into a preset relational database framework. This framework defines the field types, units of measurement, and cross-parameter association rules of each parameter through a data dictionary, and enables fast data retrieval through an index structure, thereby constructing a well-structured initial database that can be directly used for model training.
[0084] The specific process is as follows: First, the standardized dataset is classified according to four levels of parameters: solid waste attributes, production process, material properties, and service environment. For example, the physical properties (e.g., moisture content), chemical properties (e.g., MnO content), process parameters (e.g., stirring rate), material properties (e.g., compressive strength), and service environment (e.g., temperature, humidity) of solid waste materials such as steel slag and fly ash in the dataset are classified. Next, the classified data is imported into a pre-defined relational database framework. This framework defines the field types, units of measurement, and cross-parameter association rules for each parameter through a data dictionary. For example, the data dictionary defines the "moisture content" field as a floating-point type with the unit of percentage (%), and specifies its association with the name of the solid waste material. A fast data retrieval is achieved through an index structure, ensuring efficient data storage and retrieval. For example, an index is created for the "solid waste name" field to quickly query all data records related to a specific solid waste. Finally, a well-structured initial database that can be directly used for model training is constructed. Through the above classification and import process, the initial database has a clear data structure, well-defined field types, consistent units of measurement, and clear relationships between parameters. For example, each record in the initial database contains fields such as solid waste name, moisture content, MnO content, stirring rate, and compressive strength, and the relationships between these fields are defined in a standardized manner through a data dictionary and index structure.
[0085] Step S320: Using the data in the initial database as training samples, a BP neural network is used as the preset machine learning model. Solid waste composition parameters and process parameters are used as input layers, and material performance parameters are used as output layers. The network weights are iteratively adjusted through the error backpropagation algorithm to learn the nonlinear correlation between composition, process and performance until the model's prediction error on the validation set is lower than a preset threshold, thus forming a trained model that can accurately map parameter correlations.
[0086] The BP neural network model consists of one input layer, at least one hidden layer, and one output layer. The number of nodes in the input layer corresponds to the number of selected solid waste component parameters and process parameters (e.g., steel slag MnO content, fly ash SiO2 content, stirring rate, etc., totaling n). The number of nodes in the output layer corresponds to the number of material performance parameters to be predicted (e.g., compressive strength, gamma-ray attenuation coefficient, etc., totaling m). The number of hidden layers and the number of nodes per layer can be determined through conventional cross-validation based on the initial database size; for example, two hidden layers with 64 and 32 nodes respectively. During model training, mean squared error (MSE) is used as the loss function, the Adam optimizer is employed, the learning rate is set to 0.001, the batch size is set to 32, and the number of epochs is increased until the validation set error no longer decreases significantly.
[0087] The specific process is as follows: First, the data in the initial database is divided into a training set and a validation set. For example, 80% of the data is used as the training set, and 20% as the validation set. The training set is used to train the BP neural network model, and the validation set is used to evaluate the model's predictive performance. Next, the BP neural network model is constructed. Solid waste composition parameters (such as the MnO content of steel slag and the SiO2 content of fly ash) and process parameters (such as stirring rate and curing temperature) are used as the input layer, and material performance parameters (such as compressive strength and gamma-ray attenuation coefficient) are used as the output layer. For example, the input layer includes features such as the MnO content of steel slag, the SiO2 content of fly ash, and the stirring rate, while the output layer includes compressive strength and gamma-ray attenuation coefficient. Then, the network weights are iteratively adjusted using the error backpropagation algorithm. During training, the BP neural network calculates the predicted value through forward propagation, then calculates the error between the predicted value and the actual value through the error backpropagation algorithm, and adjusts the network weights according to the error. For example, the mean squared error (MSE) is used as the loss function, and the weights are adjusted using gradient descent to minimize the prediction error. Finally, the model is continuously trained until its prediction error on the validation set falls below a preset threshold. For example, a prediction error threshold of 0.05 is set; when the model's prediction error on the validation set is below 0.05, the model training is considered complete. At this point, the model can accurately map the nonlinear correlation between solid waste composition parameters, process parameters, and material performance parameters, forming the trained model.
[0088] Step S330: Based on the correlation patterns captured by the trained model and the sample distribution characteristics of the initial database, a generative adversarial network algorithm is used to generate supplementary samples: the generator is trained with the real samples in the initial database as a benchmark to generate virtual samples that conform to the correlation patterns of composition-process-performance; at the same time, the consistency between the virtual samples and the real samples is verified by the discriminator, and only the virtual samples whose probability of being judged as real samples reaches a preset threshold are temporarily stored as candidate supplementary samples.
[0089] In this Generative Adversarial Network (GAN), the generator employs a fully connected neural network structure. Its input is a random noise vector (100 dimensions), and its output is a simulated solid waste component-process parameter vector. The discriminator also uses a fully connected neural network structure to distinguish whether the input parameter vector comes from a real database or the generator. During adversarial training, the generator and discriminator are optimized alternately until the discriminator can no longer effectively distinguish between real and fake samples (e.g., the discrimination accuracy is close to 50%). After training stabilizes, only virtual samples generated by the generator and judged as "real" by the discriminator with a probability exceeding a preset threshold (e.g., 0.9) are adopted as candidate supplementary samples.
[0090] The specific process is as follows: First, the nonlinear correlation between solid waste composition parameters, process parameters, and material performance parameters is captured using the trained BP neural network model. These correlations serve as prior knowledge for the Generative Adversarial Network (GAN), ensuring that the generated supplementary samples conform to the learned parameter correlations. For example, the model has learned the relationship between the MnO content and compressive strength of steel slag, as well as the effect of stirring rate on the gamma-ray attenuation coefficient. Next, the GAN is constructed. The GAN consists of a generator and a discriminator. The generator aims to generate virtual samples that conform to the correlation between solid waste composition parameters, process parameters, and material performance parameters, while the discriminator aims to distinguish between the generated virtual samples and real samples. Through adversarial training, the generator gradually generates increasingly realistic samples. Then, the generator is trained using real samples from the initial database as a benchmark. By learning the sample distribution characteristics in the initial database, the generator generates virtual samples that conform to the composition-process-performance correlation. For example, the generator produces new virtual samples based on steel slag and fly ash samples in the initial database. These virtual samples share similar distribution characteristics with real samples in terms of composition, process, and performance. Simultaneously, a discriminator verifies the consistency between the virtual and real samples. The discriminator evaluates the generated virtual samples to determine their high degree of consistency with the real samples. For example, the discriminator judges the authenticity of the virtual samples by comparing the feature distributions of the virtual and real samples. Only when the authenticity of a virtual sample reaches a preset accuracy rate (e.g., 90%) is it temporarily stored as a candidate supplementary sample.
[0091] Step S340: Integrate the initial database with the candidate supplementary samples verified by the discriminator, use the K-means clustering algorithm for consistency verification, remove abnormal samples whose clustering deviation values exceed the preset range, and finally form a complete database with balanced data distribution and significant correlation patterns.
[0092] The specific process is as follows: First, the real samples in the initial database are integrated with the candidate supplementary samples verified by the discriminator. For example, the steel slag and fly ash samples in the initial database are merged with the generated virtual samples to form a larger dataset. Next, the K-means clustering algorithm is used to verify the consistency of the integrated dataset. The K-means clustering algorithm divides the data points into K clusters, calculates the centroid of each cluster, and evaluates the deviation of each data point from the centroid. For example, the integrated dataset is divided into several clusters, each cluster representing a group of data points with similar characteristics. Then, the deviation value of each data point from the centroid of its cluster is calculated. For example, Euclidean distance is used to calculate the deviation of each data point from the centroid. A preset range is set (e.g., the deviation value does not exceed 10% of the centroid), and data points with deviation values exceeding the preset range are marked as outliers. This process ensures the consistency between the data points and the cluster centroids, eliminating possible outliers or isolated data points. Finally, the samples marked as outliers are removed, forming the final complete database. For example, samples with deviation values exceeding a preset range are removed from the dataset, while data points with high consistency are retained. The resulting comprehensive database not only has a balanced data distribution but also exhibits significant correlation patterns.
[0093] By using a pre-defined gene coding model to mine and refine the structure-activity relationship of solid waste components, proportions, processes, and performance in the database, a quantitative structure-activity relationship between performance and solid waste components, proportions, and processes is established, including:
[0094] Step S410: Extract data with dynamic association tags from the complete database and construct a four-dimensional feature matrix according to the dimensions of solid waste composition, proportioning parameters, process parameters, and performance indicators. The dynamic association tags record the source information of the data and the causal relationship information between the parameters.
[0095] The specific process is as follows: First, data is extracted from a comprehensive database. This data not only includes the composition, proportion, and process parameters of solid waste, but also carries dynamic correlation tags. These tags record the source and causal relationship information of the data. For example, the extracted data records may include the MnO content of steel slag, the SiO2 content of fly ash, the ratio of steel slag to fly ash, the stirring rate, the curing temperature, and corresponding performance indicators such as compressive strength and gamma-ray attenuation coefficient, while also carrying a dynamic correlation tag stating that "increased MnO content in steel slag leads to increased compressive strength." Next, the extracted data is organized and structured according to four dimensions: solid waste composition, proportion parameters, process parameters, and performance indicators. For example, the MnO content of steel slag and the SiO2 content of fly ash are categorized as solid waste composition, the ratio of steel slag to fly ash as proportion parameters, the stirring rate and curing temperature as process parameters, and the compressive strength and gamma-ray attenuation coefficient as performance indicators. This process ensures the structure and orderliness of the data, providing a foundation for the subsequent construction of the feature matrix. Then, a four-dimensional feature matrix is constructed. The processed data is arranged according to the four dimensions mentioned above, forming a four-dimensional feature matrix. For example, each row of the feature matrix represents a data sample, and each column represents a feature dimension. In this way, the information of each data sample in the four dimensions of solid waste composition, proportioning parameters, process parameters, and performance indicators is completely recorded.
[0096] Step S420: Based on the four-dimensional feature matrix, proceed to the parameter encoding stage: convert solid waste components, ratio parameters, and process parameters into standardized gene sequences, and embed the causal association information in the dynamic association tags into the corresponding coding positions to form a gene coding dataset with causal tags.
[0097] Among these, the gene sequence refers to the standardized numerical sequence converted from solid waste components, proportioning parameters, and process parameters, similar to gene coding in biological genetics, used to represent the characteristics of different parameters. Dynamic association tags are labels that record the causal relationships and correlations between parameters in the data, used to trace and explain the impact of parameter changes on performance in subsequent steps. Causal markers are causal information embedded in the gene sequence, used to identify the specific impact of parameter changes on performance indicators, such as "increasing the MnO content of steel slag leads to increased compressive strength."
[0098] The specific process is as follows: First, the solid waste components, proportioning parameters, and process parameters in the four-dimensional feature matrix are standardized. For example, component parameters such as the MnO content of steel slag and the SiO2 content of fly ash, as well as process parameters such as the proportion of steel slag and fly ash, stirring rate, and curing temperature, are converted into standardized values between 0 and 1 using a normalization method. This process ensures the comparability and consistency between different parameters, providing a foundation for subsequent coding work. Next, the standardized parameters are converted into gene sequences. For example, the standardized parameters such as the MnO content of steel slag, the SiO2 content of fly ash, the proportion of steel slag and fly ash, stirring rate, and curing temperature are converted into gene sequences according to certain rules. Each parameter corresponds to a gene locus, and each value in the gene sequence represents the characteristic of the corresponding parameter. This process converts complex parameter information into a concise gene sequence form, facilitating subsequent processing and analysis. Then, causal information is extracted from the dynamic association tags and embedded into the corresponding gene coding loci. For example, if the dynamic association tag records the causal information that "increasing the MnO content in steel slag leads to increased compressive strength," then the gene locus corresponding to the MnO content in the steel slag will be labeled as "positively affecting compressive strength" in the gene sequence. This process ensures that the gene sequence contains not only numerical information about the parameters but also information about the causal impact of parameter changes on performance indicators. Finally, the gene sequences with causal labels are integrated into a gene coding dataset. For example, the gene sequences of all samples and their corresponding causal labels are integrated into one dataset to form a gene coding dataset with causal labels.
[0099] Step S430: Using the encoded gene-encoded dataset as input, proceed to the relationship mining stage: use the ridge regression algorithm to control parameter collinearity and quantify the linear correlation coefficient between a single parameter and the performance index; simultaneously use the random forest algorithm to mine the nonlinear relationship of multi-parameter interaction, verify and correct the linear and nonlinear analysis results based on the embedded causal association information, and output the complete structure-activity relationship coefficient.
[0100] "Gene encoding" is the process of mapping continuous solid waste component parameters (such as content percentage), proportion parameters (such as admixture ratio), and process parameters (such as temperature and rate) to a fixed-length numerical sequence through interval discretization and normalization. Each bit in the sequence represents the encoding of a parameter within a specific interval. Dynamic association labels (such as "parameter A is positively correlated with performance B") are converted into weight symbols or constraints for the encoded bits of that parameter and embedded into subsequent regression and random forest analyses. When using ridge regression, the regularization coefficient α is optimally determined within the interval of 0.1, 10, 0.1, and 10 through cross-validation. The number of decision trees in the random forest model is set to 100, and the maximum depth of the trees is dynamically adjusted according to the number of features. Ridge regression algorithm: a linear regression algorithm that improves the stability and generalization ability of the model by adding an L2 regularization term to the loss function to control the collinearity of parameters. Structure-property relationship coefficient: a coefficient that quantifies the relationship between solid waste components, proportion parameters, process parameters, and material performance indicators, including linear correlation coefficients and nonlinear correlation coefficients.
[0101] The specific process is as follows: First, using the encoded gene dataset as input, the Ridge Regression algorithm is employed to control parameter collinearity. Ridge Regression effectively controls collinearity among parameters by adding an L2 regularization term to the loss function, improving the model's stability and generalization ability. For example, when analyzing the linear relationship between MnO content and compressive strength in steel slag, Ridge Regression can provide a stable linear correlation coefficient, quantifying the impact of MnO content changes on compressive strength. Next, the Random Forest algorithm is simultaneously used to mine nonlinear relationships involving multiple parameters. By constructing multiple decision trees and integrating their prediction results, Random Forest can effectively capture complex nonlinear relationships. For example, when analyzing the combined influence of multiple parameters such as MnO content in steel slag, SiO2 content in fly ash, and stirring rate on the γ-ray attenuation coefficient, Random Forest can provide a nonlinear relationship coefficient, quantifying the impact of these parameter interactions on performance indicators.
[0102] Then, the embedded causal association information is used to verify and correct the initially obtained linear and nonlinear relationships. Specifically: 1. Causal consistency verification: The parameter-performance relationships (including the direction and magnitude of influence) obtained from ridge regression and random forest analysis are compared with the causal association tags embedded in the gene encoding. For example, if the causal tag indicates that "the increase in steel slag MnO content has a positive effect on compressive strength", but the analysis results show a negative correlation or insignificant effect, it is marked as an item to be verified. 2. Conflict analysis and correction: For items with conflicts, the original experimental data or production records in the database are traced back to verify the data quality and boundary conditions. If the original data is confirmed to be reliable, the model assumptions are re-examined or the algorithm parameters are adjusted; if the data is found to have limitations or biases, the analysis results are weighted and corrected or constraints are introduced and recalculated according to the priority of causal knowledge. 3. Knowledge fusion and confirmation: The verified and corrected linear association coefficients and nonlinear relationship coefficients are integrated to form a set of self-consistent structure-activity relationship coefficients that are based on data-driven principles and conform to domain causal knowledge.
[0103] Finally, the complete structure-property relationship coefficients are output, providing a foundation for the subsequent mapping generation stage. These structure-property relationship coefficients not only include linear relationships between single parameters and performance indicators, but also nonlinear relationships involving multiple parameter interactions. Furthermore, they have been verified through causal knowledge, providing a more reliable and interpretable quantitative basis for constructing accurate function mapping models.
[0104] Step S440: Based on the structure-property relationship coefficients obtained in the relationship mining stage, proceed to the mapping generation stage: Using solid waste composition, ratio, and process parameters as independent variables and performance indicators as dependent variables, construct a function mapping model. Through this model, clarify the quantitative correspondence between each parameter combination and the corresponding performance indicator, and form a quantitative structure-property relationship between performance and solid waste composition, ratio, and process that integrates data sources and causal relationships.
[0105] Among them, the function mapping model is a mathematical model used to describe the quantitative relationship between input parameters (such as solid waste composition, proportion, and process parameters) and output performance indicators (such as compressive strength and gamma-ray attenuation coefficient). This model predicts performance indicators through input parameters, providing a scientific basis for the design and optimization of concrete materials in nuclear engineering.
[0106] The specific process is as follows: First, a function mapping model is constructed using solid waste composition, proportioning parameters, and process parameters as independent variables, and performance indicators as dependent variables. For example, parameters such as the MnO content of steel slag, the SiO2 content of fly ash, the ratio of steel slag to fly ash, stirring rate, and curing temperature are used as independent variables, while performance indicators such as compressive strength and gamma-ray attenuation coefficient are used as dependent variables to construct a mathematical model. This model can predict performance indicators based on the input parameters. Next, using the structure-property relationship coefficients obtained in step S430, the quantitative correspondence between each parameter combination and the corresponding performance indicator is clarified. For example, based on the linear correlation coefficient obtained by the ridge regression algorithm and the nonlinear relationship coefficient obtained by the random forest algorithm, the contribution of each parameter is quantified into the model. Then, the accuracy and reliability of the model are ensured through model training and validation. For example, the model is trained using a training dataset and then validated using a validation dataset to ensure that the model's prediction error is within an acceptable range. This process ensures that the model can accurately predict performance indicators, providing a reliable tool for subsequent applications. Finally, a mapping relationship between performance and solid waste composition, proportioning, and process is established. For example, through model prediction, the quantitative correspondence between combinations of parameters such as the MnO content of steel slag, the SiO2 content of fly ash, the ratio of steel slag to fly ash, the stirring rate, and the curing temperature, and performance indicators such as compressive strength and gamma-ray attenuation coefficient, has been clarified. This mapping relationship provides a scientific basis for the design and optimization of concrete materials for nuclear engineering.
[0107] Step S450: Based on the established quantitative structure-property relationship, first calibrate its application boundary: Combine with the industry standard for nuclear engineering concrete, define the effective application range of each parameter in the quantitative structure-property relationship, eliminate invalid parameter associations that exceed the boundary, and obtain the effective quantitative structure-property relationship within the boundary.
[0108] Application boundaries refer to the effective range of parameters under specific application scenarios. For concrete materials in nuclear engineering, application boundaries are typically defined by industry standards, design specifications, and actual engineering needs to ensure that material performance meets safety and functional requirements. Industry standards refer to standardized regulations regarding material performance, design, and construction within the nuclear engineering field. These standards ensure the safety and reliability of materials and structures.
[0109] The specific process is as follows: First, based on industry standards for nuclear engineering concrete, the effective application range of each parameter is determined. For example, according to the International Atomic Energy Agency (IAEA) and relevant national standards, the effective ranges of parameters such as the MnO content of steel slag, the SiO2 content of fly ash, the mixing rate, and the curing temperature are determined. These ranges are usually based on the physicochemical properties of the materials, engineering experience, and safety requirements. Next, the established quantitative structure-property relationships are calibrated against the application boundaries. The value of each parameter is compared with the effective range defined in the industry standard. For example, if the industry standard stipulates that the MnO content of steel slag should be between 5% and 15%, then parameter values in the quantitative structure-property relationship that exceed this range will be considered invalid. Then, invalid parameter associations that exceed the application boundaries are eliminated. For example, if the MnO content of steel slag in the quantitative structure-property relationship is 20%, exceeding the 5% to 15% range stipulated by the industry standard, then this parameter association will be eliminated. This process ensures that all retained quantitative structure-property relationships are effective and reliable in actual engineering. Finally, the effective quantitative structure-property relationships within the boundaries are obtained. These relationships not only conform to industry standards but also have practical application value in engineering. For example, calibrated quantitative structure-property relationships can be used to design and optimize concrete materials for nuclear engineering to ensure that their performance meets safety and functional requirements.
[0110] Step S460: Initiate causal interpretability verification for valid quantified structure-property relationships within the boundary: quantify the impact weight of each parameter on performance using SHAP values, generate a rule-based interpretation through LIME, and compare it with the preset rules in the nuclear engineering expert database.
[0111] SHAP Score: SHAP (SHapley Additive Explanations) is a game theory-based method used to quantify the contribution of each feature (parameter) to the model's predictions. It provides a fair and interpretable way to understand the model's decision-making process. LIME: LIME is a method for interpreting individual predictions and the importance of their features. It explains the predictions of complex models by locally fitting a simple model to the prediction points. Causal Interpretability: This refers to the ability to explain model predictions through causal relationships, i.e., to clearly define the specific impact of parameter changes on performance metrics, rather than just a statistical correlation.
[0112] The specific process is as follows: First, SHAP values are used to quantify the weight of each parameter's influence on performance. For example, for parameters such as the MnO content of steel slag, the SiO2 content of fly ash, and the stirring rate, their SHAP values for compressive strength and gamma-ray attenuation coefficient are calculated. These SHAP values represent the contribution of each parameter to the performance index, thus quantifying the importance of the parameter. Next, a rule-based interpretation is generated using LIME. LIME locally fits a simple model to explain the prediction results of complex models. For example, for a specific concrete sample, LIME can explain why its compressive strength reaches 45 MPa, specifically which parameters (such as the MnO content of steel slag and the stirring rate) contribute the most to this result. Then, the SHAP values and the interpretation generated by LIME are compared with the preset rules of the nuclear engineering expert database. The nuclear engineering expert database contains the experience and knowledge of industry experts and presets causal relationship rules between parameters and performance indicators. For example, the expert database may preset a rule that "an increase in the MnO content of steel slag will lead to an increase in compressive strength." By comparison, it is verified whether the interpretation of the quantified structure-property relationship is consistent with the expert rules. Finally, if the interpretation results conflict with the expert rules, the subsequent correction steps will be initiated (step S470). If the interpretation results are consistent with the expert rules, the quantified structure-activity relationship is considered to have good causal interpretability and can be used in practical applications.
[0113] Step S470: If the interpretation result conflicts with the expert rules, trace back to the original data in the improved database by dynamically associating labels, correct the conflicting items, and then re-optimize the quantified structure-property relationship until the quantified structure-property relationship within the boundary simultaneously meets the application boundary requirements and expert causal knowledge.
[0114] Conflicting items refer to those items in the causal interpretability verification (step S460) where the interpretation result of the quantified structure-activity relationship is inconsistent with the preset rules of the nuclear engineering expert database. These conflicting items may indicate a deviation in the quantified structure-activity relationship or that the expert rules need to be updated. Dynamic association tag tracing: Using the data source and causal information recorded by the dynamic association tags, the original data of the conflicting items is traced to determine the cause of the conflict and make corrections. Quantified structure-activity relationship optimization: Based on the tracing results, the quantified structure-activity relationship is adjusted and optimized to ensure that its prediction results are consistent with the expert rules, improving its reliability and accuracy.
[0115] The specific process is as follows: First, when the causal interpretability verification (step S460) finds a conflict between the interpretation result of the quantified structure-property relationship and the preset rules of the nuclear engineering expert database, the source tracing of the conflict item is initiated. For example, the quantified structure-property relationship interpretation shows that "increasing the MnO content of steel slag leads to a decrease in compressive strength," while the preset expert rule is "increasing the MnO content of steel slag leads to an increase in compressive strength." In this case, it is necessary to trace the cause of the conflict. Next, the original data of the conflict item is traced using the data source and causal information recorded by the dynamic association tags. The dynamic association tags record the source and causal relationship of the data, and the corresponding original data record in the database can be quickly located through these tags. For example, tracing back to an experimental record shows that under specific conditions, an increase in the MnO content of steel slag does indeed lead to a decrease in compressive strength, which may be because the experimental conditions are different from the conditions assumed by the expert rule. Then, based on the source tracing results, the cause of the conflict is analyzed. Possible causes include differences in experimental conditions, deviations in the quantified structure-property relationship, and limitations of the expert rule. For example, it is found that the curing temperature used in the experiment is different from the temperature assumed by the expert rule, leading to different results. Next, the conflict item is corrected. Based on the source tracing results, adjust the parameters of the quantified structure-activity relationship (SPR) or update the expert rules to resolve conflicts. For example, if the conflict is found to be caused by differences in experimental conditions, the SPR can be adjusted to consider the impact of different curing temperatures on performance, or the expert rules can be updated to reflect the new experimental conditions. Finally, the SPR is re-optimized. After correcting the conflict terms, the mining and construction process of steps S410 to S440 is re-executed to ensure that the newly obtained SPR simultaneously meets the application boundary requirements (step S450) and expert causal knowledge (step S460). For example, the model is retrained based on the corrected data or rules to verify whether the newly obtained SPR is consistent with the updated expert rules and to ensure that its prediction results are valid and reliable within the application boundary.
[0116] Step S480: The verified quantified structure-property relationship is standardized and encapsulated according to the parameter effective range, performance compliance threshold, causal explanation label, and source traceability code to form an interpretable quantified structure-property relationship with boundary constraints, and stored in the knowledge base model layer.
[0117] Among them, the effective range of parameters refers to the reasonable range of values for parameters (such as solid waste composition, proportion, and process parameters) in practical applications, ensuring that material performance meets design requirements. Performance compliance thresholds refer to the minimum standards that material properties (such as compressive strength and gamma-ray attenuation coefficient) must meet to satisfy the safety and functional requirements of nuclear engineering concrete. Causal explanation labels are used to mark the causal relationship between parameter changes and performance indicators, facilitating the understanding and interpretation of model prediction results. Source traceability codes are unique identifiers used to trace the source of data, ensuring data traceability and transparency.
[0118] The specific process is as follows: First, the validated quantitative structure-property relationships are standardized and encapsulated. This includes integrating the effective range of each parameter, the corresponding performance threshold, the causal explanation label, and the source traceability code. For example, for the MnO content of steel slag, its effective range (5% to 15%), performance threshold (e.g., compressive strength ≥ 45 MPa), causal explanation label (e.g., "positively affects compressive strength"), and corresponding data traceability code are encapsulated. Next, based on the above encapsulation, an interpretable quantitative structure-property relationship with boundary constraints is formed. This relationship not only defines the mathematical association between parameters and performance but also integrates its application boundaries, performance targets, and causal knowledge. Then, the encapsulated, boundary-constrained, interpretable quantitative structure-property relationship is stored in the knowledge base model layer. The knowledge base model layer is a structured data storage system used to store validated and standardized model knowledge. This model knowledge can be used for subsequent model training, performance prediction, and material design. For example, it can be stored in a relational database or knowledge graph for easy querying and application.
[0119] Inputting the target performance requirements of nuclear engineering concrete, based on the above-mentioned quantified structure-property relationship and preset derivation constraints, multiple sets of initial parameter combinations are automatically generated and reverse-derived to generate multiple initial optimization schemes, including:
[0120] Step S510: Invoke the interpretable quantitative structure-property relationship with boundary constraints in the knowledge base model layer, and extract the effective range of parameters, performance attainment threshold and causal explanation label associated with the target performance requirements.
[0121] The specific process is as follows: First, the target performance requirements for nuclear engineering concrete are defined. For example, target performance requirements may include compressive strength ≥45MPa and gamma-ray attenuation coefficient ≥0.55 / cm. These requirements are set according to specific engineering application scenarios and safety standards. Next, the interpretable quantitative structure-property relationship with boundary constraints in the knowledge base model layer is invoked. This relationship contains verified and standardized quantitative structure-property information, providing information such as the effective range of parameters, performance achievement thresholds, and causal explanation labels. For example, this relationship may record that the effective range of MnO content in steel slag is 5% to 15%, corresponding to a compressive strength achievement threshold of 45MPa. Then, the effective range of parameters and performance achievement thresholds associated with the target performance requirements are extracted from this quantitative structure-property relationship. For example, for a target compressive strength ≥45MPa, the effective range of MnO content in steel slag (5% to 15%) and the corresponding compressive strength achievement threshold (45MPa) are extracted. At the same time, causal explanation labels are extracted, such as "increasing the MnO content in steel slag leads to an increase in compressive strength". Finally, the extracted parameter effective range, performance attainment threshold, and causal explanation labels are used for subsequent initial parameter combination generation and back-derivation. This information provides crucial boundary conditions and guidance for generating an initial optimization scheme that meets the target performance requirements.
[0122] Step S520: Based on the effective range of the extracted parameters, the Latin hypercube sampling method is used to generate initial parameter combinations: within the effective range of the parameters of the effective quantitative structure-property relationship within the boundary, combinations of solid waste components, proportioning parameters, and process parameters are uniformly selected. Each combination must meet the safety constraints preset by the nuclear engineering expert database to ensure that all initial combinations fall within the solution space covered by the effective quantitative structure-property relationship within the boundary, and a preset number of initial parameter combinations are generated.
[0123] Latin hypercube sampling: A statistical sampling technique used to generate sample points in a multidimensional space. It improves the representativeness and efficiency of sampling by dividing each dimension into equally probable intervals and randomly selecting a sample point within each interval, ensuring a uniform stratified distribution of sample points across all dimensions. Safety constraints: Based on pre-defined rules from a nuclear engineering expert database, it ensures that the generated parameter combinations meet the safety and functional requirements of nuclear engineering concrete. These constraints typically include reasonable ranges for parameters and the relationships between parameters. Solution space: Refers to the space comprised of all possible parameter combinations. In this step, the solution space refers to the range of parameter combinations covered by the effective quantified structure-property relationships within the boundaries, ensuring that the generated parameter combinations are effective and reliable in practical applications.
[0124] The specific process is as follows: First, based on the effective range of the parameters extracted in step S510, the value range of each parameter is determined. For example, the effective range of MnO content in steel slag is 5% to 15%, the effective range of SiO2 content in fly ash is 20% to 30%, and the effective range of stirring rate is 50 to 150 rpm. Next, the Latin hypercube sampling method is used to uniformly select parameter combinations within these effective ranges. The Latin hypercube sampling method divides the value range of each parameter into equally probable intervals and randomly selects a sample point within each interval, ensuring the stratification uniformity of the sample points across all parameter dimensions. For example, the MnO content range of steel slag is divided into 10 equally probable intervals, and a sample point is randomly selected within each interval; similarly, parameters such as SiO2 content and stirring rate of fly ash are sampled. Then, during the sampling process, each generated parameter combination is verified to ensure that it meets the preset safety constraints of the nuclear engineering expert database. For example, the expert database might pre-determine the ratio between the MnO content of steel slag and the SiO2 content of fly ash, or the synergistic relationship between stirring rate and curing temperature. Verification ensures that all initial combinations fall within the solution space covered by the effective quantified structure-property relationships within the boundaries. Finally, a predetermined number of initial parameter combinations are generated. For example, if 100 initial parameter combinations need to be generated, 100 combinations are uniformly selected within the effective parameter range using Latin hypercube sampling, ensuring that each combination meets safety constraints.
[0125] Step S530: Using the target performance requirement as the output constraint, perform reverse derivation on the initial parameter combination based on the function mapping relationship in the interpretable quantified structure-property relationship.
[0126] Target performance requirements: Performance indicators such as compressive strength and gamma-ray attenuation coefficient are set according to the specific application scenarios of nuclear engineering concrete. Quantified structure-property relationship: The quantified structure-property relationship with boundary constraints describes the calculable relationship between input parameters (solid waste composition, mix proportion, process parameters) and output performance indicators (such as compressive strength and gamma-ray attenuation coefficient).
[0127] The specific process is as follows: First, define the target performance requirements. For example, target performance requirements may include compressive strength ≥ 45 MPa and gamma-ray attenuation coefficient ≥ 0.55 / cm. These requirements are set according to specific engineering application scenarios and safety standards. Next, invoke an interpretable quantitative structure-property relationship with boundary constraints. This relationship contains the calculable relationship between input parameters (solid waste composition, proportion, process parameters) and output performance indicators (such as compressive strength and gamma-ray attenuation coefficient). For example, this relationship may record the specific relationship between parameters such as MnO content of steel slag, SiO2 content of fly ash, and stirring rate and compressive strength and gamma-ray attenuation coefficient. Then, using the target performance requirements as output constraints, perform reverse derivation for each set of initial parameter combinations. For example, for a target compressive strength ≥ 45 MPa, derive the specific values of solid waste composition, proportion, and process parameters that meet this performance requirement from the initial parameter combinations based on this quantitative structure-property relationship. This process may involve numerical optimization methods, such as gradient descent or genetic algorithms, to find parameter combinations that satisfy the constraints. Finally, the reverse derivation results for each initial parameter combination are recorded. These results include the specific values of the parameters and the corresponding performance estimates. For example, the compressive strength corresponding to a certain parameter combination (10% MnO content in steel slag, 25% SiO2 content in fly ash, and a stirring speed of 100 rpm) is recorded as 47 MPa, and the gamma-ray attenuation coefficient is 0.58 / cm.
[0128] Step S540: After completing the reverse derivation for each set of initial parameter combinations, organize them to form the corresponding initial optimization scheme. The scheme includes the specific values of the parameters, the performance target prediction, and the derivation logic explanation based on the causal explanation label.
[0129] The specific process is as follows:
[0130] First, organize the reverse derivation results for each set of initial parameter combinations. This includes recording the specific values of each parameter and the corresponding performance estimates. For example, for a certain parameter combination (10% MnO content in steel slag, 25% SiO2 content in fly ash, and a stirring speed of 100 rpm), record its corresponding compressive strength as 47 MPa and gamma-ray attenuation coefficient as 0.58 / cm.
[0131] Next, based on the causal explanation labels, the rationale for the selection of each initial parameter combination and the performance prediction is explained. The causal explanation labels record the specific impact of parameter changes on performance indicators. This information can be used to explain why specific parameter combinations are chosen and how these combinations affect performance indicators. For example, it explains that "the 10% MnO content in steel slag is chosen because, according to model predictions, this content can increase the compressive strength to 47 MPa, meeting the target performance requirements."
[0132] Then, the detailed information and derivation logic of each initial parameter combination are compiled into a complete initial optimization scheme. Each scheme includes not only the specific values of the parameters and the performance estimates, but also a detailed explanation of these values and estimates. For example, a complete initial optimization scheme might look like this:
[0133] Parameter values:
[0134] MnO content in steel slag: 10%;
[0135] SiO2 content of fly ash: 25%;
[0136] Stirring speed: 100 rpm;
[0137] Performance target estimate:
[0138] Compressive strength: 47 MPa;
[0139] Gamma-ray attenuation coefficient: 0.58 / cm;
[0140] Explanation of the derivation logic: The steel slag MnO content of 10% was chosen because, according to model predictions, this content can increase the compressive strength to 47 MPa, meeting the target performance requirements. The fly ash SiO2 content of 25% was chosen because this content can optimize the gamma-ray attenuation coefficient to 0.58 / cm, meeting the target performance requirements. The mixing rate of 100 rpm was chosen because this rate ensures the uniformity and stability of the concrete, meeting the process requirements for nuclear engineering concrete.
[0141] Step S550: Summarize all initial optimization schemes to form a preset number of initial optimization scheme sets.
[0142] Using the target performance requirement as the output constraint, the reverse derivation of the initial parameter combination based on the interpretable quantified structure-property relationship includes:
[0143] Step S531: Decompose the target performance requirements into key sub-targets of nuclear engineering, allocate the weights of the sub-targets according to the safety priority of nuclear engineering, and clarify the minimum compliance threshold for each sub-target.
[0144] Target Performance Requirements: Performance indicators set based on the specific application scenarios of nuclear engineering concrete, such as compressive strength and gamma-ray attenuation coefficient. Key Sub-Objectives for Nuclear Engineering: The target performance requirements are decomposed into multiple key sub-objectives, each corresponding to a specific performance indicator, such as compressive strength, durability, and radiation shielding performance. Sub-Objective Weights: Weights assigned to each sub-objective based on nuclear engineering safety priorities, reflecting the importance of different sub-objectives within the overall objective. Minimum Compliance Thresholds: The minimum performance standards that each sub-objective must meet to ensure that nuclear engineering concrete materials meet safety and functional requirements.
[0145] The specific process is as follows:
[0146] First, the overall target performance requirements for concrete in nuclear engineering should be clearly defined. For example, the overall target performance requirements may include compressive strength ≥45MPa, gamma-ray attenuation coefficient ≥0.55 / cm, and durability ≥300 freeze-thaw cycles.
[0147] Next, the overall target performance requirements are decomposed into several key sub-targets. For example, the target performance requirements are decomposed into the following sub-targets: compressive strength ≥ 45 MPa; gamma-ray attenuation coefficient ≥ 0.55 / cm;
[0148] Durability ≥ 300 freeze-thaw cycles.
[0149] Then, weights are assigned to each sub-objective according to the safety priority of nuclear engineering. For example, based on the safety requirements of nuclear engineering, the weights are assigned as follows: compressive strength: weight 0.4; gamma-ray attenuation coefficient: weight 0.3; durability: weight 0.3.
[0150] Finally, the minimum compliance thresholds for each sub-target are defined. These thresholds are set according to the safety standards and functional requirements of nuclear engineering to ensure that the material performance meets the needs of practical applications. For example, the minimum compliance threshold for compressive strength is set at 45 MPa; the minimum compliance threshold for gamma-ray attenuation coefficient is 0.55 / cm; and the minimum compliance threshold for durability is 300 freeze-thaw cycles.
[0151] Step S532: Based on the functional mapping relationship in the interpretable quantified structure-property relationship, establish the correlation equation between the initial parameter combination and the performance of each sub-target, and embed the parameter influence direction in the causal explanation label in the equation.
[0152] The specific process is as follows: First, an interpretable, quantified structure-property relationship with boundary constraints is invoked. This relationship describes the calculable relationship between input parameters (such as solid waste composition, proportion, and process parameters) and output performance indicators (such as compressive strength and gamma-ray attenuation coefficient). For example, this relationship may record the specific relationship between parameters such as the MnO content of steel slag, the SiO2 content of fly ash, and the stirring rate, and compressive strength and gamma-ray attenuation coefficient. Next, based on the sub-objectives decomposed in step S531, correlation equations are established between the initial parameter combination and the performance of each sub-objective. For example, for the compressive strength sub-objective, a correlation equation is established to describe the relationship between parameters such as the MnO content of steel slag, the SiO2 content of fly ash, and the stirring rate, and compressive strength. Similarly, corresponding correlation equations are established for sub-objectives such as gamma-ray attenuation coefficient and durability. Then, in each correlation equation, the coefficients or structures of the corresponding parameters in the equation are constrained or given symbolic guidance according to the parameter influence direction indicated by the causal interpretation label. The causal interpretation label records the specific impact of parameter changes on performance indicators, and this information directly guides the setting of the parameter action direction in the equation. For example, in the correlation equation for compressive strength, if the causal explanation label is "increased MnO content in steel slag leads to increased compressive strength," then it is essential to ensure that the contribution coefficient of the term characterizing MnO content in the equation to performance is positive. Finally, verify whether the established correlation equation accurately reflects the relationship between the initial parameter combination and the sub-target performance. By substituting known parameter combinations and performance data, check whether the predicted results of the equation are consistent with the actual data. If there is a significant deviation between the predicted results of the equation and the actual data, it is necessary to trace back to the process of establishing the quantified structure-property relationship or the causal explanation label for verification and correction.
[0153] Step S533: Set dual constraints for parameter adjustment. Use the effective range of parameters in the interpretable quantitative structure-property relationship as a hard constraint, and combine it with the sub-target weights to set soft constraints. The single adjustment range of the associated parameters of high-weight sub-targets is ≤ a preset value, and the single adjustment range of the associated parameters of low-weight sub-targets is ≤ another preset value.
[0154] Hard constraints refer to the strict limitations that must be followed during parameter adjustment, typically set within the effective range of parameters based on physicochemical properties or engineering safety standards. For example, the MnO content of steel slag must be between 5% and 15%. Soft constraints refer to relatively flexible limitations during parameter adjustment, typically set based on the weight of sub-objectives. For example, the adjustment range of a single parameter associated with a high-weight sub-objective may be limited to a smaller range to ensure the stability and reliability of the adjustment. Parameter adjustment range refers to the amount of change in the parameter value each time it is adjusted. For example, the adjustment range of the MnO content in steel slag may be limited to within 1% each time.
[0155] The specific process is as follows: First, set hard constraints. Based on the effective range of the parameters extracted from the quantified structure-property relationship, set hard constraints for each parameter. These hard constraints ensure that the parameter adjustment process does not exceed the range allowed by physicochemical properties and engineering safety standards. For example, the effective range of MnO content in steel slag is 5% to 15%, so the MnO content must always be kept within this range during the adjustment process. Next, set soft constraints. Based on the sub-objective weights, set soft constraints for each parameter. The single adjustment range of parameters associated with high-weight sub-objectives is limited to a small range to ensure the stability and reliability of the adjustment. For example, for a high-weight compressive strength sub-objective, the single adjustment range of associated parameters (such as the MnO content of steel slag) may be limited to within 1%. The single adjustment range of associated parameters associated with low-weight sub-objectives can be relatively large, but still needs to be within a reasonable range. For example, for a low-weight durability sub-objective, the single adjustment range of associated parameters (such as stirring rate) may be limited to within 10 rpm. Then, combine hard and soft constraints to adjust the parameters. During each adjustment process, first check whether the adjusted parameter value meets the hard constraint conditions. If the hard constraints are met, the adjustment range is then determined or verified based on the soft constraints. For example, if the adjusted MnO content in the steel slag is 16%, exceeding the hard constraint range (5% to 15%), it needs to be readjusted to ensure it remains within the effective range. If the adjusted MnO content is 14%, meeting the hard constraints, but the single adjustment exceeds the soft constraints (e.g., 1%), it needs to be adjusted to meet the soft constraints. Finally, the results of each adjustment are recorded to ensure the transparency and traceability of the parameter adjustment process. For example, the parameter values, adjustment range, and post-adjustment performance estimates for each adjustment are recorded for subsequent analysis and optimization.
[0156] Step S534: The weighted gradient descent algorithm is used to perform parameter iterative optimization. The loss function is the sum of the weights of the sub-targets multiplied by the deviation between the performance prediction of each sub-target and the minimum threshold. The parameters that have a significant impact on the high-weight sub-targets are iterated first. After each iteration, the performance prediction is calculated through the correlation equation until the loss function value is lower than the preset threshold.
[0157] Weighted Gradient Descent Algorithm: An optimization algorithm that minimizes the loss function by calculating its gradient and iteratively updating the parameters using weights. In this step, the loss function is the sum of the deviations of each sub-objective's performance prediction from the minimum acceptable threshold multiplied by the sub-objective's weight. Loss Function: A function used to measure the difference between the model's predicted value and the target value. In this step, the loss function, the sum of the deviations of each sub-objective's performance prediction from the minimum acceptable threshold multiplied by the sub-objective's weight, guides the direction of parameter optimization. Sub-objective Weights: Weights assigned according to nuclear engineering safety priorities, reflecting the importance of different sub-objectives within the overall goal. Higher-weighted sub-objectives have a larger proportion in the loss function and are optimized first. Parameter Iterative Optimization: By iteratively adjusting the parameter values multiple times, the value of the loss function is gradually reduced until the preset optimization target or stopping condition is reached.
[0158] The specific process is as follows:
[0159] First, define the loss function. The loss function is the sum of the deviations of the predicted performance values of each sub-objective from the minimum acceptable threshold, multiplied by the weight of the sub-objective. For example, for the compressive strength sub-objective, its loss function term is (predicted compressive strength - target threshold) × the weight of that sub-objective. Next, initialize the parameters. Set the initial parameter values according to the initial parameter combination generated in step S520. For example, the initial parameter combination might be 10% MnO content in steel slag, 25% SiO2 content in fly ash, and a stirring rate of 100 rpm. Then, calculate the initial loss function value. Substitute the initial parameter values into the correlation equation to calculate the predicted performance values of each sub-objective, and then calculate the initial loss function value according to the definition of the loss function. For example, calculate the predicted performance values based on the initial parameter combination, and obtain the initial loss function value accordingly. Next, use a weighted gradient descent algorithm for iterative parameter optimization. Calculate the gradient of the loss function with respect to each parameter, and iteratively update the parameters in conjunction with the sub-objective weights. During iteration, based on the gradient information and weights, prioritize updating parameters that are more sensitive to the performance of high-weight sub-objectives. For example, if the compressive strength sub-objective has the highest weight, parameters that significantly affect it, such as the MnO content of the steel slag and the SiO2 content of the fly ash, are adjusted first. After each iteration, the performance estimate corresponding to the updated parameter combination is calculated using the correlation equation, and the loss function value is recalculated. For example, if the adjusted parameter combination is 11% MnO content in steel slag, 26% SiO2 content in fly ash, and a stirring rate of 105 rpm, a new performance estimate and loss function value are calculated. The above iterative process is repeated until the loss function value is lower than a preset optimization threshold. For example, when the loss function value converges to below the preset threshold, the iteration stops.
[0160] Step S535: Call the rule-based interpretation to perform parameter interaction verification. If the parameter combination after iteration triggers collaborative constraints, then adjust the associated parameters in a coordinated manner to conform to the quantitative structure-activity relationship rules.
[0161] Rule-based interpretation: A method for interpreting model predictions by generating concise rules to describe the model's decision-making process. In this step, rule-based interpretation verifies whether the adjusted parameter combinations conform to the synergistic relationships revealed by the quantified structure-activity relationship. Parameter interaction verification: The adjusted parameter combinations are verified using rule-based interpretation to ensure that the synergistic relationships between parameters conform to the quantified structure-activity relationship. Synergistic constraints: These refer to the interdependencies between parameters revealed by the quantified structure-activity relationship, which need to be considered and satisfied during parameter adjustment. For example, adjusting some parameters may affect the optimal values of other parameters. Linked adjustment: When a parameter combination triggers a synergistic constraint, the associated parameters are adjusted synchronously to ensure that the parameter combination conforms to the quantified structure-activity relationship.
[0162] The specific process is as follows:
[0163] First, the rule-based interpretation is invoked to validate the iterative parameter combinations. The rule-based interpretation describes the decision-making process for quantifying structure-property relationships by generating concise rules, which can be used to validate the rationality of parameter adjustments. For example, the rule-based interpretation might generate the following rule: "If the MnO content of steel slag increases, the SiO2 content of fly ash needs to be reduced accordingly to maintain stable compressive strength."
[0164] Next, check whether the parameter combination triggers synergistic constraints. Synergistic constraints refer to the interdependencies between parameters that need to be considered and satisfied during parameter adjustment. For example, if the adjusted steel slag MnO content is 12%, while the fly ash SiO2 content remains at 25%, this may trigger synergistic constraints because, according to the quantitative structure-property relationship, an increase in MnO content requires a corresponding decrease in SiO2 content.
[0165] Then, if the parameter combination triggers collaborative constraints, a coordinated adjustment is performed. Coordinated adjustment refers to the synchronous adjustment of related parameters to ensure that the parameter combination conforms to the quantified structure-property relationship. For example, based on the rules generated by the rule-based interpretation, the SiO2 content of fly ash is adjusted from 25% to 24% to maintain stable compressive strength.
[0166] Finally, the adjusted parameter combination was verified to conform to the quantitative structure-property relationship. By substituting the parameters into the correlation equation, the performance estimates corresponding to the adjusted parameter combination were calculated to ensure that the adjusted parameter combination not only met the synergistic constraints but also met the target performance requirements. For example, the adjusted parameter combination (steel slag MnO content 12%, fly ash SiO2 content 24%, stirring rate 105 rpm) corresponds to a compressive strength estimate of 46 MPa, a gamma-ray attenuation coefficient estimate of 0.57 / cm, and a durability estimate of 315 freeze-thaw cycles. These estimates were verified to meet the target performance requirements.
[0167] Step S536: When the parameter combination satisfies that the estimated performance of all sub-objectives is greater than or equal to the minimum threshold and does not exceed the dual constraints, stop the iteration and output the parameter combination as the reverse derivation result.
[0168] The specific process is as follows:
[0169] First, verify whether the parameter combination meets the performance forecast requirements of all sub-objectives. By substituting into the correlation equation, calculate the performance forecast corresponding to the current parameter combination, and check whether these forecasts meet the minimum compliance thresholds of all sub-objectives. For example, check whether the compressive strength is ≥45MPa, the gamma-ray attenuation coefficient is ≥0.55 / cm, and the durability is ≥300 freeze-thaw cycles.
[0170] Next, verify that the parameter combination does not exceed the dual constraints. Check whether the parameter values are within the hard constraints (the effective range of the parameters) and whether the single adjustment of the parameters meets the soft constraints (the limit on the parameter adjustment range). For example, check whether the MnO content of the steel slag is between 5% and 15%, and whether the single adjustment range of the stirring rate is ≤10 rpm.
[0171] Then, if the parameter combination simultaneously satisfies the performance forecast requirements of all sub-objectives and does not exceed the double constraints, the iteration stops. This indicates that the current parameter combination has been optimized to the point of meeting the target performance requirements, and the adjustment process complies with all constraints.
[0172] Finally, output the parameter combination as the reverse derivation result. Record the specific values of the parameters, the corresponding performance estimates, and the relevant derivation logic explanations to form a complete reverse derivation result. For example, output the parameter combination (steel slag MnO content 12%, fly ash SiO2 content 24%, stirring speed 105 rpm), the corresponding performance estimates (compressive strength 46 MPa, gamma-ray attenuation coefficient 0.57 / cm, durability 315 freeze-thaw cycles), and the derivation logic explanations.
[0173] Based on the same inventive concept, embodiments of the present invention provide a system for establishing a knowledge base of nuclear engineering concrete materials for high-value utilization of solid waste, including a memory and a processor. The memory stores information that can be run on the processor to implement, as described above. Figure 1 The procedure for the method shown.
[0174] The embodiments described in this specific implementation are preferred embodiments of this application and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A method for establishing a knowledge base for nuclear engineering concrete materials for high-value utilization of solid waste, characterized in that, include: External data is collected through a pre-set standardized search terminology database and intelligent search engine, while internal data is collected simultaneously to complete the integration and classification of dual-source heterogeneous data; Based on the preset tree-structured four-level parameter architecture, the integrated and classified data is cleaned in a structured manner, classified by parameter category, subcategory and indicator level, outliers are removed and data units are unified to form a standardized dataset. The standardized dataset is imported into a pre-defined database framework to build an initial database. A pre-defined machine learning model is used to learn the correlation between components, processes, and performance. Supplementary samples are generated to expand the data and obtain a complete database. By using a pre-defined gene coding model to mine and improve the structure-activity relationship of solid waste components, ratios, processes, and performance in the database, a quantitative structure-activity relationship between performance and solid waste components, ratios, and processes is established. Input the target performance requirements of nuclear engineering concrete, and based on the above-mentioned quantitative structure-property relationship and preset derivation constraints, automatically generate multiple sets of initial parameter combinations and perform reverse derivation to generate multiple initial optimization schemes; Substitute each initial optimization scheme into the quantified structure-property relationship to verify performance compliance; if it meets the standard, it is included in the candidate optimization scheme set. If the target is not met, the parameters are corrected according to the preset adjustment rules and the derivation is repeated until the preset number of candidate solutions are generated. Based on a preset screening index system, a preset weighted scoring model is used to score the candidate solution set, and the solution with the highest score is selected as the optimal solution. The database, structure-function relationship, candidate solution set, optimal solution and iteration process data are linked and stored according to the preset knowledge graph structure to build a knowledge base including data layer, model layer and solution layer.
2. The method for establishing a knowledge base for nuclear engineering concrete materials through high-value utilization of solid waste according to claim 1, characterized in that, The integration and classification of dual-source heterogeneous data includes: The collected external data is formatted and standardized. Natural language processing algorithms are used to extract key entities from unstructured text. Regular expressions are used to unify the data representation format. Key entities include solid waste type, performance indicators, and process description. The collected internal data is structured and transformed. The discrete data in the experimental records and production logs are reorganized according to the preset fields of solid waste name, detection date, physical property parameters, process parameters and characterization data. The Z-score method is used to preliminarily screen the reorganized data and remove invalid data that deviates from the preset normal distribution range. The cosine similarity algorithm is used to calculate the feature matching degree between key entities extracted from external data and preset fields of internal data, and a bidirectional correlation mapping is established between performance indicators mentioned in literature and experimental detection performance data, and between patent process descriptions and production process parameters. Based on the association mapping results, the dual-source data are classified according to solid waste type and nuclear engineering concrete application scenario. Redundant data with feature matching degree lower than the preset feature matching degree are removed, and finally a dual-source integrated dataset with clear classification and effective data is formed.
3. The method for establishing a knowledge base for nuclear engineering concrete materials through high-value utilization of solid waste according to claim 2, characterized in that, The process of structuring and transforming the collected internal data also includes blockchain traceability processing, including: The discrete data in the experimental records and production logs are reorganized according to the preset fields of solid waste name-blockchain traceability code-test date-physical property parameters-process parameters-characterization data, and a blockchain traceability code is added on the basis of the original fields; The blockchain traceability code links solid waste data across the entire supply chain, including origin identification, transportation process information, original testing data, and warehousing information. Smart contracts are used to encrypt and store the data across the entire supply chain to ensure it is tamper-proof. When using the Z-score method to initially screen the recombined data, internal data without blockchain traceability codes or with incomplete traceability data are simultaneously removed, while only structured data with complete traceability information is retained.
4. The method for establishing a knowledge base for nuclear engineering concrete materials through high-value utilization of solid waste according to claim 1, characterized in that, The standardized dataset is imported into a pre-defined database framework to construct an initial database. A pre-defined machine learning model is used to learn the correlation between composition, process, and performance. Supplementary samples are generated to expand the data, resulting in a complete database including: The standardized dataset is classified into four levels of parameters: solid waste properties, production process, material properties, and service environment, and imported into a pre-defined relational database framework. This framework defines the field types, units of measurement, and cross-parameter association rules for each parameter through a data dictionary, and enables fast data retrieval through an index structure, thereby constructing a well-structured initial database that can be directly used for model training. Using data from the initial database as training samples, a BP neural network is used as the preset machine learning model. Solid waste composition parameters and process parameters are used as input layers, and material performance parameters are used as output layers. The network weights are iteratively adjusted through the backpropagation algorithm to learn the nonlinear correlation between composition, process and performance until the model's prediction error on the validation set is lower than the preset threshold, thus forming a trained model that can accurately map the parameter correlation. Based on the correlation patterns captured by the trained model and the sample distribution characteristics of the initial database, a generative adversarial network algorithm is used to generate supplementary samples: the generator is trained with the real samples in the initial database as a benchmark to generate virtual samples that conform to the correlation patterns of composition-process-performance; at the same time, the consistency between the virtual samples and the real samples is verified by the discriminator, and only the virtual samples whose probability of being judged as real samples reaches a preset threshold are temporarily stored as candidate supplementary samples. By integrating the initial database with candidate supplementary samples that have passed the discriminator verification, the K-means clustering algorithm is used for consistency verification. Abnormal samples with clustering deviation values exceeding the preset range are removed, ultimately forming a complete database with balanced data distribution and significant correlation patterns.
5. The method for establishing a knowledge base for nuclear engineering concrete materials with high-value solid waste utilization according to claim 1, characterized in that, By using a pre-defined gene coding model to mine and refine the structure-activity relationship of solid waste components, proportions, processes, and performance in the database, a quantitative structure-activity relationship between performance and solid waste components, proportions, and processes is established, including: Data with dynamic association tags were extracted from the complete database, and a four-dimensional feature matrix was constructed according to the dimensions of solid waste composition, proportioning parameters, process parameters, and performance indicators. The dynamic association tags recorded the source information of the data and the causal relationship information between the parameters. Based on the four-dimensional feature matrix, the parameter encoding stage is entered: solid waste components, ratio parameters, and process parameters are converted into standardized gene sequences, and causal association information in the dynamic association tags is embedded into the corresponding coding positions to form a gene coding dataset with causal tags. Using the encoded gene dataset as input, the relationship mining stage begins: Ridge regression algorithm is used to control parameter collinearity and quantify the linear correlation coefficient between a single parameter and performance index; simultaneously, random forest algorithm is used to mine nonlinear relationships of multi-parameter interactions, and the results of linear and nonlinear analysis are verified and corrected based on the embedded causal association information, outputting complete structure-activity relationship coefficients; Based on the structure-property relationship coefficients obtained in the relationship mining stage, the mapping generation stage begins: using solid waste composition, ratio, and process parameters as independent variables and performance indicators as dependent variables, a function mapping model is constructed. Through this model, the quantitative correspondence between each parameter combination and the corresponding performance indicator is clarified, forming a quantitative structure-property relationship between performance and solid waste composition, ratio, and process that integrates data sources and causal relationships.
6. The method for establishing a knowledge base for nuclear engineering concrete materials through high-value utilization of solid waste according to claim 5, characterized in that, Following the formation of a quantitative structure-property relationship between performance and solid waste composition, proportion, and process that integrates data sources and causal correlations, the following steps are also included: Based on the established quantitative structure-property relationship, its application boundary is first calibrated: combined with the industry standard for nuclear engineering concrete, the effective application range of each parameter in the quantitative structure-property relationship is defined, and invalid parameter associations that exceed the boundary are eliminated to obtain the effective quantitative structure-property relationship within the boundary. For valid quantified structure-property relationships within the boundary, initiate causal interpretability verification: use SHAP values to quantify the impact weight of each parameter on performance, generate rule-based interpretations through LIME, and compare them with preset rules in the nuclear engineering expert database; If the interpretation results conflict with the expert rules, the original data in the improved database is traced back through dynamic association labels. After correcting the conflicting items, the quantitative structure-activity relationship is re-optimized until the quantitative structure-activity relationship within the boundary simultaneously meets the application boundary requirements and the expert causal knowledge. The validated quantified structure-property relationships are standardized and encapsulated according to the parameter effective range, performance compliance threshold, causal explanation label, and source traceability code to form interpretable quantified structure-property relationships with boundary constraints, and stored in the knowledge base model layer.
7. The method for establishing a knowledge base for nuclear engineering concrete materials through high-value utilization of solid waste according to claim 6, characterized in that, Inputting the target performance requirements of nuclear engineering concrete, based on the above-mentioned quantified structure-property relationship and preset derivation constraints, multiple sets of initial parameter combinations are automatically generated and reverse-derived to generate multiple initial optimization schemes, including: Call the interpretable quantified structure-property relationship with boundary constraints in the knowledge base model layer, and extract the effective range of parameters, performance achievement threshold and causal explanation label associated with the target performance requirements; Based on the effective range of the extracted parameters, the Latin hypercube sampling method is used to generate initial parameter combinations: within the effective range of the parameters of the effective quantitative structure-property relationship within the boundary, combinations of solid waste components, proportioning parameters, and process parameters are uniformly selected. Each combination must meet the safety constraints preset by the nuclear engineering expert database to ensure that all initial combinations fall within the solution space covered by the effective quantitative structure-property relationship within the boundary, and a preset number of initial parameter combinations are generated. Using the target performance requirement as the output constraint, the initial parameter combination is inversely derived based on the function mapping relationship in the interpretable quantified structure-property relationship; After completing the reverse derivation for each set of initial parameter combinations, the corresponding initial optimization scheme is compiled. The scheme includes the specific values of the parameters, the performance target estimate, and the derivation logic explanation based on the causal explanation label. All initial optimization schemes are compiled to form a preset set of initial optimization schemes.
8. A system for establishing a knowledge base for nuclear engineering concrete materials through high-value utilization of solid waste, characterized in that, The system includes a memory, a processor, and a program stored in the memory and executable on the processor, which, when loaded and executed by the processor, implements a method for establishing a knowledge base for high-value nuclear engineering concrete materials from solid waste as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Knowledge representation model creation method
CN110516808A
Large language model special for solid waste cementing material and innovative hypothesis generation method of large language model
CN120632038A
Electric power design knowledge base construction method fusing multi-modal data and RAG technology
CN120929611A