Method and system for constructing material database based on high-throughput experimental multi-modal data
By standardizing data acquisition and processing, and combining multi-device synchronization and a three-layer logical architecture, the problem of integrating high-throughput experimental multimodal data was solved, an efficient materials database was built, the standardization and reliability of data were achieved, and the efficiency and accuracy of new materials research and development were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI UNIV
- Filing Date
- 2025-05-27
- Publication Date
- 2026-05-26
AI Technical Summary
Existing methods for constructing materials databases cannot effectively integrate multimodal data generated by high-throughput experiments, especially dynamic process data, resulting in insufficient data synchronization, heterogeneity, and intrinsic correlation, which affects the integrity of the data and the accuracy of the prediction model.
A standardized experimental data acquisition protocol is adopted. Through sample encoding, multi-device time synchronization and adaptive acquisition frequency, infrared thermal imaging and sensor time series data are acquired and processed to extract structured temperature features. A material database is constructed using a three-layer logic architecture and a hybrid storage strategy.
A materials database with standardized data, accurate features, and easy management has been constructed, supporting efficient data retrieval and performance prediction, and improving the efficiency of new material screening and the reliability of the model.
Smart Images

Figure CN120523797B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of materials database construction technology, and in particular to a method and system for constructing materials databases based on high-throughput experimental multimodal data. Background Technology
[0002] Materials databases, as a core component of materials genome engineering, aim to systematically collect, store, manage, and share multi-dimensional information on materials, including their composition, structure, processing, and properties. By constructing comprehensive materials databases and combining them with technologies such as data mining and machine learning, the research and development of new materials can be accelerated, material properties optimized, and intelligent material design and screening achieved. Traditional materials database construction often relies on literature collection, manual data entry, and the integration of various experimental data. Its core lies in establishing effective data models to describe material information and support efficient data retrieval and analysis, enabling researchers to quickly locate and compare the properties of different materials.
[0003] With the widespread application of high-throughput experimental techniques in new materials fields such as catalysis, energy, and biology, the rate and volume of experimental data generation have increased dramatically, exhibiting typical multimodal characteristics. These data are not only high-dimensional and highly heterogeneous, but also require extremely high standards for temporal synchronization, spatial correspondence, and data integrity verification to ensure accurate capture of subtle changes in materials during dynamic reaction processes and their correlation with performance. Current methods for constructing materials databases are insufficient to meet these requirements.
[0004] Chinese invention patent application CN116049160A discloses a method for constructing a machine learning database based on multi-series aluminum alloy data. This method, targeting the aluminum alloy field, integrates data from multiple heat-treatable aluminum alloy series, such as Al-Zn-Mg-Cu, Al-Cu, and Al-Li, to address the issues of data centrality and scarcity of specific element data in single-series modeling. Its database construction primarily relies on obtaining alloy composition, processing techniques (such as heat treatment regimes), and corresponding material property data from existing published literature. This method effectively addresses the characteristics of the aluminum alloy field, namely the complex relationship between the properties, composition, and processing of multiple alloy series. It employs data cleaning and utilizes machine learning algorithms (such as support vector regression) to perform boundary learning on the cleaned data to fill in data gaps, attempting to construct a three-dimensional integrated dataset encompassing alloy composition, heat treatment regimes, and material properties. However, the technical means of constructing the database in this method mainly rely on the secondary processing of existing literature data and machine learning to fill it in. This may result in insufficient data in terms of the granularity of the original experimental process, the synchronicity of multi-source heterogeneous data (especially dynamic process data), and the in-depth mining of the intrinsic correlation. At the same time, its data acquisition method also limits the direct integration and utilization of real-time, multimodal raw data (such as in-situ characterization data and sensor network time series data) during the experiment. The accuracy and reliability of the data filled in by machine learning also strongly depend on the quality and coverage of the initial literature data. If the original literature data itself lacks a comprehensive record of complex experimental details, it may affect the integrity of the final database and the accuracy of the prediction model. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for constructing a materials database for the large amount of complex and heterogeneous multimodal data generated by high-throughput experiments. It aims to overcome the shortcomings of existing technologies in terms of data acquisition standardization, multi-source data synchronization and integration, accuracy of key feature extraction, and database structuring and management efficiency, thereby efficiently and systematically constructing a comprehensive and highly available materials database that includes sample information, experimental conditions, process data, core performance characteristics, and raw data indexes.
[0006] The technical solution of this invention is implemented as follows:
[0007] On the one hand, this invention provides a method for constructing a materials database based on high-throughput experimental multimodal data, including:
[0008] S1. Establish a standardized experimental data acquisition protocol, including: assigning unique identifiers to samples using a sample coding system and recording the chemical composition, preparation method, and channel location information of the samples; adopting a multi-device time synchronization mechanism; and setting adaptive acquisition frequency parameters for the infrared thermal imager.
[0009] S2. Based on the standardized experimental data acquisition protocol, synchronously collect and record experimental metadata, infrared thermal imaging data and sensor time series data, and perform integrity verification on the collected data to obtain the original multimodal experimental dataset with timestamps and verification.
[0010] S3. Process the infrared thermal imaging data in the original multimodal experimental dataset. Through image preprocessing, channel calibration and feature calculation, extract the structured temperature features of each catalyst channel. The structured temperature features include time-series temperature features and spatial temperature distribution features.
[0011] S4. Integrate experimental metadata and sensor time-series data from the original multimodal experimental dataset, and combine them with structured temperature features. Through time alignment and standardization, generate standardized sample-experiment-performance correlation data records.
[0012] S5. Based on the standardized sample-experiment-performance correlation data records and the index information of the original multimodal experimental dataset, a three-layer logical architecture and a hybrid physical storage strategy are adopted to construct a material database.
[0013] Preferably, the sample coding system includes assigning a globally unique sample identifier to each catalyst sample and recording the sample's chemical composition, preparation method, and channel position coordinates in the screening device; the multi-device time synchronization mechanism includes using a network time protocol to synchronize the infrared thermal imager, temperature sensor, mass flow controller, pressure sensor, and central control system at the millisecond level before the experiment; the adaptive acquisition frequency parameters of the infrared thermal imager include reducing the acquisition frequency during the stable reaction phase and increasing the acquisition frequency during the temperature abrupt change phase.
[0014] Preferably, the experimental metadata that is synchronously collected and recorded includes: a unique experimental identifier, experimental date, experimental personnel information, experimental device identifier, sample batch information, experimental batch information, temperature program setting value, target flow rate and reaction pressure of each gas component; the infrared thermal imaging data includes raw temperature matrix data recording the temperature changes of the catalyst channel; and the sensor time-series data includes platform thermocouple temperature value, gas flow rate value and pressure value.
[0015] Preferably, the image preprocessing includes geometric correction of the infrared image sequence; the channel calibration includes aligning each frame of infrared image with the template through image registration using a pre-stored standard channel template image of the screening device, and accurately segmenting the region of interest for each channel; the feature calculation includes extracting temporal and spatial distribution features of the temperature data within the region of interest for each channel.
[0016] Preferably, the temporal temperature characteristics include ignition time, ignition temperature, maximum temperature rise, peak temperature and its arrival time, average temperature during the stable reaction phase, and standard deviation of temperature fluctuation; the spatial temperature distribution characteristics include the temperature standard deviation and hot spot area ratio within the region of interest.
[0017] Preferably, the ignition time in the time-series temperature characteristics is calculated using a normalized dynamic ignition criterion. To determine, when The ignition time is determined when the time continuously exceeds a preset ignition criterion threshold for at least a preset time period. Calculated using the following formula:
[0018] ,
[0019] in, Let be the channel average temperature at time i. This represents the average channel temperature at the previous moment. Let i be the time interval between time i and time i-1. To normalize the temperature rise rate parameter, For reference baseline temperature, For characteristic activation temperature parameters, It is an exponential adjustment factor.
[0020] Preferably, the time alignment includes precisely aligning the infrared thermal imaging data and the sensor time-series data on the time axis based on high-precision timestamps; the standardization process includes converting data from different sources into a unified International System of Units (SI), eliminating sensor fault noise points, extracting statistics of time-varying parameters in key reaction stages, and forming standardized records.
[0021] Preferably, the three-layer logical architecture includes a basic metadata layer, a derived feature data layer, and a raw data index layer; the hybrid physical storage strategy includes storing structured metadata and feature data through a relational database, and storing large-capacity raw infrared thermal imaging video files and sensor log files through a distributed object storage system.
[0022] Preferably, the basic metadata layer includes a sample information table, a preparation method table, an experimental configuration table, and a channel mapping table; the derived feature data layer includes a catalyst performance table and a time series data summary table; the raw data index layer includes a raw data registry and a keyframe table; and a standardized application programming interface is provided to support data manipulation and integration.
[0023] Another method, the present invention also provides a material database construction system based on high-throughput experimental multimodal data, the system being used to implement the method described in any of the above claims, the system comprising:
[0024] The data protocol management module is used to establish standardized experimental data acquisition protocols. It includes a sample coding submodule, which is used to assign unique identifiers to samples and record the chemical composition, preparation method and channel position information of the samples; a time synchronization submodule, which is used to realize millisecond-level clock synchronization of experimental equipment; and an acquisition frequency control submodule, which is used to set the adaptive acquisition frequency parameters of the infrared thermal imager.
[0025] The multimodal data acquisition module is used to synchronously acquire and record experimental metadata, infrared thermal imaging data and sensor time-series data according to a standardized experimental data acquisition protocol, as well as to perform data integrity verification and generate a timestamped and verified original multimodal experimental dataset.
[0026] The infrared image processing module is used to process infrared thermal imaging data in the original multimodal experimental dataset. It includes an image preprocessing unit for geometric correction of infrared image sequences; a channel calibration unit for accurately aligning infrared images with standard channel templates and segmenting channel regions; and a feature calculation unit for extracting structured temperature features, including temporal temperature features and spatial temperature distribution features.
[0027] The data integration and association module is used to integrate experimental metadata and sensor time-series data from the original multimodal experimental dataset, and combine structured temperature features to generate standardized sample-experiment-performance association data records through time alignment and standardization.
[0028] The database management module is used to construct and maintain a materials database based on standardized sample-experiment-performance related data records and index information of the original multimodal experimental dataset, using a three-layer logical architecture and a hybrid physical storage strategy. The three-layer logical architecture includes a basic metadata layer, a derived feature data layer, and an original data index layer.
[0029] The present invention has the following advantages over the prior art:
[0030] (1) This invention systematically collects, processes, integrates, and stores multimodal data generated from high-throughput experiments, thereby constructing a materials database that is standardized, accurate in features, clearly structured, and easy to manage. This database can effectively support data retrieval, performance comparison, and model building in materials research and development, thereby improving the screening efficiency of new materials and the reliability of performance prediction;
[0031] (2) This invention implements a standardized data acquisition protocol during the data acquisition stage, including assigning a globally unique sample identifier to each catalyst sample and recording its channel position information, using the network time protocol to synchronize multiple devices at the millisecond level clock, and dynamically adjusting the acquisition frequency of the infrared thermal imager according to the reaction stage, thereby ensuring the accuracy, spatiotemporal correspondence and integrity of the original data from the source.
[0032] (3) The present invention employs image preprocessing and calibration techniques, including geometric correction, channel precision calibration based on pre-stored standard template images and region of interest segmentation, for processing infrared thermal imaging data. Based on these techniques, the temporal and spatial distribution characteristics are calculated, thereby improving the accuracy and reliability of performance indicators obtained from complex infrared data.
[0033] (4) This invention precisely aligns infrared thermal imaging data and sensor time-series data on the time axis based on high-precision timestamps, converts data from different sources into a unified international unit system, eliminates sensor fault noise points, and extracts statistics of time-varying parameters in key reaction stages, forming a standardized sample-experiment-performance correlation record. This systematic data integration and standardized processing effectively solves the inconsistency problem of multi-source heterogeneous data in terms of format, units, and time reference, ensuring the correct correlation and comparability between data of different dimensions;
[0034] (5) This invention employs a three-layer logical architecture and a hybrid physical storage strategy combining relational databases and distributed object storage to construct the material database. Furthermore, the catalyst performance table in the derived feature data layer explicitly stores key performance parameters such as ignition time determined by the normalized dynamic ignition criterion. This specific database architecture design and storage method not only efficiently manages and stores structured metadata, derived feature data, and large-capacity raw data files, ensuring data integrity and accessibility, but also optimizes data retrieval efficiency through clear logical layering and direct storage of key features, facilitating users to quickly obtain the required information and conduct in-depth analysis. Attached Figure Description
[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 This is a flowchart of the method of the present invention;
[0037] Figure 2 This is a diagram illustrating the technical implementation of the present invention;
[0038] Figure 3 This is a schematic diagram of data acquisition according to the present invention;
[0039] Figure 4 This is a schematic diagram of the feature extraction process of the present invention;
[0040] Figure 5 This is a schematic diagram of the data record generation process of the present invention;
[0041] Figure 6 This is a schematic diagram of the system modules of the present invention. Detailed Implementation
[0042] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0043] like Figure 1 As shown, this invention provides a method for constructing a materials database based on high-throughput experimental multimodal data, including:
[0044] S1. Establish a standardized experimental data acquisition protocol, including: assigning unique identifiers to samples using a sample coding system and recording the chemical composition, preparation method, and channel location information of the samples; adopting a multi-device time synchronization mechanism; and setting adaptive acquisition frequency parameters for the infrared thermal imager.
[0045] S2. Based on the standardized experimental data acquisition protocol, synchronously collect and record experimental metadata, infrared thermal imaging data and sensor time series data, and perform integrity verification on the collected data to obtain the original multimodal experimental dataset with timestamps and verification.
[0046] S3. Process the infrared thermal imaging data in the original multimodal experimental dataset. Through image preprocessing, channel calibration and feature calculation, extract the structured temperature features of each catalyst channel. The structured temperature features include time-series temperature features and spatial temperature distribution features.
[0047] S4. Integrate experimental metadata and sensor time-series data from the original multimodal experimental dataset, and combine them with structured temperature features. Through time alignment and standardization, generate standardized sample-experiment-performance correlation data records.
[0048] S5. Based on the standardized sample-experiment-performance correlation data records and the index information of the original multimodal experimental dataset, a three-layer logical architecture and a hybrid physical storage strategy are adopted to construct a material database.
[0049] like Figure 2 As shown, the technical implementation process of this invention includes: 1) Establishment and execution of a standardized data acquisition protocol, that is, firstly, a globally unique sample identifier is set for each catalyst sample and its composition, preparation method, and precise channel position in the screening device are recorded. Before the experiment, the infrared thermal imager, various sensors, and the central control system are synchronized at the millisecond level using the network time protocol. At the same time, the adaptive acquisition frequency of the infrared thermal imager is set so that it can reduce the frequency during the stable reaction stage and increase the frequency during the rapid temperature change stage; 2) Synchronous acquisition and preliminary verification of multimodal experimental data, that is, detailed experimental metadata, the original temperature matrix sequence of infrared thermal imaging, and sensor time series data are synchronously acquired and recorded during the experiment, and preliminary integrity verification is performed after acquisition; 3) The deep processing and key feature extraction of infrared thermal imaging data involves geometric correction of the infrared image sequence, precise segmentation of the region of interest for each catalyst sample channel using pre-stored standard channel template images through image registration, and calculation of the temporal temperature characteristics for each channel, particularly the ignition time. This ignition time is determined by calculating a normalized dynamic ignition criterion, which comprehensively considers the temperature rise rate and the normalization degree of the current temperature relative to the reference baseline and characteristic activation temperature, and introduces an exponential adjustment factor. The ignition moment is determined when this criterion continuously exceeds a preset threshold for a preset time. In addition, other temporal features such as ignition temperature, maximum temperature rise, peak temperature, and spatial temperature distribution features such as hot spot area ratio are extracted; 4) The integration, standardization, and formation of associated records of multi-source data involves integrating extracted structured temperature features with experimental metadata and sensor time-series data. Based on high-precision timestamps, infrared thermal imaging data and sensor time-series data are precisely aligned on the time axis. Unit unification, noise removal, and extraction of time-varying parameter statistics for key reaction stages are performed. Finally, standardized and structured sample-experiment-performance associated data records containing sample information, experimental conditions, processed sensor data, and key performance characteristics are formed. 5) Construction and management of the material database: a three-layer logical architecture consisting of a basic metadata layer, a derived feature data layer, and a raw data index layer is adopted. For physical storage, a relational database is used to store structured metadata and derived feature data. A distributed object storage system is used to store a large capacity of raw infrared video and sensor logs. Key performance parameters such as ignition time determined according to normalized dynamic ignition criteria are stored in the catalyst performance table of the derived feature data layer. Finally, a standardized application programming interface is provided to support data operation and application analysis of the database.
[0050] Specifically, such as Figure 3As shown, in one embodiment of the present invention, step S1 includes:
[0051] A sample coding system was constructed, including assigning a globally unique sample identifier to each catalyst sample and recording the sample's chemical composition, preparation method, and channel position coordinates in the screening device. This sample identifier is designed as a structured code, integrating, for example, project code, experimental batch number, material classification code, and the sample's serial number within that batch. It can further incorporate its physical location information in the screening device, such as reaction plate number and well row and column number. This structured identifier not only ensures the absolute uniqueness of each sample but also embeds easily parsed and preliminarily screenable metadata within the identifier itself. Simultaneously recorded with the sample identifier is detailed chemical composition information, accurate to the percentage content, purity, and source of each component, as well as the precise three-dimensional coordinates of the sample's channel position in the high-throughput screening device.
[0052] A multi-device time synchronization mechanism is employed, including millisecond-level clock synchronization of the infrared thermal imager, temperature sensors, mass flow controllers, pressure sensors, and central control system using network time protocols before the experiment. Specifically, before each round of high-throughput experiments officially begins, the central control system deployed on the experimental platform will perform mandatory clock synchronization calibration on all devices involved in data acquisition using network time protocols, such as NTP, or, in scenarios with higher time accuracy requirements, the precise time protocol PTP / IEEE 1588 standard. These devices include the infrared thermal imager capturing the temperature distribution on the sample surface, the thermocouples or resistance temperature sensors in each channel used to monitor the temperature at specific points, the mass flow controller precisely controlling the feed of the reaction gas, the pressure sensor monitoring the environmental pressure of the reaction system, and the central data recording and control system itself responsible for summarizing and initially processing the data. The goal is to control the timestamp error of all devices to the millisecond level or even lower, thereby providing a unified and reliable time reference for subsequent multimodal data fusion analysis.
[0053] As one implementation method, a hardware synchronization pulse signal is designed, which is uniformly issued by the central control system when the experiment starts or a specific triggering event occurs. This signal is distributed in parallel to all key data acquisition devices, serving as a common zero-time reference point or periodic calibration signal. This further improves the time synchronization accuracy of data recording across devices and effectively eliminates the slight time deviation that may be introduced by factors such as network latency jitter.
[0054] The infrared thermal imager's adaptive acquisition frequency parameters are set, including reducing the acquisition frequency during the stable reaction phase and increasing it during temperature abrupt changes. The aim is to achieve a balance between fully capturing key dynamic features of the reaction process and effectively controlling the total amount of data stored. Specifically, when the infrared thermal imager determines, through real-time image analysis or correlated sensor data, that the reaction process of all sample channels is in a relatively stable phase—that is, the temperature change rate is low or within a preset stability threshold—the system automatically adjusts the image acquisition frequency to a lower baseline level, such as one frame per second or lower, to avoid generating a large amount of redundant thermal imaging data during the stable reaction period. However, once the system detects that the temperature of any one or more sample channels begins to change drastically, or when a temperature abrupt change phase, such as catalyst ignition, extinguishing, or intense exothermic or endothermic reactions, is anticipated through a preset experimental procedure, the infrared thermal imager's acquisition frequency is rapidly and automatically increased to a preset higher level, such as tens of frames per second or even higher, ensuring that the entire dynamic evolution and subtle features of these key reaction events can be accurately recorded with sufficient temporal resolution.
[0055] In this embodiment, the acquisition frequency of the infrared thermal imager is... Designed to be based on real-time monitoring of the average temperature change rate of the sample area The function is dynamically adjusted, as follows:
[0056] ,
[0057] The calculation result It will be constrained to a preset minimum sampling frequency. and highest sampling frequency between. The minimum data collection frequency that represents the plateau period of the response. This represents the highest acquisition frequency allowed by the instrument or data processing capabilities. This is an adjustable sensitivity coefficient used to amplify or reduce the influence of the temperature change rate on the sampling frequency. It refers to the rate of temperature change obtained by real-time differential calculation of temperature data of a specific region of interest of an infrared thermal imager, while abs represents the absolute value of the rate of temperature change.
[0058] Specifically, in one embodiment of the present invention, in step S2, the experimental metadata that is synchronously collected and recorded includes: a unique experimental identifier, experimental date, experimental personnel information, experimental device identifier, sample batch information, experimental batch information, temperature program setting value, target flow rate and reaction pressure of each gas component; the infrared thermal imaging data includes raw temperature matrix data recording the temperature change of the catalyst channel; and the sensor time series data includes platform thermocouple temperature value, gas flow rate value and pressure value.
[0059] Specifically, after being entered, the metadata is stored in a structured file format, JSON or XML. After storage, a hash value is generated for each set of metadata records.
[0060] In this embodiment, during the experiment, the infrared thermal imager continuously captures infrared radiation from the sample area at an adaptive acquisition frequency and converts it into a two-dimensional temperature matrix sequence, i.e., the raw thermal imaging video stream. The acquired data is recorded in raw binary format with detailed timestamps or in a standard image sequence format.
[0061] In this embodiment, sensor time-series data is collected at a preset higher fixed frequency or an equally configurable adaptive frequency, each data point is accompanied by a timestamp, and the data is stored in CSV format.
[0062] In this embodiment, after the initial completion of the data acquisition task, a data verification process is also included to quickly assess the quality of the acquired data. The verification includes checking whether all expected experimental metadata fields have been correctly filled and their values are within a reasonable range; verifying whether the infrared thermal imaging data file is complete, whether the file size meets expectations, and whether it can be correctly parsed; confirming the existence of all sensor time-series data streams, whether the timestamps are continuous and aligned with the start and end times of the experiment, and whether the sensor readings are within the preset physical meaning or safety threshold range.
[0063] Specifically, set a data health score. This is used to quantify the initial quality of the data. If the verification finds any missing, corrupt, or obviously abnormal data, such as a data health score below a preset threshold, the system will automatically generate an alarm message and mark the relevant data records, prompting the experimenters to conduct manual review or decide whether to repeat part of the experiment or data collection.
[0064] In a specific example, data health score The calculation process is as follows:
[0065] ,
[0066] In the formula, , , These represent the quality scores for metadata, infrared data, and sensor data, respectively. , , These are the corresponding weighting coefficients.
[0067] ,
[0068] In the formula, This indicates the total number of fields in the metadata. This indicates the number of fields that have been successfully populated. It is a hyperparameter that controls the degree of penalty for missing fields. It is the quality score of the i-th field, with a value range of [0,1]. It is evaluated by verifying whether the value of this field is within the expected range, whether the format is correct, and whether it is logically consistent with other related fields. , These are the sensitivity parameter and threshold parameter for quality assessment.
[0069] ,
[0070] In the formula, This represents the actual number of infrared image frames acquired. To determine the expected number of frames to be obtained based on the experiment duration and the set acquisition frequency, This is an exponential parameter that adjusts the integrity weight. The second term assesses the continuity of infrared data, where... It is the total number of time sampling points. This indicates key data features at time t, such as the average temperature value. This is a parameter that controls the importance of continuity. The third parameter evaluates the signal-to-noise ratio (SNR). , These are the sensitivity parameter and reference threshold for signal-to-noise ratio evaluation, respectively.
[0071] ,
[0072] In the formula, K is the total number of sensors. It represents the number of outliers for the k-th sensor. This is the total number of data points of the sensor. The penalty level for controlling the proportion of outliers. The second evaluation item assesses the collaborative consistency between sensors. It is the average absolute deviation of the correlation coefficients between all sensor pairs. , These are the reference threshold and sensitivity parameter for controlling collaborative evaluation, respectively. The third term introduces KL divergence from information theory. The difference between the sensor data distribution P and the expected theoretical distribution Q is measured. Control the weight of this item.
[0073] In the calculation Then, compare it with a preset threshold. In comparison, if If this occurs, an alarm mechanism will be triggered. The alarm level is determined based on the following formula:
[0074] ,
[0075] The alarm levels are quantified into 1 to 4 levels according to this formula, with level 4 being the most serious, requiring immediate intervention from the experimenters and possibly requiring the experiment to be repeated, while level 1 is only a warning. This represents the floor function.
[0076] Specifically, such as Figure 4 As shown, in one embodiment of the present invention, step S3, the processing of infrared thermal imaging data includes image preprocessing, channel calibration, and feature calculation. The image preprocessing includes geometric correction of the infrared image sequence; the channel calibration includes aligning each frame of the infrared image with the template using a pre-stored standard channel template image from a screening device, accurately segmenting the region of interest for each channel; the feature calculation includes extracting temporal and spatial distribution features of the temperature data within the region of interest for each channel. The temporal temperature features include ignition time, ignition temperature, maximum temperature rise, peak temperature and its arrival time, average temperature during the stable reaction phase, and standard deviation of temperature fluctuation; the spatial temperature distribution features include the standard deviation of temperature within the region of interest and the hot spot area ratio.
[0077] In a specific example, step S3 is implemented as follows:
[0078] S3.1 Image Preprocessing
[0079] Geometric correction is applied to the acquired infrared image sequence. This process utilizes distortion parameters pre-obtained by photographing a standard calibration target plate to establish and apply a polynomial transformation model, inversely mapping the pixel coordinates of each frame. This step corrects image distortions introduced by factors such as infrared thermal imager lens optical distortion, non-perpendicular mounting angles, and sample stage tilt, ensuring the correct spatial relationships and geometric shapes of objects in the image. Furthermore, a mean square filter is applied to smooth the corrected image to suppress sensor noise and preserve image edges and details, thereby improving the image signal-to-noise ratio.
[0080] S3.2 Channel Calibration
[0081] Based on the geometrically corrected infrared images, precise calibration of catalyst sample channels and automatic segmentation of the Region of Interest (ROI) for each channel are performed. This process employs image registration based on the Scale Invariant Feature Transform (SIFT) algorithm, precisely aligning each frame of real-time infrared image with a pre-stored standard channel template image corresponding to the high-throughput screening device. The standard channel template image is pre-made with idealized geometric contours, center positions, and unique logical numbers for each catalyst sample channel. Through image registration, a precise affine or perspective transformation matrix is determined. This transformation matrix is used to accurately map the predefined channel region masks or coordinate lists on the template image onto the real-time infrared image, thereby automatically and accurately segmenting the ROI actually occupied by each catalyst sample. As an implementation method, a deep learning instance segmentation model based on the Mask R-CNN architecture, pre-trained on a large dataset of infrared images with manually labeled channel positions and contours, is used to directly identify and segment the precise boundaries of each sample channel from the input infrared image. This method provides higher segmentation accuracy and robustness even under conditions of tight channel spacing, partial occlusion, or complex backgrounds.
[0082] S3.3 Feature Calculation
[0083] For each precisely segmented sample channel ROI, its temporal temperature characteristics and spatial temperature distribution characteristics are calculated. The ignition time is determined by calculating the following Normalized Dynamic Ignition Criterion (NDIC):
[0084] ,
[0085] in, The normalized dynamic ignition criterion value calculated at time i. Let i be the average temperature within the ROI of the sample channel. This represents the average channel temperature at the previous moment. Let i be the time interval between time i and time i-1, which is the sampling period of the infrared thermal imager. To normalize the temperature rise rate parameter, As a baseline temperature, the average ROI temperature during a steady-state period after the start of the reaction was taken. The characteristic activation temperature parameter is a typical temperature threshold set based on prior knowledge of the catalyst system, representing the temperature at which the catalyst is significantly activated or undergoes a violent reaction. This is an exponential adjustment factor, set to a constant greater than 1, used to amplify the criterion response when the temperature is close to or exceeds the characteristic activation temperature.
[0086] The preset minimum number of confirmation points is For example, the number of sampling frames corresponding to 0.5 seconds. When the calculated... Value continuous All sampling points consistently exceeded the preset judgment threshold. When that happens, the first one that meets this condition will be selected. The time point corresponding to the value The ignition time of the catalyst sample was determined. This NDIC criterion integrates the rate of temperature rise and the normalization of the current temperature relative to the reference baseline and characteristic activation temperature, and introduces an exponential adjustment factor to robustly and adaptively identify the ignition timing.
[0087] In addition to ignition time, other extracted temporal features include: ignition temperature (i.e., the average temperature within the ROI at the determined ignition time point), maximum temperature rise (the difference between the peak temperature recorded within the ROI during the experiment and the initial baseline temperature), peak temperature (the highest temperature value recorded within the ROI during the experiment) and its arrival time.
[0088] Extracting spatial temperature distribution features involves calculating the temperature standard deviation and hotspot area ratio within the Region of Interest (ROI). The temperature standard deviation within the ROI characterizes the uniformity of temperature distribution. To further quantify the non-uniformity of spatial temperature distribution, the hotspot area ratio is calculated: a dynamic high-temperature threshold is set (e.g., the average temperature within the ROI at the current time point plus twice its standard deviation). All pixel regions within the ROI with temperatures exceeding this threshold are identified as hotspot regions. The ratio of the total area of these hotspot regions to the total area of the ROI is the hotspot area ratio. This ratio reflects the proportion of high-temperature regions on the sample surface.
[0089] Specifically, such as Figure 5 As shown, in one embodiment of the present invention, step S4 includes:
[0090] The sensor time-series data collected in step S2, such as platform thermocouple temperature values, gas flow rates, and pressure values, are precisely matched and aligned with the various time-series temperature features extracted from the infrared thermal imaging data in step S3, such as ignition time, ignition temperature, maximum temperature rise, peak temperature, and their arrival time, on a unified time axis. To address the potential inconsistency in sampling frequencies between different data sources, an interpolation algorithm is employed. For sensor data with relatively low sampling frequencies, Lagrange interpolation or cubic spline interpolation methods are applied at key time points, such as before and after catalyst ignition, to generate synchronized data points that precisely correspond in time to the high-frequency infrared data. As one implementation method, a Dynamic Time Warping (DTW) algorithm based on cross-correlation analysis is introduced. This algorithm can more flexibly handle the nonlinear time delays and distortions that may exist between different sensor signals, thereby achieving more robust and accurate time alignment.
[0091] Data from different sources were standardized in terms of units, ensuring that all physical quantities were expressed using the International System of Units (SI) or industry standard units. For example, temperature was standardized to Kelvin, gas flow rate to standard milliliters per minute, and pressure to kilopascals. Secondly, to guarantee data quality, the system implemented a noise data removal process using statistical methods, specifically the Local Outlier Factor (LOF) algorithm. This algorithm identified and marked significant anomalous data points caused by sensor malfunctions or strong external interference. Subsequently, for parameters that changed over time during the experiment, such as actual gas flow rate, reactor pressure, and programmed base temperature, core statistics were calculated at key reaction stages, including reactant introduction, programmed temperature rise, catalyst ignition, reaction stabilization, and cooling. These statistics included the mean, median, standard deviation, integral value, maximum value, minimum value, and slope of the linear fit. These statistics, as a condensed representation of the original time-series data, were incorporated into the standardized data records.
[0092] After time alignment and standardization, a standardized sample-experiment-performance correlation data record is generated for each catalyst sample. This record is a structured dataset that logically and clearly correlates sample information, the experimental conditions it experienced, and the catalytic performance it exhibits. Specific content includes: a unique identifier for the sample, its chemical composition, preparation method, and channel location coordinates in the screening device; a unique identifier for the experiment, experiment date, personnel information, experimental device identification, sample batch information, experimental batch information, detailed temperature program settings, target flow rates for each gas component, and total reaction pressure; and all temporal temperature characteristics extracted from infrared thermal imaging data (such as ignition time, ignition temperature, maximum temperature rise, peak temperature and its arrival time, average temperature during the reaction stabilization phase, and standard deviation of temperature fluctuations) and spatial temperature distribution characteristics (such as temperature standard deviation and hot spot area ratio within the region of interest).
[0093] Specifically, in one embodiment of the present invention, step S5 includes:
[0094] First, in terms of the three-layer logical architecture design, the database is divided into a basic metadata layer, a derived feature data layer, and a raw data index layer, with each layer having a clear functional positioning and data content.
[0095] The basic metadata layer includes: a sample information table, which records the globally unique identifier, detailed chemical composition, specific preparation method or process parameters, and precise channel position coordinates of each catalyst sample in the screening device; a preparation method table, which serves as a detailed supplement to the preparation methods in the sample information table, providing a structured description of different preparation processes, parameters, and raw material batches; an experiment configuration table, which stores detailed settings for each high-throughput experiment, such as the experiment's unique identifier, experiment execution date, operator information, unique identifier of the experimental apparatus used, sample batch information, experiment batch information, preset temperature program curve parameters, target flow rates and concentrations of each gas component, and the system's total reaction pressure setting; and a channel mapping table, which precisely defines the correspondence between physical channels and logical samples in the screening device.
[0096] The derived feature data layer stores data extracted from the raw data and processed to directly reflect key features of material properties and experimental processes. This layer includes: a catalyst performance table, which is linked to the sample information table and experimental configuration table via foreign keys, storing various catalytic performance indicators exhibited by each sample under specific experimental conditions. These indicators mainly originate from the structured temperature features extracted in step S3, such as ignition time, ignition temperature, maximum temperature rise, peak temperature and its arrival time, average temperature and temperature fluctuation standard deviation during the reaction stabilization phase, temperature standard deviation within the region of interest, and hot spot area ratio; and a time-series data summary table, which stores statistics obtained after standardizing the sensor time-series data in step S4 for key reaction stages, such as the actual average flow rate of each gas component, the actual average pressure within the reactor, and the average temperature recorded by the platform thermocouples.
[0097] The raw data index layer is responsible for managing and linking to massive amounts of raw experimental data. This layer includes: a raw data registry, which assigns a unique internal identifier to each raw data file (such as infrared thermal imaging video, sensor log files) and records its storage path (which may point to a location in a distributed object storage system), file type, size, creation time, associated experiment ID, and sample ID, among other metadata; and a keyframe table, for infrared thermal imaging video data, which records the time frame index or precise timestamp corresponding to key experimental events in the video (such as the moment of ignition, reaching peak temperature, and the start of stable reaction), as well as the offset or direct access link of these keyframes in the raw video file. This allows for the rapid location and extraction of raw image data at important time points without having to fully decode the entire video file.
[0098] In a specific example, a simplified example of a three-tier logical architecture is as follows:
[0099] First layer: Basic metadata layer
[0100] Table 1 Sample Information Table
[0101] field name Data types Constraints / Explanations Example SampleID text Primary key, unique identifier for a sample "CAT-20240512-001" Chemical Composition text Chemical composition description <![CDATA["Pt 0.5% / Al2O3"]]> PreparationMethodID text Foreign bonds, linked to the preparation method table "PM-SprayDry-003" SynthesisDate date Sample preparation date "2024-03-15" BatchNumber text Sample batch number "SN20240315-A" PhysicalForm text Sample physical morphology "powder" Supplier text raw material suppliers "Sigma-Aldrich" Notes text Other notes "High specific surface area carrier"
[0102] Table 2 Preparation Methods
[0103] field name Data types Constraints / Explanations Example PreparationMethodID text Primary key, a unique identifier for the preparation method. "PM-SprayDry-003" MethodName text Method Name "Spray drying method" Method Description text Detailed process steps description "The precursor solution is atomized and dried in a high-temperature gas flow..." KeyParameters JSON / Text Key process parameters "{'inlet_temp': '180C', 'feed_rate': '5ml / min'}"
[0104] Table 3 Experimental Configuration Table
[0105] field name Data types Constraints / Explanations Example ExperimentID text Primary key, a unique identifier for an experiment. "EXP-HTS-20240513-A-01" ExperimentDate date Experiment execution date "2024-05-13" OperatorName text Experiment Operator Name "Zhang San" DeviceID text Unique identifier of the screening device used "HTS-Device-002" Temperature Program Text / JSON Preset temperature program "Increase from 50°C to 600°C at a rate of 10°C / min" GasComposition Text / JSON Gas components and target flow rates "{'CH4': '50sccm', 'O2': '100sccm', 'N2': 'balance'}" TotalFlowRateSet numerical values Total target gas velocity (sccm) 170 PressureSet numerical values System target response pressure (kPa) 101.3 ExperimentNotes text Experimental Notes System pressure fluctuates slightly.
[0106] Table 4 Channel Mapping Table
[0107] field name Data types Constraints / Explanations Example MappingID Integer primary key 1001 ExperimentID text Foreign key, associated with the experiment configuration table "EXP-HTS-20240513-A-01" ReactorChannelNo Integer Physical channel numbering in the screening device 5 SampleID text Foreign key, linked to the sample information table "CAT-20240512-001" LoadDate date Sample loading date "2024-05-12"
[0108] Second layer: Derived feature data layer
[0109] Table 5 Catalyst Performance Table
[0110] field name Data types Constraints / Explanations Example PerformanceID Integer primary key 2001 ExperimentID text Foreign key, associated with the experiment configuration table "EXP-HTS-20240513-A-01" SampleID text Foreign key, linked to the sample information table "CAT-20240512-001" IgnitionTime numerical values Ignition time (seconds) 125.5 IgnitionTemperature numerical values Ignition temperature (°C or K) 350.2 MaxTemperatureRise numerical values Maximum temperature rise (°C or K) 150.7 Peak Temperature numerical values Peak temperature (°C or K) 500.9 TimeAtPeakTemperature numerical values Time to reach peak temperature (in seconds) 180.3 SteadyStateTemperature numerical values Average temperature during the steady-state phase of the reaction 480.5 ROITemperatureStdDev numerical values Temperature standard deviation within the region of interest 5.3 HotspotAreaRatio numerical values Hot spot area ratio 0.15
[0111] Table 6 Time Series Data Summary Table
[0112] field name Data types Constraints / Explanations Example SummaryID Integer primary key 3001 ExperimentID text Foreign key, associated with the experiment configuration table "EXP-HTS-20240513-A-01" SampleID text Foreign key, linked to the sample information table "CAT-20240512-001" Stage text Experimental phase "steady state period" ActualAvgFlowRate numerical values Actual average flow velocity (sccm) 49.8 ActualAvgPressure numerical values Actual average reaction pressure (kPa) 101.1 PlatformAvgTemperature numerical values Average temperature of thermocouples on the platform (°C) 485.2 DurationOfStage numerical values Phase duration (seconds) 600
[0113] Third layer: Raw data index layer
[0114] Table 7 Original Data Registry
[0115] field name Data types Constraints / Explanations Example RawDataID Integer Primary key, a unique identifier for original data. 4001 ExperimentID text Foreign key, associated with the experiment configuration table "EXP-HTS-20240513-A-01" SampleID text Foreign key, linked to the sample information table "CAT-20240512-001" DataType text Data types "IRVideo" FileName text Original filename "EXP-HTS-20240513-A-01_Sample001_IR.mp4" StoragePath text URI or path of data in distributed storage "s3: / / my-hts-data / videos / 2024 / 05 / ir_vid_001.mp4" FileSizeMB numerical values File size (MB) 1024.5 CreationTimestamp Timestamp File creation time "2024-05-13 15:30:00" Checksum text File verification and "a1b2c3d4e5f6..."
[0116] Table 8 Keyframe Table
[0117] field name Data types Constraints / Explanations Example KeyframeID Integer primary key 5001 RawDataID Integer Foreign keys, linked to the original data registry 4001 EventName text Key event name "IgnitionStart" EventTimestampVideo numerical values The timestamp (in seconds or frames) of the event in the video. 125.5 (seconds) CorrespondingExperimentTime numerical values The duration (in seconds) of the event over the entire experimental timeline. 305.5 VisualFeatures BLOB / Text Path to store extracted visual feature descriptors " / features / kf_5001_sift.bin" ThumbnailPath text Storage path of keyframe thumbnails " / thumbnails / kf_5001.jpg"
[0118] As shown in the eight tables above, the basic metadata layer defines "what it is" and "under what conditions"; the derived feature data layer records "how it performs"; and the original data index layer provides "where the original evidence is." This hierarchical structure facilitates data management, querying, analysis, and system maintainability.
[0119] Specifically, structured metadata and derived feature data—the content of the basic metadata layer and derived feature data layer—are stored in high-performance relational database management systems (RDBMS), such as PostgreSQL or MySQL, due to their relatively small data volume, regular structure, and frequent query and update operations. Large-volume raw multimodal experimental datasets, especially raw infrared thermal imaging video files and detailed sensor log files, are stored in lower-cost, more scalable distributed object storage systems (such as Amazon S3, MinIO, or Ceph) or network attached storage (NAS) due to their large size and access patterns that are mostly sequential or random large-block reads. The storage path information in the raw data index layer points to specific objects in these distributed storage systems. This hybrid storage strategy effectively avoids the performance bottlenecks and management complexities associated with directly storing large binary files in relational databases.
[0120] Specifically, this embodiment also includes standardized application programming interfaces, such as RESTful APIs or GraphQL APIs. These APIs encapsulate CRUD (Create, Read, Update, Delete) operations on the three-tier logical structure of the database and support complex queries, data export, and seamless integration with other external systems, such as Electronic Lab Notebooks (ELN), Laboratory Information Management Systems (LIMS), and data analysis platforms.
[0121] In addition, such as Figure 6 As shown, the present invention also provides a material database construction system based on high-throughput experimental multimodal data for implementing the method described above, characterized in that the system comprises:
[0122] The data protocol management module is used to establish standardized experimental data acquisition protocols. It includes a sample coding submodule, which is used to assign unique identifiers to samples and record the chemical composition, preparation method and channel position information of the samples; a time synchronization submodule, which is used to realize millisecond-level clock synchronization of experimental equipment; and an acquisition frequency control submodule, which is used to set the adaptive acquisition frequency parameters of the infrared thermal imager.
[0123] The multimodal data acquisition module is used to synchronously acquire and record experimental metadata, infrared thermal imaging data and sensor time-series data according to a standardized experimental data acquisition protocol, as well as to perform data integrity verification and generate a timestamped and verified original multimodal experimental dataset.
[0124] The infrared image processing module is used to process infrared thermal imaging data in the original multimodal experimental dataset. It includes an image preprocessing unit for geometric correction of infrared image sequences; a channel calibration unit for accurately aligning infrared images with standard channel templates and segmenting channel regions; and a feature calculation unit for extracting structured temperature features, including temporal temperature features and spatial temperature distribution features.
[0125] The data integration and association module is used to integrate experimental metadata and sensor time-series data from the original multimodal experimental dataset, and combine structured temperature features to generate standardized sample-experiment-performance association data records through time alignment and standardization.
[0126] The database management module is used to construct and maintain a materials database based on standardized sample-experiment-performance related data records and index information of the original multimodal experimental dataset, using a three-layer logical architecture and a hybrid physical storage strategy. The three-layer logical architecture includes a basic metadata layer, a derived feature data layer, and an original data index layer.
[0127] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for constructing a materials database based on high-throughput experimental multimodal data, characterized in that, include: S1. Establish a standardized experimental data acquisition protocol, including: assigning unique identifiers to samples using a sample coding system and recording the chemical composition, preparation method, and channel location information of the samples; adopting a multi-device time synchronization mechanism; and setting adaptive acquisition frequency parameters for the infrared thermal imager. The sample coding system includes assigning a globally unique sample identifier to each catalyst sample and recording the sample's chemical composition, preparation method, and channel position coordinates in the screening device; the multi-device time synchronization mechanism includes using a network time protocol to synchronize the infrared thermal imager, temperature sensor, mass flow controller, pressure sensor, and central control system at the millisecond level before the experiment; the adaptive acquisition frequency parameters of the infrared thermal imager include reducing the acquisition frequency during the steady reaction phase and increasing the acquisition frequency during the temperature abrupt change phase. S2. Based on a standardized experimental data acquisition protocol, synchronously collect and record experimental metadata, infrared thermal imaging data, and sensor time-series data, and perform integrity verification on the collected data to obtain a timestamped and verified original multimodal experimental dataset; the integrity verification is achieved by calculating a data health score. To quantify the initial quality of multimodal data: ; ; ; ; In the formula, , , These represent the quality scores for metadata, infrared data, and sensor data, respectively. , , These are the corresponding weight coefficients; This indicates the total number of fields in the metadata. This indicates the number of fields that have been successfully populated. It is a hyperparameter that controls the degree of penalty for missing fields; It is the quality score of the i-th field, with a value range of [0,1]. It is evaluated by verifying whether the value of this field is within the expected range, whether the format is correct, and whether it is logically consistent with other related fields. , These are the sensitivity parameter and threshold parameter for quality assessment, respectively. This represents the actual number of infrared image frames acquired. To determine the expected number of frames to be obtained based on the experiment duration and the set acquisition frequency, It is an exponential parameter that adjusts the integrity weight; It is the total number of time sampling points. This represents the key data features at time t. It is a parameter that controls the importance of continuity; , These are the sensitivity parameter and reference threshold for signal-to-noise ratio evaluation, respectively; K is the total number of sensors. It represents the number of outliers for the k-th sensor. This is the total number of data points of the sensor. The degree of punishment for controlling the proportion of outliers; It is the average absolute deviation of the correlation coefficients between all sensor pairs. , These are the reference threshold and sensitivity parameters for controlling collaborative evaluation; The difference between the sensor data distribution P and the expected theoretical distribution Q is measured. Denotes KL divergence, Control the weight of this item; S3. Process the infrared thermal imaging data in the original multimodal experimental dataset. Through image preprocessing, channel calibration and feature calculation, extract the structured temperature features of each catalyst channel. The structured temperature features include time-series temperature features and spatial temperature distribution features. The image preprocessing includes geometric correction of the infrared image sequence; the channel calibration includes aligning each frame of infrared image with the template through image registration using a pre-stored standard channel template image of the screening device, and accurately segmenting the region of interest for each channel; the feature calculation includes extracting temporal and spatial distribution features of the temperature data within the region of interest for each channel. The temporal temperature characteristics include ignition time, ignition temperature, maximum temperature rise, peak temperature and its arrival time, average temperature during the steady-state phase of the reaction, and standard deviation of temperature fluctuation; the spatial temperature distribution characteristics include the standard deviation of temperature within the region of interest and the hot spot area ratio. The ignition time in the time-series temperature characteristics is calculated using a normalized dynamic ignition criterion. To determine, when The ignition time is determined when the time continuously exceeds a preset ignition criterion threshold for at least a preset time period. Calculated using the following formula: ; in, Let be the channel average temperature at time i. This represents the average channel temperature at the previous moment. Let i be the time interval between time i and time i-1. To normalize the temperature rise rate parameter, For reference baseline temperature, For characteristic activation temperature parameters, It is an index adjustment factor; S4. Integrate experimental metadata and sensor time-series data from the original multimodal experimental dataset, and combine them with structured temperature features. Through time alignment and standardization, generate standardized sample-experiment-performance correlation data records. S5. Based on the standardized sample-experiment-performance correlation data records and the index information of the original multimodal experimental dataset, a material database is constructed using a three-layer logical architecture and a hybrid physical storage strategy. The three-layer logical architecture includes a basic metadata layer, a derived feature data layer, and a raw data index layer; the hybrid physical storage strategy includes storing structured metadata and feature data through a relational database, and storing large-capacity raw infrared thermal imaging video files and sensor log files through a distributed object storage system.
2. The method for constructing a materials database based on high-throughput experimental multimodal data according to claim 1, characterized in that, The synchronously acquired and recorded experimental metadata includes: a unique experimental identifier, experimental date, experimental personnel information, experimental device identification, sample batch information, experimental batch information, temperature program settings, target flow rates and reaction pressures of each gas component; the infrared thermal imaging data includes raw temperature matrix data recording the temperature changes in the catalyst channel; and the sensor time-series data includes platform thermocouple temperature values, gas flow rates, and pressure values.
3. The method for constructing a materials database based on high-throughput experimental multimodal data according to claim 1, characterized in that, The time alignment includes precisely aligning infrared thermal imaging data and sensor time-series data on the time axis based on high-precision timestamps; the standardization process includes converting data from different sources into a unified International System of Units (SI), eliminating sensor fault noise points, extracting statistics of time-varying parameters in key reaction stages, and forming standardized records.
4. The method for constructing a materials database based on high-throughput experimental multimodal data according to claim 1, characterized in that, The basic metadata layer includes a sample information table, a preparation method table, an experimental configuration table, and a channel mapping table; the derived feature data layer includes a catalyst performance table and a time series data summary table; the raw data index layer includes a raw data registry and a keyframe table; and a standardized application programming interface is provided to support data manipulation and integration.
5. A materials database construction system based on high-throughput experimental multimodal data, used to implement the method as described in any one of claims 1-4, characterized in that, The system includes: The data protocol management module is used to establish standardized experimental data acquisition protocols. It includes a sample coding submodule, which is used to assign unique identifiers to samples and record the chemical composition, preparation method and channel position information of the samples; a time synchronization submodule, which is used to realize millisecond-level clock synchronization of experimental equipment; and an acquisition frequency control submodule, which is used to set the adaptive acquisition frequency parameters of the infrared thermal imager. The multimodal data acquisition module is used to synchronously acquire and record experimental metadata, infrared thermal imaging data and sensor time-series data according to a standardized experimental data acquisition protocol, as well as to perform data integrity verification and generate a timestamped and verified original multimodal experimental dataset. The infrared image processing module is used to process infrared thermal imaging data in the original multimodal experimental dataset. It includes an image preprocessing unit for geometric correction of infrared image sequences; a channel calibration unit for accurately aligning infrared images with standard channel templates and segmenting channel regions; and a feature calculation unit for extracting structured temperature features, including temporal temperature features and spatial temperature distribution features. The data integration and association module is used to integrate experimental metadata and sensor time-series data from the original multimodal experimental dataset, and combine structured temperature features to generate standardized sample-experiment-performance association data records through time alignment and standardization. The database management module is used to construct and maintain a materials database based on standardized sample-experiment-performance related data records and index information of the original multimodal experimental dataset, using a three-layer logical architecture and a hybrid physical storage strategy. The three-layer logical architecture includes a basic metadata layer, a derived feature data layer, and an original data index layer.