Database construction and data standardization method based on MOF proton conductor material

Through web crawlers, distributed databases and blockchain technology, the problems of scattered and non-standard data on MOF proton conductor materials have been solved, the systematic integration and security management of data have been achieved, the comparability and reliability of data have been improved, and the development of materials science has been promoted.

CN120636627AInactive Publication Date: 2025-09-12XI'AN POLYTECHNIC UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510758624.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The data on MOF proton conductor materials are scattered, non-standardized and difficult to effectively manage and utilize, making it difficult to compare and analyze data on the same type of materials. Existing databases cannot cover all important information, and there is a lack of a unified data storage and sharing platform, making it difficult to guarantee data quality and authenticity.

Method used

Use web crawler technology to collect multi-channel data, design multi-level data structures, build distributed database systems, use blockchain technology to ensure data security and consistency, implement data standardization methods, build a multi-level data verification system, and monitor and update data in real time.

Benefits of technology

It has achieved systematic integration and standardized management of MOF proton conductor material data, improved the comparability and reliability of data, ensured the security and timeliness of data, promoted data sharing and cooperation among scientific researchers, and promoted the development of materials science.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636627A_ABST
    Figure CN120636627A_ABST
Patent Text Reader

Abstract

The invention discloses a database construction and data standardization method based on an MOF proton conductor material, and belongs to the crossing field of material science and database technology. Aiming at the problems of data dispersion and non-uniform formats of the current MOF proton conductor material, data is collected through multiple channels, and a web crawler is innovatively adopted to improve the efficiency. In the aspect of database construction, a distributed database is selected and combined with a block chain technology. In the aspect of data standardization, multi-dimensional data formats are standardized, complex representation data are processed through artificial intelligence, quality control is enhanced, data reasonability is predicted through machine learning, and real-time updating and correction are achieved through big data. According to the method, dispersed data are integrated, consistency and comparability are ensured, data quality and intelligent level are improved, and powerful support is provided for related research and application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of material science, database technology and information processing technology, and in particular to a database construction and data standardization method based on MOF proton conductor materials. Background Art

[0002] In the current development of materials science, MOF (Metal-Organic Framework) proton conductor materials have shown great application potential in many fields due to their unique physical and chemical properties, especially in energy storage and conversion, sensors, and catalysis, which have attracted widespread attention. However, with the continuous deepening of research on MOF proton conductor materials, the related data has exploded, and these data are scattered among various sources. On the one hand, a large number of scientific research teams have published their research results in various academic journals, conference proceedings and literature. Different teams have significant differences in experimental conditions, material performance testing methods and data recording methods, making it difficult to directly compare and analyze data for the same type of MOF proton conductor materials. For the measurement of proton conductivity rate, some teams conducted tests in high humidity environments, while others operated in low humidity environments. The measuring instruments, test temperature ranges and sample preparation methods were all different, resulting in a lack of comparability in the reported proton conductivity rate data.

[0003] On the other hand, while specialized materials databases like the Cambridge Crystal Structure Database (CCDC) store some crystal structure information, they are not comprehensive databases specifically for MOF proton conductors and cannot cover all the important information about these materials, including the various experimental conditions during synthesis and performance data under different application scenarios. Furthermore, the lack of a unified data storage and sharing platform across different laboratories makes it difficult to efficiently integrate and share research results. This not only wastes data resources but also significantly hinders the further research and development of MOF proton conductors.

[0004] Furthermore, existing databases and data management methods cannot guarantee data quality and authenticity, and lack effective mechanisms for data updating and maintenance. Regarding data formats, there is no unified standard for material names, chemical composition, performance data, and crystal structure information, resulting in a low level of data standardization and further exacerbating the data chaos. These issues urgently require innovative solutions to systematically integrate and standardize dispersed data, enabling efficient data management and sharing, thereby advancing research in the field of MOF proton conductor materials to a new level.

[0005] Therefore, the present invention is committed to developing a database construction and data standardization method based on MOF proton conductor materials to overcome the above-mentioned problems of data dispersion, non-standardization, and difficulty in effective management and utilization, and to provide a more complete, convenient and reliable data platform for scientific researchers and related practitioners in this field, thereby promoting the research, development and application of MOF proton conductor materials.

[0006] In the current development of materials science, MOF (Metal-Organic Framework) proton conductors, due to their unique physical and chemical properties, have shown great potential for application in numerous fields, particularly in energy storage and conversion, sensors, and catalysis. However, as research on MOF proton conductors continues to deepen, the amount of data related to them has exploded. This data is scattered across various sources, leading to a series of urgent challenges.

[0007] On the one hand, numerous research teams publish their findings in various academic journals and conference proceedings. However, significant differences exist between different teams regarding experimental conditions, material performance testing methods, and data recording methods, making it difficult to directly compare and analyze data from the same type of MOF proton conductor material. For example, when measuring proton conductivity, some teams conduct tests in high-humidity environments, while others operate in low-humidity environments. Furthermore, the measurement instruments, test temperature ranges, and sample preparation methods vary, resulting in a lack of comparability in the reported proton conductivity data. Summary of the Invention

[0008] This invention discloses a method for building a database and standardizing data based on MOF proton conductor materials. This method belongs to the fields of materials science, database technology, and information processing technology, and aims to address the issues of MOF proton conductor material data being dispersed, non-standardized, and difficult to effectively manage and utilize. The specific contents are as follows:

[0009] The database construction steps are as follows

[0010] Multi-channel data collection: Using web crawler technology, we set "MOF-based proton conductor", "MOF-based proton-conducting material", "MOF-based proton conductive materials", "proton-conducting MOF", "proton-conductive MOF", and "proton conduction in MOF" as keywords, and automatically searched multiple mainstream academic databases such as Web of Science, ScienceDirect, ACS Publications, Wiley, RSC, and Elsevier. The crawler program has intelligent self-adaptation capabilities and can flexibly adjust the preset extraction rules according to the unique page structure and document format of different databases. Key data was accurately extracted from the title, abstract, text, figures, and supporting information of the literature, including the material name or chemical formula of the MOF proton conductor material, the number of bound water and crystal water, and the synthesis method (detailed records of the solvents and reactants used, the reaction temperature was controlled within an error range of ±1°C, and the reaction time was accurate to ±1 hour). However, data such as chemical composition and performance test results (such as proton conduction rate, activation energy, water adsorption, water contact angle, and acidity at different temperatures and humidities) will be screened out, and the extracted data will be temporarily stored in a temporary database. In addition, two professional crystal structure databases, the Cambridge Crystal Structure Database (CCDC) and the Crystallographic Open Database (COD), were manually searched in depth to carefully screen out crystal structure data related to MOF proton conductor materials. After systematic organization and necessary supplementation, the data was entered into the database. In a laboratory environment, the synthesis of MOF proton conductor materials was carried out in strict accordance with internationally accepted standard experimental procedures. High-precision instruments and equipment were used to accurately measure the various properties of the materials. Taking the measurement of proton conduction rate by AC impedance spectroscopy as an example, the test temperature was strictly controlled within an error range of ±0.5°C, and the measurement humidity was controlled within an error range of ±2% relative humidity. At the same time, every operation step in the experimental process, detailed information on the experimental equipment used, information on the experimenters, and measurement data were recorded in detail to ensure the integrity and reliability of the experimental data in all aspects.

[0011] Design data hierarchy: Carefully design a data hierarchy containing multiple key information layers. The basic information layer of the material stores the unique identifier of the material, the name of the material that follows the common naming rules in the industry, and the chemical composition formula written strictly in accordance with the regulations of the International Union of Pure and Applied Chemistry (IUPAC), providing accurate and standardized basic identification for the material to facilitate subsequent data association and query. The synthesis information layer records the specific synthesis route of each material in detail, accurate to the name, purity (≥99%), dosage of the chemical reagents used, and precise reaction conditions (reaction temperature accurate to ±1°C, reaction time accurate to ±1 hour), fully presenting the synthesis process of the material. The structural information layer specifically stores crystal structure data, in which the average hydrogen bond length is accurate to The porosity is accurate to 0.1%, and the unit cell parameters are stored strictly in the standard crystallographic format. The space group uses the symbols specified by the International Union of Crystallography (IUCr), providing a solid and reliable basis for subsequent material structure analysis and performance correlation. The performance information layer comprehensively stores proton conductivity, activation energy, water adsorption, water contact angle, and acidity performance data at different temperatures (range set from 25°C to 80°C, intervals of 5°C) and humidity (range from 0% to 100% relative humidity, selected based on common saturated salt solutions). It also records the application performance parameters of the material in different practical application scenarios (such as various types of fuel cells such as proton exchange membrane fuel cells and direct methanol fuel cells, as well as different sensors such as gas sensors and humidity sensors), including output power density, response time, sensitivity, and stability data. Each layer of data is closely linked through an optimized association algorithm, enabling rapid association from basic material information to synthesis information, structural information, and performance information, and vice versa. The material's unique identifier allows for easy retrieval of its performance in different application scenarios, and its synthesis and structural information can be accurately traced based on the performance data, providing users with comprehensive and systematic data association and query capabilities, facilitating in-depth understanding and research of materials from all aspects.

[0012] Database System Construction: Distributed database management systems such as Apache Cassandra and MongoDB were selected to build a powerful distributed database cluster. This system utilizes distributed storage and parallel processing technologies to efficiently handle the storage and processing needs of massive amounts of data. It supports horizontal and vertical data scalability, automatically and intelligently adjusting storage and computing resources based on dynamic data growth, ensuring system performance consistently meets established performance targets and guaranteeing data query response times of less than one second, significantly improving data storage and access efficiency. A consortium chain model based on blockchain technology was adopted, rationally assigning different permissions to different types of participating nodes. Nodes belonging to research institutions were granted data entry and modification permissions to ensure timely updates of specialized data; ordinary user nodes only had data query permissions to ensure data security. Blockchain hashing algorithms and the PBFT consensus mechanism ensured data consistency and immutability. During data entry, modification, and update operations, the system automatically generated immutable hash values, providing strong security for trusted data storage and use. A user-friendly database management interface was developed, integrating a variety of practical functions to support data entry, query, modification, deletion, and export. At the same time, it provides powerful data visualization functions, which can display the trend of material properties changing with temperature and humidity conditions in the form of intuitive charts, making it convenient for users to conduct in-depth analysis and comparison of data.

[0013] The data standardization method comprises the following steps:

[0014] Standardize data format: For numerical data, establish unified unit and precision standards. The proton conduction rate is uniformly expressed in S / cm, and is accurately retained to three decimal places according to the accuracy of the measuring instrument; the activation energy is in eV, also retained to three decimal places; the temperature is in °C, and the humidity is expressed in % relative humidity, and is accurate to the integer digit to ensure the consistency and comparability of the data in terms of values. For text data, the material names strictly follow the naming rules commonly used in the industry, and the chemical composition formulas are written in accordance with the writing standards of chemical element symbols and chemical bonds. The use of ambiguous and non-standard abbreviations is eliminated to ensure that the text information is clear and standardized, so that different users can accurately understand and use the data. In terms of structural data, crystal structure data is stored in the standard format of crystallography, in which the average hydrogen bond length is accurate to Porosity is accurate to 0.1%. Advanced artificial intelligence algorithms are used to pre-process complex material characterization data, such as X-ray photoelectron spectroscopy (XPS), infrared spectroscopy (IR), and thermogravimetric analysis (TGA) data. This algorithm automatically identifies the raw data formats generated by different instruments and equipment, intelligently selects the corresponding processing algorithm based on the characteristics of the data, accurately extracts key information such as the position, intensity, and half-peak width of characteristic peaks, and converts this information into a unified standard format for storage. The extraction error of characteristic peaks is strictly controlled to no more than ±5%, allowing direct comparison and analysis of complex characterization data measured by different sources and different instruments.

[0015] Strengthen data quality control: Build a multi-level, comprehensive data verification system. Before data entry, conduct strict format and range checks on the data. For numerical data, check whether it is within a reasonable range based on physical and chemical principles; for text data, carefully check whether it follows the naming and writing specifications; for structural data, comprehensively check whether it is complete and accurate. Utilize machine learning models to conduct in-depth training on a large number of known MOF proton conductor material data sets to make reasonable predictions on newly entered data. Once data with significant differences in characteristics from the training data is found, it will be immediately marked as suspicious data. For proton conduction rate data, its reasonable range will be predicted based on factors such as material type, structure, and synthesis conditions. Data outside this range will be subject to focused review, and a detailed error report will be automatically generated. The report content covers the error type, possible causes, and targeted modification suggestions to ensure the reliability and accuracy of the data in all aspects. Establish a scientific data quality assessment mechanism and run the data quality assessment algorithm regularly. The data in the database is comprehensively scored on a scale of 0 to 100 based on multiple indicators, including completeness (whether it contains complete information about the material), accuracy (whether the data meets measurement accuracy and range requirements), consistency (whether data from different sources corroborate each other), and comparison results with data from similar materials. The scores are then categorized into three levels: high quality (80-100 points), medium quality (40-79 points), and low quality (0-39 points). Low-quality data is specially marked and subject to focused review. Targeted improvement plans to enhance data quality are also automatically generated, enabling continuous optimization and improvement of the data.

[0016] Data Updates and Corrections: Leveraging big data analysis tools, academic databases and professional journal websites are monitored in real time. The update monitoring frequency can be flexibly adjusted based on user needs, with a minimum interval of one hour. The tool features powerful semantic analysis capabilities, accurately identifying key information related to MOF proton conductor materials. When new research findings involving data in the database are detected, the system automatically extracts relevant information and generates an update request. Intelligent algorithms analyze the differences between new and existing data, pinpointing the data requiring revision and generating detailed, well-reasoned correction recommendations based on the importance and degree of difference. For updates to key performance data, verification results from at least three replicate experiments are required to ensure data reliability. Correction recommendations include a detailed description of the changes, the rationale for the changes, and an assessment of the overall impact on the data. Update requests and correction recommendations are submitted to experts in the relevant fields for rigorous review, who conduct a comprehensive review based on the reliability of the data source, the scientific nature of the experimental methods, and the rationale for the results. Upon approval, the database is promptly updated, with detailed records of the update time, content, and reviewer information, forming a complete data update history. After data is updated, the system automatically reassesses its quality to ensure that the updated data meets quality requirements, ensuring continuous database optimization and data quality stability. The database construction and data standardization methods of the present invention enable the systematic integration, standardized management, and efficient utilization of MOF proton conductor material data, providing strong data support for research and development in this field and significantly promoting the further development of materials science and related application fields.

[0017] The above technical solution can bring the following technical effects:

[0018] 1. Data integration and comprehensiveness improvement: Through multi-channel data collection methods, including web crawlers, manual retrieval of professional databases and laboratory's own experimental data accumulation, the present invention can systematically integrate the MOF proton conductor material data originally scattered in various academic literature, different laboratories and professional databases. This integration covers everything from basic information, synthesis information, structural information, performance information to application information of the material, providing a comprehensive data resource library for scientific researchers. Scientific researchers no longer need to search for the required information in many different documents and databases, but can easily obtain the proton conduction rate and activation energy data of MOF proton conductor materials under different temperature and humidity conditions in the database of the present invention, and can directly link to the material's synthesis route, crystal structure and performance in fuel cell and sensor applications, greatly improving the efficiency and convenience of research. This helps to avoid duplication of research due to incomplete and scattered information, saves scientific researchers a lot of time and energy, and promotes data sharing and communication between different research teams.

[0019] 2. Data standardization and enhanced comparability: The data standardization method of the present invention ensures that data from different sources have a unified format and specification. For numerical data, the units and precision of proton conduction rate and activation energy are clearly defined; for text data, universal naming rules and chemical writing specifications are adopted; for complex structural and characterization data, artificial intelligence algorithms are used for preprocessing and feature extraction, and converted into a standard format. This standardization makes the data generated by different research teams and different experiments comparable. In the past, the proton conduction rate test results of different laboratories for the same type of MOF proton conductor material may not be directly compared due to differences in test environment, measuring instruments and data recording methods. However, in the database of the present invention, all data follow a unified standard, allowing scientific researchers to accurately compare the performance advantages and disadvantages of different materials, analyze the relationship between material performance and structure and synthesis conditions, and then more accurately evaluate the performance potential of the material, providing a reliable basis for material screening and optimization. At the same time, the standardized data format also facilitates further analysis and mining of the data, providing a high-quality data foundation for subsequent research and development.

[0020] 3. Data Quality and Security: This invention establishes a multi-level data verification system and data quality assessment mechanism. Using machine learning models, it predicts the rationality of newly entered data, assigns quality scores to the data, and manages the data at different levels, ensuring high data quality in the database. For suspicious or low-quality data, the system automatically generates error reports and improvement plans, helping researchers promptly identify and resolve data issues. Furthermore, by adopting a consortium chain model based on blockchain technology, different permissions are assigned to different nodes, ensuring data consistency and immutability. During data entry, modification, and update processes, hashing algorithms and consensus mechanisms ensure data security and prevent malicious tampering. During collaboration among research institutions, all participants can confidently contribute data to the database without worrying about data authenticity and security. This provides reliable guarantees for cross-institutional research collaboration and data sharing, promoting collaborative innovation in the field of MOF proton conductor materials. Furthermore, real-time monitoring by big data analysis tools and a data update and correction mechanism powered by intelligent algorithms ensure that the data in the database is always up-to-date, ensuring that researchers use the most cutting-edge and accurate data, contributing to the reliability and innovation of research results. The above technical effects have jointly promoted the development of the MOF proton conductor material field, making research and development in this field more efficient, scientific and orderly, and providing more solid data support and guarantee for the realization of related applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a multi-channel data collection step diagram of the present invention

[0022] Figure 2 This is the database component step diagram of the present invention

[0023] Figure 3 This is a diagram of the steps for standardizing the multi-dimensional data format of the present invention.

[0024] Figure 4 This is a step diagram of the database construction and data standardization method of the present invention DETAILED DESCRIPTION

[0025] The following is a combination of the embodiments of the present invention Figures 1 to 4 This document clearly and completely describes the technical solutions in the embodiments of the database construction and data standardization method of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.

[0026] Example 1: Data Collection and Integration Advantages

[0027] Background: A research team planned to conduct research on the application of novel MOF proton conductor materials in fuel cells. They needed to collect a large amount of data on the synthesis, structure, and properties of MOF proton conductor materials. However, facing numerous academic databases, traditional manual data collection was not only time-consuming and labor-intensive, but also prone to missing key information, seriously hindering research progress.

[0028] Implementation process:

[0029] 1. Web crawler search: Using web crawler technology, we set "MOF-based proton conductor," "MOF-based proton-conducting material," "MOF-based proton conductive materials," "proton-conducting MOF," "proton-conductive MOF," and "proton conduction in MOF" as keywords to automatically search mainstream academic databases such as Web of Science, ScienceDirect, ACS Publications, Wiley, RSC, and Elsevier. The crawler program adaptively adjusts the extraction rules based on the page structure and document format of different databases. For example, in the ScienceDirect database, the extraction algorithm is optimized based on the HTML structure of the document, extracting data from the document title, abstract, text, figures, and supporting information. After three days of searching, data from 500 relevant documents were obtained.

[0030] 2. Data extraction and storage: According to preset rules, data such as the material name or chemical formula of the MOF proton conductor material, the number of bound water and crystal water, and the synthesis method (including the solvent used, reactants, reaction temperature accurate to ±1°C, and reaction time accurate to ±1 hour) were extracted from the literature. After removing data such as chemical composition and performance test results, the extracted data was stored in a temporary database. Of the 500 documents, valid data was successfully extracted from 420, covering 300 different MOF proton conductor materials.

[0031] 3. Professional Database Search: Professionals were assigned to manually search the Cambridge Crystal Structure Database (CCDC) and the Crystallographic Open Database (COD). After two days of screening and sorting, 80 crystal structure data items related to MOF proton conductor materials were obtained from the CCDC and 60 from the COD. These data were then added to the temporary database.

[0032] Results: Using traditional manual data collection methods, it would have taken at least four weeks to complete the same amount of data. However, using the method presented in this paper, data collection was completed in just five days, significantly saving time and labor costs. Comprehensive web crawler searches and manual supplementation of specialized databases ensured data integrity. Compared to manual data collection, which may have missed some of the literature, the data obtained by this method covers a wider range of MOF proton conductor materials and related information.

[0033] Example 2: Data hierarchy design advantages

[0034] Background: A research institute specializing in the development of proton conductors planned to develop a novel proton exchange membrane based on MOF proton conductor materials for fuel cell applications. However, due to a lack of systematic understanding of the relationship between material properties, synthesis, and structure, research and development progress was slow and the direction was unclear.

[0035] Implementation process:

[0036] 1. Constructing a data hierarchy: Building a database based on the data hierarchy designed by the present invention. The basic material information layer stores the unique identifier, name and chemical composition formula of the material written in accordance with the provisions of the International Union of Pure and Applied Chemistry (IUPAC). For the material MOF-801, its name and chemical formula [Zr6O4(OH)4(fumarate)6] are accurately recorded. The synthesis information layer records the synthesis route in detail, such as the name, purity (≥99%), dosage of the chemical reagents used, and the reaction conditions (reaction temperature accurate to ±1°C, reaction time accurate to ±1 hour). When synthesizing MOF-801, the fumaric acid and ZrOCl2·8H2O used are of high purity, respectively analytical grade (generally ≥99%) and above 99.5%. The dosage is determined according to the scale of the synthesis (e.g., when preparing microcrystals, 1 mmol of fumaric acid and 1 mmol of ZrOCl2·8H2O), the reaction temperature is 130°C, and the reaction time is 6 hours. The structural information layer stores crystal structure data, average hydrogen bond length (accurate to ), porosity (accurate to 0.1%), etc. The performance information layer stores performance data such as proton conductivity, activation energy, water adsorption, water contact angle, acidity, etc. at different temperatures (289K to 334K, intervals of 5K) and humidity (0% to 100% relative humidity, selected based on common saturated salt solutions), as well as the application performance parameters of the material in fuel cells (open circuit voltage, output current, power density data).

[0037] 2. Data Association and Query: Through optimized association algorithms, each layer of data is tightly linked. When researching a humidity sensor, researchers at a research institute can input the unique identifier [Zr6O4(OH)4(fumarate)6] and quickly link it to its synthesis, structure, and performance information. For example, the performance information reveals that at a temperature of 298K and a relative humidity of 98%, the material's proton conductivity is 1.8×10-3 S / cm. This information can also be traced back to its synthesis method and crystal structure.

[0038] Results: R&D personnel can fully understand the various information of the materials and their interrelationships, providing a clear direction for research and development. In the past research and development process, due to the dispersion of data, it may take several days to find relevant information. Using the database of the present invention, it only takes a few minutes to obtain the required comprehensive information, avoiding blind attempts and shortening the research and development cycle from the originally expected 6 months to 4 months. Tracing the synthesis and structural information based on material performance data helps research institutes optimize the synthesis process. For example, by analyzing the proton conductivity performance of materials under different synthesis conditions, research institutes adjusted the synthesis temperature, time, reactant and additive ratio of [Zr6O4(OH)4(fumarate)6]. The synthesized MOF gel material achieved efficient proton conduction under anhydrous conditions in a wide temperature range (such as -50°C to 200°C), solving the problem that the proton conduction rate of traditional MOF materials in this environment is extremely low (usually <10 -6 S / cm) problem, the conduction rate can be increased to 10 -3 -10 -2 The S / cm level provides strong support for the R&D decisions of scientific research institutes.

[0039] Example 3: Advantages of a distributed database system

[0040] Background: A large research institution has accumulated a vast amount of data from its long-term research on MOF proton conductor materials. Traditional centralized databases experienced severe performance bottlenecks when storing and processing this data, resulting in data query response times of several minutes, making them unable to meet the needs of researchers.

[0041] Implementation process:

[0042] 1. Build a distributed database cluster: We selected Apache Cassandra and MongoDB distributed database management systems to build a distributed database cluster. Data was stored across 10 nodes, each equipped with 16GB of memory and a 1TB hard drive, enabling distributed data storage and parallel processing.

[0043] 2. Data migration and optimization: Migrate existing data to a distributed database and optimize it based on the characteristics of the distributed system. For example, partition data and store related data on adjacent nodes to improve data access efficiency.

[0044] 3. Performance Testing and Adjustment: After the data migration was complete, performance testing was conducted. By simulating a large number of query requests, we found that the average data query response time was 0.5 seconds, significantly less than the previous several minutes. Furthermore, as the data volume continued to grow, the distributed system was able to automatically adjust storage and computing resources. When the data volume increased by 50%, the system automatically added two nodes to ensure that the data query response time remained below 1 second.

[0045] Results: The distributed system's parallel processing capabilities significantly improved data storage and access efficiency. Researchers can now quickly access the data they need, significantly improving their work efficiency. For example, a large-scale material properties data query that previously took five minutes now takes only 0.5 seconds, a 600-fold improvement. The system supports both horizontal and vertical scalability, easily accommodating continued data growth in the future. The research institute projects a threefold increase in data volume over the next five years, but the distributed database system can maintain system performance by adding nodes and upgrading hardware.

[0046] Example 4: Blockchain technology ensures data security advantages

[0047] Background: In a multi-institutional collaborative research project, five research institutions and three universities shared data on MOF proton conductor materials. Because the data involved core research results and commercial secrets from each institution, ensuring data security and immutability became key issues for the collaboration.

[0048] Implementation process:

[0049] 1. Constructing a consortium chain model: Using blockchain technology, the consortium chain model assigns different permissions to different participating nodes. Nodes belonging to scientific research institutions have data entry and modification permissions, while ordinary user nodes (such as non-core researchers in enterprises) only have data query permissions.

[0050] 2. Data security: The blockchain's hashing algorithm and PBFT consensus mechanism ensure data consistency and immutability. The system automatically generates an immutable hash value when data is entered, modified, or updated. For example, when a scientific research institution enters new material performance data, the system generates a unique hash value for the data and records it on the blockchain. Any subsequent attempt to tamper with the data will result in a change in the hash value, which the system will immediately detect and reject.

[0051] 3. Permission Management and Auditing: Establish a strict permission management and auditing mechanism. Each node's operation will be recorded on the blockchain for easy audit and traceability. For example, when a scientific research institution node modifies data, it must provide detailed reasons for the modification and relevant experimental evidence, and the operation will be reviewed by other nodes.

[0052] Results: The blockchain's hashing algorithm and consensus mechanism effectively prevented malicious data tampering, ensuring data security and credibility. Throughout the collaborative project, no data tampering occurred, ensuring the security of each institution's core data. A rational permissions allocation mechanism ensured that different nodes could only perform operations consistent with their permissions. Regular enterprise users could only query publicly available data, protecting the privacy and confidentiality of research institutions. Furthermore, a rigorous audit mechanism promoted standardized data use and enhanced trust in cross-institutional collaboration.

[0053] Example 5: Advantages of Data Standardization

[0054] Background: Different research teams use different formats and units when recording data on MOF proton conductors, making it difficult to directly compare and analyze the data. For example, some teams use mS / cm to express proton conductivity, while others use S / cm; and when recording temperature, some use thermodynamic temperature, while others use degrees Celsius. This seriously hinders the communication and application of scientific research results.

[0055] Implementation Process

[0056] 1. Standardization of numerical data: Standardize the units and precision of numerical data. Proton conductivity data is standardized in S / cm and rounded to three decimal places based on the precision of the measuring instrument. For example, convert proton conductivity data originally expressed in mS / cm, such as 200 mS / cm, to 0.200 S / cm (rounded to three decimal places). Activation energy is expressed in eV and rounded to three decimal places; temperature is expressed in °C; and humidity is expressed in % relative humidity, rounded to the nearest integer.

[0057] 2. Text data standardization: Text data adheres to industry-standard naming conventions and chemical formula writing standards. For material names, uniformly use internationally recognized naming conventions and avoid ambiguous and non-standard abbreviations. Chemical formulas are written according to standard notation for chemical element symbols and chemical bonds. For example, the non-standard abbreviation material name "MOF-801" was corrected to the correct [Zr6O4(OH)4(fumarate)6].

[0058] 3. Standardization of structural data: Crystal structure data are stored in the standard format of crystallography, and the average hydrogen bond length is accurate to The porosity is accurate to 0.1%. Crystal structure data provided by different research teams are converted in format and adjusted in accuracy to ensure data consistency.

[0059] Results: Standardized data allows for direct comparison and analysis, facilitating communication and collaboration between different research teams. For example, at an international academic exchange conference, teams were able to quickly and accurately compare the performance data of different MOF proton conductor materials, promoting the sharing and application of scientific research results. Standardized data formats and writing requirements reduce data ambiguity and errors. Through data standardization, approximately 15% of errors and irregularities in the original data were identified and corrected, improving data quality and usability.

[0060] Example 6: Data Update and Correction Advantages

[0061] Background: During their ongoing research on MOF proton conductors, a research team discovered that some data in their previous database had become outdated due to new research findings. For example, the proton conductivity rates of certain materials under new experimental conditions differed significantly from those in the database. However, traditional methods struggled to quickly and accurately locate and correct these data.

[0062] Implementation process:

[0063] 1. Real-time Monitoring and Information Extraction: Using big data analysis tools, we monitor updates to academic databases and professional journal websites in real time. We set the update monitoring interval to one hour and used semantic analysis to accurately identify key information related to MOF proton conductor materials. During a week of monitoring, we discovered 20 new research findings involving data in the database.

[0064] 2. Discrepancy Analysis and Correction Suggestion Generation: When new research results are discovered that involve data in the database, the system automatically extracts relevant information and generates an update request. An intelligent algorithm analyzes the differences between the new and old data, locates the data that requires modification, and generates detailed correction suggestions based on the importance and degree of difference. For example, for the proton conductivity data of a certain material at 25°C and 98% relative humidity, the new research result is 2.5×10⁻³ S / cm, while the database data is 1.8×10⁻³ S / cm. After analysis, the intelligent algorithm recommends updating the data to the new measured value and provides detailed experimental methods and data sources as a basis for modification.

[0065] 3. Expert Review and Data Update: Update requests and correction suggestions are submitted to three experts in the relevant fields for review. The experts conduct a comprehensive review based on the reliability of the data source, the scientific nature of the experimental methods, and the rationality of the results. During the review, the experts requested further verification of five data points. After the research team supplemented the verification results with at least three replicate experiments, the review was approved. The database is updated and detailed records of the update time, content, and reviewer information are recorded to form a complete data update history. After the update, the system automatically re-evaluates the data quality to ensure that the updated data meets quality requirements.

[0066] Results: Outdated data in the database can be promptly identified and updated, ensuring its timeliness and accuracy. Through real-time monitoring, the research team can update the database shortly after new research results are published, ensuring researchers have access to the latest material data. A combination of intelligent algorithms and expert review ensures the scientific and rational nature of data updates. During the update process, a strict review mechanism and repeated experimental verification requirements enhance the reliability and practicality of the database and avoid research bias caused by erroneous data.

[0067] This example demonstrates the security advantages of blockchain technology within the present invention. By assigning permissions to different nodes and utilizing hashing algorithms and consensus mechanisms, the security and immutability of data during storage and use are ensured, preventing malicious tampering and misuse. This ensures the reliability of key information in the database, provides a reliable foundation of trust for cross-institutional scientific research collaboration and data sharing, and promotes cooperation and information exchange among research teams.

Claims

1. A database construction and data standardization method based on MOF proton conductor materials, characterized in that: The following steps are involved: S101 Multi-channel data collection: Using web crawler technology, the keywords "MOF-based protonconductor", "MOF-based proton-conducting material", "MOF-based proton conductive materials", "proton-conducting MOF", "proton-conductive MOF", and "proton conduction inMOF" were set to automatically search multiple mainstream academic databases including Web of Science, ScienceDirect, ACS Publications, Wiley, RSC, and Elsevier. According to preset extraction rules, the crawler program extracted the material name or chemical formula, the number of bound water and crystalline water, the synthesis method (including the solvent used, reactants, reaction temperature, reaction time, and reaction pressure), the material name or chemical formula, the number of bound water and crystalline water of MOF proton conductor materials from the title, abstract, text, figures, and supporting information of the literature. The chemical composition and performance test results (including proton conduction rate, activation energy, water adsorption, water contact angle, and acidity at different temperatures and humidities) were removed and stored in a temporary database. s102. Manually search the Cambridge Crystal Structure Database (CCDC) and Crystallographic Open Database (COD) professional crystal structure databases to screen out crystal structure data related to MOF proton conductor materials, and organize and supplement them. s103. In the laboratory, MOF proton conductor materials are synthesized according to internationally accepted standard experimental procedures. High-precision instruments and equipment are used to accurately measure the various properties of the materials. The experimental process and results are recorded in detail. When measuring the proton conduction rate using AC impedance spectroscopy, the test temperature is precisely controlled within an error range of ±1°C and the measured humidity is within an error range of ±2% relative humidity. The corresponding experimental conditions and data are recorded. s201. Design a data hierarchy, including a basic material information layer, which stores the material's unique identifier, name, and chemical composition formula, ensuring that the chemical composition formula is written in accordance with the International Union of Pure and Applied Chemistry (IUPAC) regulations; a synthesis information layer, which records in detail the specific synthesis route of each material, including the name, purity, and dosage of the specific chemical reagents used, and the reaction conditions (reaction temperature accurate to ±1°C, reaction time accurate to ±1 hour); a structure information layer, which stores crystal structure data, including the average hydrogen bond length accurate to The porosity is accurate to 0.1%; the performance information layer stores proton conductivity, activation energy, water adsorption, water contact angle, and acidity performance data at different temperatures (ranging from 25°C to 80°C, with an interval of 5°C) and humidity (ranging from 0% to 100% relative humidity, selected according to common saturated salt solutions); and records the application performance parameters of the material in different actual application scenarios (in different types of fuel cells, including proton exchange membrane fuel cells, direct methanol fuel cells, different sensors, gas sensors, and humidity sensors), including output power, response time, sensitivity, and stability data. s202. Data at each level is tightly linked through optimized association algorithms, enabling rapid linkage from basic material information to synthesis, structure, and performance information, and vice versa. A material's unique identifier can be used to retrieve its performance in different application scenarios, while performance data can be used to trace its synthesis and structure information. s301. Select distributed database management systems, Apache Cassandra and MongoDB, to build a distributed database cluster and store data on multiple nodes to achieve distributed storage and parallel processing of data, thereby improving the efficiency of data storage and access. s302. Using a consortium chain model based on blockchain technology, different permissions are assigned to different participating nodes. Nodes belonging to research institutions have data entry and modification permissions, while ordinary user nodes only have data query permissions. The blockchain's hash algorithm and consensus mechanism (PBFT algorithm) ensure data consistency and immutability. An immutable hash value is generated during data entry, modification, and update, guaranteeing data security. s303. Develop a user-friendly database management interface that supports data entry, query, modification, deletion, and export operations, while providing data visualization capabilities to display the trends of material properties changing with temperature and humidity in the form of charts. s401. For numerical data, the proton conductivity rate is uniformly expressed in S / cm and is rounded to three decimal places according to the accuracy of the measuring instrument. The activation energy is in eV and is rounded to three decimal places. The temperature is in °C and the humidity is expressed in % relative humidity, accurate to the integer. s402. For text data, material names should follow the commonly used naming rules in the industry, and chemical composition formulas should follow the writing standards for chemical element symbols and chemical bonds, avoiding the use of vague and non-standard abbreviations. s403. For structural data, crystal structure data are stored in the standard format of crystal structure, where the average hydrogen bond length is accurate to The porosity is accurate to 0.1%. s501. Build a multi-level data verification system. Before data entry, first check whether the data format complies with regulations, whether the numerical data is within a reasonable range, whether the text data follows the naming and writing standards, and whether the structural data is complete and accurate. s502. Utilize machine learning models to train on known MOF proton conductor material data sets to predict the rationality of newly entered data. Data with significant differences in characteristics from the training data will be marked as suspicious. For proton conduction rate data, its reasonable range will be predicted based on material type, structure, and synthesis conditions. Data outside the range will be subject to key review. s503. Establish a data quality assessment mechanism, regularly run the data quality assessment algorithm, score the data in the database according to the data completeness, accuracy, and consistency indicators, and classify the data into high-quality, medium-quality, and low-quality data based on the score, and store and manage them separately, with low-quality data being reviewed and processed in a focused manner. s601. Using big data analysis tools, monitor updates to academic databases and professional journal websites in real time. Through text mining and data comparison techniques, promptly discover new research results related to MOF proton conductor materials stored in the database. s602. When new research results are found involving data in the database, relevant information is automatically extracted and an update request is generated. Intelligent algorithms are used to analyze the differences between the new and old data, locate the data that needs to be modified, and generate corresponding correction suggestions based on the importance and degree of difference of the data. For updates to key performance data, verification results from at least three repeated experiments are required. s603. Update requests and correction suggestions are submitted to experts for review. Experts will review them based on the reliability of the data source, the scientific nature of the experimental methods, and the rationality of the results. After the review is passed, the database will be updated and the time, content, and reviewer information of the update will be recorded to form a complete data update history.

2. The method for database construction and data standardization based on MOF proton conductor materials according to claim 1, characterized in that: The web crawler technology automatically searches and extracts key data related to MOF proton conductor materials in multiple academic databases by setting specific keywords, wherein the keywords at least include "MOF-based proton conductor""MOF-based proton-conducting material""MOF-based protonconductive materials""proton-conducting MOF""proton-conductive MOF" "proton conduction in MOF", and can adaptively adjust the extraction rules according to the page structure and document format of different databases.

3. The method for database construction and data standardization based on MOF proton conductor materials according to claim 1, characterized in that: The distributed database management system adopts distributed storage and computing technology to meet the storage and processing needs of massive data, while supporting horizontal and vertical expansion of data. It can automatically adjust storage and computing resources according to the growth of data volume, ensuring that system performance is not lower than the set performance indicators and the data query response time does not exceed 1 second.

4. The method for database construction and data standardization based on MOF proton conductor materials according to claim 1, characterized in that: The consortium chain model of the blockchain technology assigns different permissions to each node, and ensures data consistency and immutability through a consensus mechanism. The consensus mechanism can be dynamically adjusted according to the number of participating nodes and network performance to ensure efficient data processing and system scalability.

5. The method for database construction and data standardization based on MOF proton conductor materials according to claim 1, characterized in that: The use of artificial intelligence algorithms to process complex material characterization data includes but is not limited to feature extraction and format conversion of X-ray photoelectron spectroscopy (XPS), infrared spectroscopy (IR), and thermogravimetric analysis (TGA) data. The artificial intelligence algorithms can automatically identify the original data formats generated by different instruments and equipment, and adopt corresponding processing algorithms based on the data characteristics. The extraction error of characteristic peaks does not exceed ±5%.

6. The method for database construction and data standardization based on MOF proton conductor materials according to claim 1, characterized in that: The multi-level data verification system includes checks on data format and scope, as well as predictive analysis of data rationality using machine learning models. For abnormal data, it can automatically generate detailed error reports, including error types, possible causes, and modification suggestions.

7. The method for database construction and data standardization based on MOF proton conductor materials according to claim 1, characterized in that: The data quality assessment mechanism runs a data quality assessment algorithm to score the data in the database based on the data's completeness, accuracy, consistency, and comparison results with data of the same type of materials. The scoring range is 0 to 100 points, and the data is divided into three levels: high quality (80-100 points), medium quality (40-79 points), and low quality (0-39 points). Low-quality data is marked and reviewed in key areas, and improvement plans to improve data quality are automatically generated.

8. The method for database construction and data standardization based on MOF proton conductor materials according to claim 1, characterized in that: The big data analysis tool monitors updates to academic databases and professional journal websites in real time. The update monitoring frequency can be set according to user needs, with a minimum update monitoring interval of 1 hour. It also has the function of semantic analysis of new research results and can accurately identify key information related to MOF proton conductor materials.

9. The method for database construction and data standardization based on MOF proton conductor materials according to claim 1, characterized in that: The intelligent algorithm locates the data that needs to be modified in the database based on the information of new research results, and generates correction suggestions for expert review. The correction suggestions include a detailed description of the modification content, the basis for the modification, and an assessment of the overall impact on the data. After the data is updated, the data quality is automatically re-evaluated to ensure that the updated data meets the quality requirements.

Citation Information

Cited By

  • Artificial intelligence monitoring method and system for material performance test data of heavy duty gas turbine

    CN121561376A