一种面向原子对分布函数的数据库构建方法

By constructing a database based on the atomic pair distribution function, the problems of inconsistent PDF data format and poor reproducibility were solved, enabling automated data processing and efficient retrieval, and improving the data management capabilities of materials science research.

CN121658686BActive Publication Date: 2026-07-17TONGJI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TONGJI UNIV
Filing Date
2025-12-02
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing PDF experimental data lacks a unified format, processing parameters cannot be automatically extracted, data is not reproducible, quality is difficult to quantify and assess, and there is a lack of efficient retrieval and traceability management. This results in inconsistent data, poor reproducibility, and low retrieval efficiency, making it difficult to meet the needs of high-throughput scientific research and artificial intelligence applications in materials science.

Method used

By receiving and parsing atomic pair distribution function data, performing parameter consistency checks and standardized storage, calculating quality scores using cosine similarity, verifying data reproducibility, evaluating structure fit through weighted residual factors, supporting logical expression retrieval and multi-condition filtering, and providing traceable management.

Benefits of technology

It enables automated parameter extraction and standardized storage of PDF data, improves data reproducibility and comparability, enhances retrieval efficiency, supports efficient structure fitting and model evaluation, and meets the needs of high-throughput scientific research and artificial intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121658686B_ABST
    Figure CN121658686B_ABST
Patent Text Reader

Abstract

本发明涉及一种面向原子对分布函数的数据库构建方法,方法包括以下步骤:接收实测的原子对分布函数数据.gr以及对应的可选结构文件.cif和相关原始散射数据.iq / .sq / .fq,从PDF数据文件头.gr中解析实验参数,并对参数进行一致性检验,通过检验的参数和对应的PDF数据.gr、相关原始散射数据.iq / .sq / .fq及标准化元数据统一存入数据库中;当数据库接收到验证任务,则依据上传元数据重新执行I(Q)→S(Q)→F(Q)→G(r)再处理以获得验证PDF,并依据余弦相似度生成质量评分,对质量评分低于阈值的PDF数据进行标记并生成差异定位报告。与现有技术相比,本发明具有提高数据管理与查询效率,有利于定位数据差异来源,满足局域结构分析和人工智能等应用对高质量PDF数据的需求等优点。
Need to check novelty before this filing date? Find Prior Art