A method for constructing a data set for evaluating applicability of a synthetic feasibility scoring model in the field of energetic materials

By constructing a stable and representative set of energetic molecules in the field of energetic materials, the applicability of the synthesis feasibility scoring model in the field of energetic materials is solved, the robustness and reproducibility of the evaluation are improved, and the potential energetic characteristics and structural legitimacy of the candidate set are ensured.

CN122157865APending Publication Date: 2026-06-05XIAN MODERN CHEM RES INST

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN MODERN CHEM RES INST
Filing Date
2026-01-30
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

The applicability of existing synthesis feasibility scoring models in the field of energetic materials has not been systematically verified, resulting in evaluation conclusions that are sensitive to data noise and difficult to reproduce in large-scale molecular searches. Furthermore, the number of high-energy molecules is scarce and their property distribution is unstable.

Method used

A stable and representative candidate set of energetic molecules is constructed by combining structural constraint screening, energy-related property calculation, and statistical monotonicity test. This includes molecule retrieval, energy-related property calculation, threshold scanning, and monotonicity test to generate the target candidate set.

Benefits of technology

The synthesis feasibility scoring model is improved in the field of energetic materials, ensuring the potential energetic characteristics and structural legitimacy of the candidate set, suppressing local noise, and reducing subjective threshold bias.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122157865A_ABST
    Figure CN122157865A_ABST
Patent Text Reader

Abstract

The application provides a construction method of a data set for evaluating the applicability of a synthetic feasibility scoring model in the field of energetic materials. The method retrieves and standardizes the initial candidate set of 154,432 energetic molecules from the PubChem database according to element composition, functional group and molecular weight range; calculates or predicts energy-related properties such as oxygen balance, number of nitro groups, nitrogen content and energy factor; performs three-dimensional threshold scanning on oxygen balance, energy factor and nitrogen content, and performs monotonicity test on continuous properties and discrete properties respectively; screens and determines the target threshold with the smallest statistical deviation degree, and then outputs the target candidate set of 28,195 energetic molecules which are statistically stable and representative. The application can conveniently and repeatedly evaluate the applicability of any synthetic feasibility scoring model in the field of energetic materials, and also provides a new idea for data screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of cheminformatics and energetic materials data engineering, specifically to a method and system that combines structural constraint screening, energy-related property calculation, and statistical monotonicity testing to construct a statistically stable and representative set of energetic molecules, thereby supporting the verification of the applicability of a synthesis feasibility scoring model in the field of energetic materials. Background Technology

[0002] In recent years, data-driven high-throughput combinatorial design methods have become an important technology for accelerating the discovery of new molecules in the field of energetic materials. However, due to the unique characteristics of energetic materials, the synthesis of novel energetic materials presents enormous challenges. The development cycle of a new energetic molecule often takes more than ten years. Therefore, it is foreseeable that when faced with a large number of candidate molecules, the difficulty of verifying their synthesis one by one will increase exponentially. In order to reduce the risk of resource depletion, it is urgent to incorporate synthesis feasibility assessment into core indicators. Thanks to the development of cheminformatics, synthesis feasibility scoring models based on data accumulation and machine learning have gradually emerged. By setting calculation methods, they can quickly score using simple inputs, and can effectively balance accuracy and efficiency.

[0003] However, most of these synthetic feasibility scoring models are based on drug molecules and have only seen preliminary applications in the field of energetic materials, without systematic validation of their applicability. This is mainly because the molecular design and screening of energetic materials typically faces challenges such as a vast chemical space, a scarcity of available labeled samples, and data distribution noise. Especially in synthetic feasibility evaluation tasks, if the candidate molecule set used for model evaluation does not possess stable statistical properties that conform to domain priors, the evaluation conclusions will be sensitive to data noise and difficult to reproduce. Empirically, high energy levels often correspond to higher structural complexity and functional group density, resulting in a scarcity of high-energy molecules in the chemical space, with a general trend of "the higher the energy, the fewer the number of molecules." Therefore, there is an urgent need for a method that can construct a stable and interpretable candidate set of energetic molecules based on large-scale molecular retrieval, combined with energy-related property characterization and statistical monotonicity constraints. Summary of the Invention

[0004] In view of the fact that there is currently no dataset suitable for evaluating the applicability of synthesis feasibility scoring models in the field of energetic materials, the purpose of this invention is to provide a method for constructing a candidate set of energetic molecules based on "structural constraints-property calculation-statistical monotonicity test" to automatically screen out statistically stable, reasonably distributed and representative target candidate sets in a large-scale chemical space, thereby improving the robustness and reproducibility of subsequent synthesis feasibility scoring model evaluation.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A method for constructing a dataset to evaluate the applicability of a synthesis feasibility scoring model in the field of energetic materials includes the following steps: Step 1: Molecular search and basic structure screening to obtain an initial candidate set: Search for candidate molecule records from the PubChem database, set the structural constraints to contain only C, H, O and N elements and at least one -NO2 functional group, and limit the molecular weight range to 150 to 450 Da; Step 2, Calculation of energy-related properties: For each molecule in the initial candidate set, calculate the oxygen balance OB, the number of nitro groups Nitro_count, the nitrogen content N_content, and the energy factor EF. The energy factor EF is obtained through a machine learning prediction model. Step 3, Threshold Scanning and Statistical Monotonicity Test: Set lower threshold combinations for OB, EF, and N_content, and generate several threshold combinations using a three-dimensional threshold grid search; for the molecular set retained under each threshold combination, perform monotonicity tests on the continuous properties of OB, EF, and N_content, and perform monotonicity tests on the discrete properties of Nitro_count. Step 4, Threshold Selection and Output: Screen out threshold combinations that satisfy the preset conditions for both continuous and discrete properties monotonicity tests, and select the threshold combination with the smallest statistical deviation as the target threshold; based on the target threshold, screen out the target energetic molecule candidate set from the initial candidate set and output it.

[0007] Optionally, the monotonicity test of the continuous property is performed by calculating viol using the following formulas (1) and (2) to measure whether the continuous property is increasing; ; ; The collected dataset is divided and plotted according to its properties. The vertical axis represents the number of molecules contained, and the horizontal axis represents the four energy-related properties. Since the first three properties are continuous, the bin values ​​are used to divide the regions. Where b represents the number of bars in the histogram, and i represents the bar number in the histogram. This represents the height difference between the (i+1)th pillar and the ith pillar. The closer viol is to 0, the more monotonic the property distribution and the smaller the local noise. Optionally, the monotonicity test of the continuous property also includes calculating the Spearman correlation coefficient ρ between the property value and the interval index, as shown in the following formula (3), and using viol and ρ together to determine the global monotonicity; ; in This represents the center point of each bin in the histogram. This indicates the number of molecules in the corresponding bin. yes rank, yes rank, , It is the average of the ranks.

[0008] Spearman ρ ≤ -0.7 and vio ≤ 0.2 are considered to be monotonic True.

[0009] Optionally, the monotonicity test of the continuous property includes: for each set of threshold combinations, dividing the property value into 20 equally wide intervals (b=20), and counting the number of molecules corresponding to each interval. ,…, ; Define the counting difference between adjacent intervals ; If the property increases, a reverse growth will occur ( > 0), which means deviating from the expected trend of "the larger the quality, the smaller the quantity", and calculate the rate of increase viol.

[0010] Optionally, for discrete properties (Nitro_count), use formula (4) to perform a strict count monotonicity test; ; k represents the number of nitro groups; If equation (4) is satisfied, then it is denoted as monotonic True; otherwise, it is False.

[0011] Optionally, in step 1, the search results are further standardized and deduplicated. The standardization and deduplication process includes at least the removal of salts and ionic compounds, valence state verification, aromaticity verification, and structure deduplication based on InChIKey to obtain an initial candidate set.

[0012] A system for constructing a dataset for evaluating the applicability of a synthesis feasibility scoring model in the field of energetic materials, the system implementing the method for constructing a dataset for evaluating the applicability of a synthesis feasibility scoring model in the field of energetic materials as described in any of the present invention, specifically including: The molecular retrieval and preprocessing module is used to perform molecular retrieval, structural constraint screening, and standardized deduplication. The property calculation module is used to calculate or predict OB, Nitro_count, N_content, and EF. The monotonicity test module is used to perform monotonicity tests on continuous and discrete properties and output evaluation indicators. The threshold grid search and selection module is used to perform three-dimensional threshold scanning, filter threshold combinations that meet the conditions, and determine the target threshold. The candidate set output module is used to generate and output a candidate set of target energetic molecules according to the target threshold.

[0013] Compared with the prior art, the present invention has at least the following advantages: (1) Combine structural constraints with property calculations to ensure that the candidate set has potential energy content and a legal structure; (2) Introducing statistical monotonicity tests oriented towards distribution patterns can suppress local noise and abnormal peaks, thereby improving the stability of the dataset; (3) Automatic parameter selection is achieved through three-dimensional threshold grid search, reducing the deviation caused by subjective threshold setting; (4) The output target candidate set can be directly used for subsequent assessment of the applicability of the synthetic feasibility scoring model. Attached Figure Description

[0014] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings: Figure 1 This refers to the entire screening process in Example 1; Figure 2 The bar chart before and after screening shows the distribution of properties using the screening strategy of the present invention in Example 2. Detailed Implementation

[0015] Unless otherwise specified, the scientific and technical terms used in this article are intended for understanding by those skilled in the art.

[0016] This invention provides a method for constructing a dataset to evaluate the applicability of a synthesis feasibility scoring model in the field of energetic materials. This method automatically filters out statistically stable, reasonably distributed, and representative target candidate sets in a large-scale chemical space, thereby improving the robustness and reproducibility of subsequent synthesis feasibility scoring model evaluations.

[0017] The present invention will be further described below with reference to specific embodiments. It should be noted that the embodiments of the present invention are only used to explain the present invention and are not intended to limit the scope of protection of the present invention.

[0018] The method for constructing a dataset to assess the applicability of the synthesis feasibility scoring model of the present invention in the field of energetic materials includes the following steps: Step 1: Molecular search and basic structure screening to obtain an initial candidate set: Search for candidate molecule records from the PubChem database, set the structural constraints to contain only C, H, O and N elements and at least one -NO2 functional group, and limit the molecular weight range to 150 to 450 Da; Step 2, Calculation of energy-related properties: For each molecule in the initial candidate set, calculate the oxygen balance (OB), number of nitro groups (Nitro_count), nitrogen content (N_content), and energy factor (EF), where EF is obtained through a machine learning prediction model; Step 3, Threshold Scanning and Statistical Monotonicity Test: Set lower threshold combinations for OB, EF, and N_content, and generate several threshold combinations using a three-dimensional threshold grid search; for the molecular set retained under each threshold combination, perform monotonicity tests on the continuous properties of OB, EF, and N_content, and perform monotonicity tests on the discrete properties of Nitro_count. Among them, the following formulas (1) and (2) were used to calculate viol, which is used to measure whether the continuous properties (OB, EF and N_count) increase and remain in a monotonic state. This lays the foundation for subsequent screening thresholds.

[0019] ; ; Where b represents the number of bars in the histogram, here we take b=20, and i represents the bar number (1-20) in the histogram. The height difference between the (i+1)th pillar and the ith pillar is represented by . The closer viol is to 0, the more monotonic the property distribution and the smaller the local noise. The collected dataset is divided into regions according to properties and plotted. The vertical axis represents the number of molecules contained, and the horizontal axis represents the four energy-related properties. Since the first three properties are continuous, the bin value is used to divide the regions. Here, bin(b) = 20 is obtained.

[0020] In this invention, the monotonicity test for continuous properties also includes calculating the Spearman correlation coefficient ρ between the property value and the interval index, and using viol and ρ together to determine global monotonicity. Since viol only counts whether it increases and cannot reflect the magnitude of the increase, while Spearman ρ does not require a linear relationship, it is the standard statistical method for testing monotonicity. Therefore, two complementary indicators are used together to determine global monotonicity. Statistical analysis revealed that EF is difficult to achieve a completely monotonically decreasing trend. Therefore, the conditions were relaxed when searching for the threshold, requiring Spearman ρ ≤ -0.7 and viol ≤ 0.2 to be considered monotonic (True).

[0021] The Spearman correlation coefficient ρ between the property value and the interval index is calculated as shown in equation (3) below: ; in This represents the center point of each bin in the histogram. This indicates the number of molecules in the corresponding bin. yes rank, yes rank, , It is the average of the ranks.

[0022] For discrete properties (Nitro_count), use the following formula (4) to make a strict count monotonicity judgment.

[0023] ; Where k represents the number of nitro groups.

[0024] If equation (4) is satisfied, it is denoted as monotonic (True); otherwise, it is False.

[0025] Step 4, Threshold Selection and Output: Screen out threshold combinations that satisfy the preset conditions for both continuous and discrete properties monotonicity tests, and select the threshold combination with the smallest statistical deviation as the target threshold; based on the target threshold, screen out the target energetic molecule candidate set from the initial candidate set and output it.

[0026] The search results are standardized and deduplicated. The standardization and deduplication process includes at least the removal of salts and ionic compounds, valence state verification, aromaticity verification, and structure deduplication based on InChIKey to obtain an initial candidate set. In this invention, the monotonicity test of continuous properties includes: for each set of threshold combinations, dividing the property value into 20 equally wide intervals (b=20), and counting the number of molecules corresponding to each interval. ,…, Define the counting difference between adjacent intervals. If the property increases, a reverse growth will occur ( > 0), meaning a deviation from the expected trend of "the larger the quality, the smaller the quantity". The rate of increase, viol, is then calculated.

[0027] In this invention, the system includes: a molecular retrieval and preprocessing module for performing molecular retrieval, structural constraint screening, and standardization deduplication; a property calculation module for calculating or predicting OB, Nitro_count, N_content, and EF; a monotonicity test module for performing monotonicity tests on continuous and discrete properties and outputting evaluation indices; a threshold grid search and selection module for performing three-dimensional threshold scanning, screening threshold combinations that meet the conditions, and determining the target threshold; and a candidate set output module for generating and outputting a target energetic molecule candidate set according to the target threshold.

[0028] Example 1: Construction of a candidate set for feasibility assessment of energetic molecule synthesis This embodiment describes the construction of a dataset using the method of the present invention, which can be universally applied to assessing the applicability of synthesis feasibility scoring models in the field of energetic materials. The entire screening process is as follows: Figure 1 As shown, a dataset containing 28,195 energetic molecules was obtained.

[0029] 1. Molecular records were retrieved from the PubChem database, limited to those containing only CHON elements, at least one -NO2 functional group, and with molecular weights restricted to 150-450 Da. The search results were standardized and deduplicated, including desalting, deionization, valence state and aromaticity verification, and deduplication was performed using InChIKey, resulting in an initial candidate set of 154,432 energetic molecules.

[0030] 2. Calculate OB, Nitro_count, and N_content for each molecule in the initial candidate set, and predict EF using a machine learning model.

[0031] 3. Threshold Scanning and Statistical Monotonicity Test: Lower threshold combinations are set for OB, EF, and N_content, and several threshold combinations are generated using a three-dimensional threshold grid search. For the molecular set retained under each threshold combination, the monotonicity test of the continuous property is performed on OB, EF, and N_content, and the monotonicity test of the discrete property is performed on Nitro_count.

[0032] Among them, the following formulas (1) and (2) were used to calculate viol, which is used to measure whether the continuous properties (OB, EF and N_count) increase and remain in a monotonic state. This lays the foundation for subsequent screening thresholds.

[0033] ; ; Where b represents the number of bars in the histogram, here we take b=20, and i represents the bar number in the histogram (1-20, take the integer). This represents the height difference between the (i+1)th pillar and the ith pillar. The closer viol is to 0, the more monotonic the property distribution and the smaller the local noise.

[0034] The monotonicity test for continuous properties also includes calculating the Spearman correlation coefficient ρ between the property value and the interval index, as shown in equation (3), and using viol and ρ together to determine global monotonicity; ; in This represents the center point of each bin in the histogram. This indicates the number of molecules in the corresponding bin. yes rank, yes rank, , It is the average of the ranks.

[0035] Spearman ρ ≤ -0.7 and vio ≤ 0.2 are considered to be monotonic True.

[0036] For discrete properties (Nitro_count), use the following formula (4) to make a strict count monotonicity judgment.

[0037] ; Where k represents the number of nitro groups.

[0038] If equation (4) is satisfied, it is denoted as monotonic (True); otherwise, it is False.

[0039] 4. Threshold combinations that satisfy preset conditions for both continuous and discrete properties' monotonicity tests are selected, and the threshold combination with the smallest statistical deviation is chosen as the target threshold. Based on the target threshold, a target energetic molecule candidate set is obtained from the initial candidate set and output, resulting in a total of 28,195 energetic molecules. Finally, the distribution of the dataset across the four properties is as follows: Figure 2 The figures shown represent oxygen balance (OB), number of nitrogroups (Nitro_count), nitrogen content (N_content), and energy factor (EF), respectively.

[0040] Although the illustrative specific embodiments of the present invention have been described above to enable those skilled in the art to understand the invention, it should be understood that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of the present invention are protected.

Claims

1. A method for constructing a dataset to evaluate the applicability of a synthesis feasibility scoring model in the field of energetic materials, characterized in that, Includes the following steps: Step 1: Molecular search and basic structure screening to obtain an initial candidate set: Search for candidate molecule records from the PubChem database, set the structural constraints to contain only C, H, O and N elements and at least one -NO2 functional group, and limit the molecular weight range to 150 to 450 Da; Step 2, Calculation of energy-related properties: For each molecule in the initial candidate set, calculate the oxygen balance OB, the number of nitro groups Nitro_count, the nitrogen content N_content, and the energy factor EF. The energy factor EF is obtained through a machine learning prediction model. Step 3, Threshold Scanning and Statistical Monotonicity Test: Set lower threshold combinations for OB, EF, and N_content, and generate several threshold combinations using a three-dimensional threshold grid search; for the molecular set retained under each threshold combination, perform monotonicity tests on the continuous properties of OB, EF, and N_content, and perform monotonicity tests on the discrete properties of Nitro_count. Step 4, Threshold selection and output: Select threshold combinations that satisfy the preset conditions for both continuous and discrete properties monotonicity tests, and select the threshold combination with the smallest statistical deviation as the target threshold. Based on the target threshold, a target energetic molecule candidate set is obtained from the initial candidate set and output.

2. The method for constructing a dataset for evaluating the applicability of the synthesis feasibility scoring model in the field of energetic materials according to claim 1, characterized in that, The monotonicity test of the continuous property is performed by calculating viol using the following formulas (1) and (2) to measure whether the continuous property increases; ; The collected dataset was divided and plotted according to its properties. The vertical axis represents the number of molecules contained, and the horizontal axis represents the four energy-related properties. The bin values ​​were used to divide the regions. Where b represents the number of bars in the histogram, and i represents the bar number in the histogram. This represents the height difference between the (i+1)th pillar and the ith pillar. The closer viol is to 0, the more monotonic the property distribution and the smaller the local noise.

3. The method for constructing a dataset for evaluating the applicability of the synthesis feasibility scoring model in the field of energetic materials according to claim 2, characterized in that, The monotonicity test of the continuous property also includes calculating the Spearman correlation coefficient ρ between the property value and the interval index, and the calculation formula is as follows (3), and using viol and ρ together to determine the global monotonicity; ; in This represents the center point of each bin in the histogram. This indicates the number of molecules in the corresponding bin. yes rank, yes rank, , It is the average of the ranks; Spearman ρ ≤ -0.7 and vio ≤ 0.2 are considered to be monotonic True.

4. The method for constructing a dataset for evaluating the applicability of the synthesis feasibility scoring model in the field of energetic materials according to claim 2, characterized in that, The monotonicity test for the continuous property includes: for each set of threshold combinations, dividing the property value into 20 equally wide intervals (b=20), and counting the number of molecules corresponding to each interval. ,…, ; Define the counting difference between adjacent intervals ; If the property increases, a reverse growth will occur. > 0, which means deviating from the expected trend of "the larger the quality, the smaller the quantity", and the rate of increase viol is calculated.

5. A method for constructing a dataset for evaluating the applicability of the synthetic feasibility scoring model according to any one of claims 1-4 in the field of energetic materials, characterized in that, For the discrete property Nitro_count, use formula (4) to make a strict count monotonicity judgment; ; k represents the number of nitro groups; If equation (4) is satisfied, then it is denoted as monotonic True; otherwise, it is False.

6. The method for constructing a dataset for evaluating the applicability of the synthesis feasibility scoring model according to any one of claims 1-4 in the field of energetic materials, characterized in that, In step 1, the search results are further standardized and deduplicated. The standardization and deduplication process includes at least the removal of salts and ionic compounds, valence state verification, aromaticity verification, and structure deduplication based on InChIKey to obtain an initial candidate set.

7. A system for constructing a dataset to evaluate the applicability of a synthesis feasibility scoring model in the field of energetic materials, characterized in that, The system implements a method for constructing a dataset for assessing the applicability of the synthesis feasibility scoring model described in any one of claims 1-6 in the field of energetic materials, specifically including: The molecular retrieval and preprocessing module is used to perform molecular retrieval, structural constraint screening, and standardized deduplication. The property calculation module is used to calculate or predict OB, Nitro_count, N_content, and EF. The monotonicity test module is used to perform monotonicity tests on continuous and discrete properties and output evaluation indicators. The threshold grid search and selection module is used to perform three-dimensional threshold scanning, filter threshold combinations that meet the conditions, and determine the target threshold. The candidate set output module is used to generate and output a candidate set of target energetic molecules according to the target threshold.