Machine learning screening method for near-infrared two-zone organic photothermal co-crystals
By constructing a cascaded task framework based on the XGBoost model, the problem of low development efficiency of NIR-II organic photothermal eutectic materials was solved, enabling rapid and accurate screening of high-performance eutectic materials and promoting the application of photothermal materials.
Patent Information
- Application Number
- CN202610155827.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-04
- Publication Date
- 2026-04-28
- Estimated Expiration
- 2046-02-04
AI Technical Summary
In the existing technology, the development efficiency of NIR-II organic photothermal eutectic materials is low, screening is difficult, traditional experimental methods are difficult to screen efficiently, the structure-property relationship is complex, and there is a lack of rational design schemes for high-performance materials.
Using machine learning methods, an XGBoost model was constructed to establish a cascaded task framework of "molecular structure-NIR-II performance judgment". Through the organic photothermal eutectic identification sub-model and the NIR-II screening sub-model, high-performance eutectic materials can be screened rapidly.
It significantly shortened the screening cycle from several days to several hours, reduced costs, improved prediction accuracy to 92%, achieved one-stop screening, discovered novel NIR-II organic photothermal eutectic materials, and promoted the development of solar energy utilization, biomedicine and optoelectronics.
Smart Images

Figure CN121641278B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence-enabled material innovation technology, and in particular to a machine learning screening method for near-infrared II organic photothermal eutectic. Background Technology
[0002] Solar energy, as a clean and abundant renewable energy source (wavelength range 200-2500 nm), shows great promise in the field of photothermal conversion. Compared with traditional energy sources, solar-driven photothermal conversion technology has significant advantages such as zero pollution and wide distribution. However, the core bottleneck of this technology lies in the lack of efficient photothermal conversion materials—their performance directly determines the energy utilization efficiency of the entire system.
[0003] Currently, the absorption range of most organic photothermal materials is mainly concentrated in the near-infrared I region (NIR-I, 700-1000 nm), which severely limits their application in fields such as deep photothermal conversion. In contrast, near-infrared II (NIR-II, 1000-1700 nm) photothermal materials have significant advantages:
[0004] (1) Deeper tissue penetration: NIR-II photon scattering is weaker, which enables more efficient energy transfer;
[0005] (2) Higher photothermal conversion efficiency: Long wavelength excitation is conducive to non-radiative transition processes and reduces energy loss;
[0006] (3) Lower background interference: The NIR-II band is less affected by water molecule absorption and is suitable for complex environments.
[0007] Organic eutectic materials, due to their unique donor-acceptor (DA) interaction, can achieve strong absorption in the NIR-II region through precise control of energy level structure via molecular engineering. They also possess advantages such as simple synthesis, low cost, and tunable structure. However, their development faces two major challenges: First, the chemical space is vast, and theoretically any two organic molecules can form a eutectic. Currently, more than 10,000 eutectic combinations have been reported, making it difficult to efficiently screen them using traditional experimental methods. Second, the structure-activity relationship is complex, and the correlation mechanism between microstructural features such as molecular packing patterns and electronic coupling effects and NIR-II absorption performance is still unclear. This brings significant difficulties to the rational design of high-performance eutectic materials.
[0008] Artificial intelligence (AI) technology offers a novel approach to addressing these challenges: machine learning (ML) models can rapidly analyze massive amounts of molecular descriptors (such as energy level differences, dipole moments, and hydrogen bond networks), establish quantitative relationships between eutectic composition and NIR-II performance, and simultaneously uncover nonlinear mappings between molecular structural features and absorption, enabling virtual screening of high-performance materials. Notably, interpretable AI technologies (such as SHAP analysis) can reveal key molecular interactions, guiding rational design.
[0009] Currently, AI-enabled material screening has achieved success in fields such as catalysts and battery materials, but the intelligent development of NIR-II organic photothermal eutectics remains a gap. This invention constructs an XGBoost machine learning model to establish an intelligent screening platform that integrates multi-scale features and interpretable machine learning in a cascaded task framework (C2T-Framework) for "molecular structure-NIR-II performance judgment." This platform is expected to overcome the limitations of traditional trial-and-error methods and provide innovative solutions for the research and development of next-generation photothermal materials. Summary of the Invention
[0010] The purpose of this invention is to address the problems of low development efficiency and difficulty in screening NIR-II organic photothermal eutectic materials in the existing technology, and to provide a machine learning screening method for near-infrared II organic photothermal eutectic materials.
[0011] The technical solution adopted to achieve the purpose of this invention is:
[0012] A machine learning screening method for near-infrared II absorption organic photothermal eutectics includes the following steps:
[0013] Step 1, establish a database: collect the eutectic molecules and molecular formulas with a donor-acceptor stoichiometric ratio of 1:1 to form the first database; collect the molecular formulas and maximum absorption wavelengths of the donor and acceptor molecules in the molecular eutectic to form the second database;
[0014] Step 2, Data Processing: Convert the molecular formulas in the database into molecular descriptors;
[0015] Step 3, establish and train the fast screening model: The fast screening model is a cascaded model consisting of an organic photothermal eutectic identification sub-model and an NIR-II screening sub-model, which are established through feature knowledge transfer. Both sub-models are based on the enhanced gradient boosting algorithm.
[0016] The organic photothermal cocrystal recognition sub-model takes the molecular descriptors of donors and acceptors in the first database as input and outputs the weight ranking of the molecular descriptors and the judgment and probability of cocrystal formation. When the probability is ≥ P1, it is judged as a potential cocrystal donor or acceptor. At the same time, the key molecular descriptors are determined based on the weight ranking and experience. The NIR-II screening sub-model takes the key molecular descriptors of donors and acceptors in the second database as input and outputs the judgment and probability of whether the maximum absorption wavelength is greater than 1000 nm. When the probability is ≥ P2, it is judged as a potential NIR-II cocrystal donor or acceptor. During the training of both sub-models, the classification objective is used to optimize the hyperparameters.
[0017] Step 4: Input the molecular formula of the donor and acceptor to be tested into the trained and validated rapid screening model. The organic photothermal eutectic recognition sub-model outputs the judgment and probability p1. When p1≥P1, the NIR-II screening sub-model outputs the judgment and probability p2. When p2≥P2, it is determined that the donor and acceptor can form near-infrared II organic photothermal eutectic.
[0018] In the above technical solution, in step 1, the data samples in the first database are obtained from literature and experiments, and the data samples in the second database are obtained from experimental accumulation.
[0019] In the above technical solution, in step 1, after collecting the eutectic molecules and molecular formulas, data enhancement is performed to form a first database, and the data enhancement includes changing the donor and acceptor sequences.
[0020] In the above technical solution, step 2, the step of converting molecular formulas into molecular descriptors, is as follows: first, convert the molecular formulas in the database into ASCII characters, then use the RD-KIT tool to extract molecular descriptors from the ASCII characters, and then delete samples with missing feature values.
[0021] In the above technical solution, in step 2, the molecular descriptors converted from the molecular formula in the first database include S_L (ratio of the short axis to the long axis of the molecular unit cell), M_L (ratio of the central axis to the long axis of the molecular unit cell), S_M (ratio of the short axis to the central axis of the molecular unit cell), S (length of the short axis of the molecular unit cell), globularity (sphericity), TPSA (polar surface area of the molecular unit cell), Fr_NO (proportion of nitrogen and oxygen heteroatoms in the molecule), Fr_AA (aromatic atoms), NHA (number of hydrogen bond acceptors), NHD (number of hydrogen bond donors), NRB (number of rotatable bonds), mats5are (... Molecular surface shape and size), gats6dv (molecular topology), GATS3c (short-range charge distribution in the molecule), GATS6c (medium-range charge distribution in the molecule), GATS8c (long-range charge distribution in the molecule), GATS4s (medium-range electronic topology in the molecule), GATS6s (global electronic topology in the molecule), ATS2m (mass distribution of the local framework of the molecule), Mare (average electronegativity of all atoms in the molecule), Mp (average polarizability of all atoms in the molecule), SssNH (-NH group), ETA_5 (electronegative atom).
[0022] In the above technical solution, in step 2, the molecular descriptors converted from the molecular formula in the second database include GATS3c, ETA_5, NRB, TPSA, Kappa2 (deformability of the skeleton), Kappa1 (degree of branching of the molecule as a whole), FCSP3 (sp³ hybrid carbon ratio), BCT (degree of unsaturation of the molecule), and HAC (heavy atom count).
[0023] In the above technical solution, in step 3, both the organic photothermal eutectic recognition sub-model and the NIR-II screening sub-model are XGBoost models.
[0024] In the above technical solution, in step 3, the key molecular descriptors obtained according to the weight ranking include GATS3c, ETA_5, NRB and TPSA, and the key molecular descriptors determined by experience are Kappa2, Kappa1, FCSP3, BCT and HAC.
[0025] In the above technical solution, in step 3, the hyperparameters include at least the number of trees, learning rate, maximum depth, subsampling ratio, and feature sampling ratio.
[0026] In the above technical solution, in step 3, 80% of the data in the first database and the second database are used for model training, and 20% of the data are used for model testing.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] 1. The technology provided by this invention can be used for rapid screening of organic NIR-II eutectic materials, thereby shortening the crystal growth trial and error cycle in the initial screening stage from several days to several hours, and significantly reducing the high cost of raw materials and manpower required for experimental trial and error, thus greatly reducing time and economic costs;
[0029] 2. Through a two-stage cascade framework and feature knowledge transfer strategy, physicochemical features (key molecular descriptors) are learned from large-scale eutectic formation data (source task) in the feature knowledge transfer process. These features are then transferred and applied to model training on small-scale NIR-II performance data (target task). This results in a significantly higher prediction accuracy on the test set (reaching 92% in Example 3), effectively solving the problem of low accuracy of data-driven models in the field of organic NIR-II materials due to the scarcity of high-performance samples.
[0030] 3. Experimental verification of the prediction results confirms the practicality and reliability of the method of the present invention, providing a powerful tool for the development of high-performance NIR-II organic photothermal eutectic materials;
[0031] 4. This invention can simultaneously predict eutectic formation capability and NIR-II absorption performance, achieving one-stop screening and avoiding the complexity of switching between multiple tools. It has successfully guided the discovery of a novel NIR-II organic photothermal eutectic material DPQ, with an absorption wavelength of up to 1200 nm;
[0032] 5. This invention provides an efficient, accurate, and economical solution for the development of NIR-II organic photothermal eutectic materials, and is expected to accelerate the discovery and application of new photothermal materials, and promote the development of fields such as solar energy utilization, biomedicine, and optoelectronics. Attached Figure Description
[0033] Figure 1 The performance of different types (XGBoost, KNN, DT, SVW, MLP) of organic photothermal eutectic recognition sub-models in Example 2 is shown.
[0034] Figure 2 The score represents the performance score of the organic photothermal eutectic recognition sub-model (XGBoost) in Example 2.
[0035] Figure 3 This is the confusion matrix of the organic photothermal eutectic recognition sub-model (XGBoost) in Example 2.
[0036] Figure 4 This is a SHAP explanation of the organic photothermal eutectic recognition sub-model (XGBoost) in Example 2.
[0037] Figure 5This is a comparison of the accuracy of the NIR-II screening sub-model using the feature knowledge transfer strategy in Example 3 with that of the NIR-II screening sub-model using empirical molecular descriptors.
[0038] Figure 6 The importance ranking of molecular descriptors for the NIR-II screening sub-model in Example 3.
[0039] Figure 7 The performance score of the NIR-II screening sub-model in Example 3.
[0040] Figure 8 The PR curve is the NIR-II screening sub-model in Example 3.
[0041] Figure 9 This is a SHAP explanation of the NIR-II screening sub-model in Example 3.
[0042] Figure 10 The crystal structure of the novel NIR-II DPQ organic photothermal eutectic obtained in Example 4 was experimentally verified.
[0043] Figure 11 The XRD pattern of the novel NIR-II DPQ organic photothermal eutectic obtained in Example 4 is used to verify the experiment.
[0044] Figure 12 The UV-Vis absorption spectrum of the novel NIR-II DPQ organic photothermal eutectic obtained in Example 4 is used to verify the experiment. Detailed Implementation
[0045] The present invention will be further described in detail below with reference to specific embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention.
[0046] Example 1
[0047] A machine learning screening method for near-infrared II organic photothermal eutectic crystals includes the following steps.
[0048] Step 1, Create a database:
[0049] Establish the first database (the first database is used to train and optimize the organic photothermal cocrystal recognition sub-model): collect a large number of cocrystal molecule data samples from literature and experiments, screen cocrystal molecules with a donor-acceptor stoichiometry ratio of 1:1, and perform data augmentation including but not limited to transforming the donor-acceptor sequence.
[0050] Since comparing data from different sample preparation methods and testing processes is meaningless, the second database focuses on collecting samples with similar testing processes. Specifically, comparing solid-state absorption with liquid-state absorption is not meaningful; therefore, all samples underwent the testing process for powder samples, i.e., testing the diffuse reflectance in the UV absorption spectrum and then measuring the UV absorption.
[0051] Based on the duality of each symmetric unit DA in the organic eutectic, the data is enhanced by transforming the donor and acceptor sequences to eliminate the influence of molecular position on the model results.
[0052] A second database was established (used for training and optimizing the NIR-II screening sub-model): The data collection for the NIR-II screening sub-model in the cascade framework came from laboratory accumulation, mainly collecting the molecular formulas of the donors and acceptors that form the co-crystal and the maximum value of the co-crystal absorption wavelength.
[0053] Step 2, Data Processing:
[0054] Data preprocessing, missing value handling, and standardization. Specifically, this includes: converting the molecular formulas of donors and acceptors in the database into ASCII characters; using a simplified, standardized molecular language to clearly describe the molecular structure; rigorously filtering and directly deleting samples with missing feature values to ensure that the final dataset used for model training is complete; and standardizing feature values to eliminate the influence of dimensions.
[0055] Step 3, Build and train the fast screening model:
[0056] The rapid screening model is a cascaded model consisting of an organic photothermal eutectic identification sub-model and an NIR-II screening sub-model. Both the organic photothermal eutectic identification sub-model and the NIR-II screening sub-model are established based on the enhanced gradient boosting algorithm and are cascaded through feature knowledge transfer.
[0057] When training the organic photothermal eutectic recognition sub-model, the molecular descriptors of donors and acceptors in the first database are used as input, and the weight ranking of the molecular descriptors and the judgment and probability of whether the donor and acceptor can form a eutectic are output. When the probability is greater than or equal to the first threshold P1, the corresponding donor and acceptor are judged to be potential donors and acceptors for forming a eutectic. Preferably, the first threshold P1 is 0.5. At the same time, the key molecular descriptors are determined based on the weight ranking of the molecular descriptors and in combination with empirical molecular descriptors.
[0058] When training the NIR-II screening sub-model, the key molecular descriptors of the donor and acceptor in the second database are used as input, and the output is the judgment and probability of whether the absorption wavelength of the co-crystal molecule is greater than 1000 nm. When the probability is greater than or equal to the second threshold P2, the corresponding donor and acceptor are judged to be potential donors and acceptors for forming NIR-II organic photothermal co-crystal. Preferably, the second threshold P2 is 0.5.
[0059] During the training of both the organic photothermal eutectic identification sub-model and the NIR-II screening sub-model, a classification objective is used to optimize the hyperparameters of both sub-models. Preferably, the classification training objective is binary logistic regression. The hyperparameters include at least the number of trees (n_estimators), the learning rate (learning_rate), the maximum depth (max_depth), the subsample ratio (subsample), and the feature sampling ratio (colsample_bytree). Preferably, max_depth is 3~10, learning_rate is 0.01~0.3, n_estimators is 50~1000, subsample is 0.5~1.0, and colsample_bytree is 0.5~1.0.
[0060] Step 4: Input the molecular formulas of the acceptor and donor to be tested into the trained and validated rapid screening model. First, the organic photothermal cocrystal recognition sub-model outputs the judgment and probability p1 that a cocrystal can be formed. When p1≥P1, the NIR-II screening sub-model outputs the judgment and probability p2 that the absorption wavelength of the cocrystal molecule is greater than 1000 nm. When p2≥P2, it is determined that the acceptor and donor to be tested can form a near-infrared II organic photothermal cocrystal.
[0061] Example 2
[0062] The organic photothermal eutectic recognizer model in Example 1 was obtained through the following method:
[0063] The RD-KIT tool was used to extract molecular descriptors from the donor and acceptor datasets in the first database. The following three basic attributes were selected: molecular structure, electrical properties, and physicochemical properties. Based on chemical principles (driving forces for eutectic formation: shape matching, charge transfer, hydrogen bonding, π-π stacking), 23 relevant molecular descriptors were pre-selected.
[0064] The 23 relevant molecular descriptors are: S_L (ratio of short axis to long axis of molecular unit cell), M_L (ratio of central axis to long axis of molecular unit cell), S_M (ratio of short axis to central axis of molecular unit cell), S (length of short axis of molecular unit cell), globularity (sphericity of molecular unit cell), TPSA (polar surface area of molecular unit cell), Fr_NO (proportion of nitrogen and oxygen heteroatoms in molecule), Fr_AA (aromatic atoms), NHA (number of hydrogen bond acceptors), NHD (number of hydrogen bond donors), NRB (number of rotatable bonds), and mats5are (shape and size of molecular surface). gats6dv (molecular topology), GATS3c (short-range charge distribution in the molecule), GATS6c (medium-range charge distribution in the molecule), GATS8c (long-range charge distribution in the molecule), GATS4s (medium-range electronic topology in the molecule), GATS6s (global electronic topology in the molecule), ATS2m (mass distribution of the local framework of the molecule), Mare (average electronegativity of all atoms in the molecule), Mp (average polarizability of all atoms in the molecule), SssNH (-NH group), ETA_5 (electronegative atom).
[0065] Organic photothermal eutectic recognition sub-models of different types (XGBoost, KNN, DT, SVW, MLP) were trained using the above molecular descriptors, and their performance varied. Among them, XGBoost was the best, such as... Figure 1 As shown.
[0066] Figure 2 The organic photothermal eutectic recognition sub-model (XGBoost) demonstrates high performance in identifying whether organic donor-acceptor pairs can form eutectics, with both accuracy and precision exceeding 94.00%. This also provides molecular descriptors and model foundations for the NIR-II screening sub-model.
[0067] Figure 3 To evaluate the performance of the eutectic recognition function of the organic photothermal eutectic recognition sub-model (XGBoost), under five-fold cross-validation, the confusion matrix shows that the recognition accuracy is 97.07% for those that can form eutectic and 97.58% for those that cannot form eutectic.
[0068] like Figure 4 As shown, the high-performance organic photothermal eutectic recognizer model is interpreted using the SHAP diagram. It can be seen that the molecular descriptors ETA_5 and GATS4s corresponding to electronic properties have the most significant impact on the results.
[0069] Example 3
[0070] In this embodiment, based on the cascaded framework constructed in Example 1, we achieved the transfer of feature knowledge from eutectic formation capability prediction to optical performance prediction. Specifically, the organic photothermal eutectic recognition sub-model not only completed the screening task but also identified key molecular descriptors closely related to charge transfer and intermolecular interactions through methods such as SHAP analysis. Guided by this knowledge, the NIR-II screening sub-model in this embodiment does not start feature selection from scratch but focuses on the electronic structure feature descriptors that determine the absorption range, i.e., the optical band gap.
[0071] The key molecular descriptors selected after screening by the organic photothermal eutectic recognition model are GATS3c, ETA_5, NRB, and TPSA. The subsequent supplementary empirical molecular descriptors are Kappa2 (deformability of the skeleton), Kappa1 (degree of branching of the molecule as a whole), FCSP3 (proportion of sp³ hybrid carbons), BCT (degree of unsaturation of the molecule), and HAC (heavy atom count).
[0072] like Figure 5 As shown, under the cascaded framework, the NIR-II screening sub-model (XGBoost) with the feature knowledge transfer strategy achieved a prediction accuracy of 92.0%, while the NIR-II screening sub-model using empirical molecular descriptors (Kappa2, Kappa1, FCSP3, BCT, HAC) achieved an accuracy of 88.0%. The relative error rate of the NIR-II screening sub-model was reduced by approximately 33.33%, achieving a significant performance improvement. This fully demonstrates that molecular feature knowledge transferred from the source task can effectively overcome the model performance bottleneck caused by insufficient data, providing a reliable technical path for high-performance material screening under small sample conditions.
[0073] Figure 6 The importance ranking of each molecular descriptor in the NIR-II screening sub-model for the feature knowledge transfer strategy is presented, and the structure-activity relationship of the near-infrared absorption organic photothermal eutectic design is explained in detail. This is of great significance for subsequent experimental verification of screening near-infrared absorption organic photothermal eutectics.
[0074] Figure 7 The performance of the NIR-II screening sub-model under the cascaded framework is shown. The recall, F1 score, and precision are all above 92%. Under the cascaded task framework, the model can still achieve an accuracy of up to 92% despite using a small sample of NIR-II organic photothermal eutectic.
[0075] Figure 8 The PR curves for the NIR-II screening sub-model within the cascaded framework are shown, despite the lack of data samples. The PR curves demonstrate that the model maintains an accuracy close to 1.0 even with high recall, validating its high robustness in positive class detection.
[0076] Figure 9 This represents the absolute value of the average SHAP value of the NIR-II screening sub-models within the cascaded framework. This explains why the black-box model simultaneously provides a mapping relationship between eigenvalues and performance.
[0077] Example 4
[0078] A rapid screening model was used to identify and predict the organic cocrystal potential and NIR-II of TCNQ (7,7,8,8-tetracyano-p-benzoquinone dimethylane) and 2,3-DAP (2,3-diaminonaphthalene). The results showed they were positive samples, indicating they could form a cocrystal and possess NIR-II absorption wavelengths. A novel DPQ cocrystal was successfully prepared via solvent evaporation. Crystal structure analysis was performed using a Bruker SMARTAPEX II instrument (radiation target Cu Kα, λ = 0.154 nm, 100 K), and the resulting crystal structure was analyzed using Olex2 software. The crystal structure obtained is shown below. Figure 10 As shown in the figure, a, b, and c represent the lattice vectors a-axis, b-axis, and c-axis in the unit cell, respectively.
[0079] Figure 11 The image shown is the XRD pattern of the DPQ eutectic, which demonstrates the successful synthesis of the crystal material and its high quality.
[0080] like Figure 12 As shown, the UV-Vis absorption spectrum reveals that this organic photothermal eutectic exhibits NIR-II light absorption capability, with an absorption wavelength up to approximately 1200 nm.
[0081] The above description is only a preferred embodiment of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A machine learning screening method for near-infrared II absorption organic photothermal eutectics, characterized in that, Includes the following steps: Step 1, establish a database: collect the eutectic molecules and molecular formulas with a donor-acceptor stoichiometric ratio of 1:1 to form the first database; collect the molecular formulas and maximum absorption wavelengths of the donor and acceptor molecules in the molecular eutectic to form the second database; Step 2, Data Processing: Convert the molecular formulas in the database into molecular descriptors; Step 3, establish and train the fast screening model: The fast screening model is a cascaded model consisting of an organic photothermal eutectic identification sub-model and an NIR-II screening sub-model, which are established through feature knowledge transfer. Both sub-models are based on the enhanced gradient boosting algorithm. The organic photothermal cocrystal recognition sub-model takes the molecular descriptors of donors and acceptors in the first database as input and outputs the weight ranking of the molecular descriptors and the judgment and probability of cocrystal formation. When the probability is ≥ P1, it is judged as a potential cocrystal donor or acceptor. At the same time, the key molecular descriptors are determined based on the weight ranking and experience. The NIR-II screening sub-model takes the key molecular descriptors of donors and acceptors in the second database as input and outputs the judgment and probability of whether the maximum absorption wavelength is greater than 1000 nm. When the probability is ≥ P2, it is judged as a potential NIR-II cocrystal donor or acceptor. During the training of both sub-models, the classification objective is used to optimize the hyperparameters. Step 4: Input the molecular formula of the donor and acceptor to be tested into the trained and validated rapid screening model. The organic photothermal eutectic recognition sub-model outputs the judgment and probability p1. When p1≥P1, the NIR-II screening sub-model outputs the judgment and probability p2. When p2≥P2, it is determined that the donor and acceptor can form near-infrared II organic photothermal eutectic.
2. The machine learning screening method for near-infrared II absorption organic photothermal eutectic as described in claim 1, characterized in that, In step 1, the data samples in the first database are obtained from literature and experiments, while the data samples in the second database are obtained from experimental accumulation.
3. The machine learning screening method for near-infrared II absorption organic photothermal eutectic as described in claim 1, characterized in that, In step 1, after collecting the co-crystal molecules and their molecular formulas, data enhancement is performed to form a first database. The data enhancement includes changing the donor and acceptor sequences.
4. The machine learning screening method for near-infrared II absorption organic photothermal eutectic as described in claim 1, characterized in that, In step 2, the steps for converting molecular formulas into molecular descriptors are as follows: first, convert the molecular formulas in the database into ASCII characters, then use the RD-KIT tool to extract molecular descriptors from the ASCII characters, and finally delete samples with missing feature values.
5. The machine learning screening method for near-infrared II absorption organic photothermal eutectic as described in claim 1, characterized in that, In step 2, the molecular descriptors converted from the molecular formulas in the first database include S_L, M_L, S_M, S, globularity, TPSA, Fr_NO, Fr_AA, NHA, NHD, NRB, mats5are, gats6dv, GATS3c, GATS6c, GATS8c, GATS4s, GATS6s, ATS2m, Mare, Mp, SssNH, and ETA_5.
6. The machine learning screening method for near-infrared II absorption organic photothermal eutectic as described in claim 1, characterized in that, In step 2, the molecular descriptors converted from the molecular formulas in the second database include GATS3c, ETA_5, NRB, TPSA, Kappa2, Kappa1, FCSP3, BCT, and HAC.
7. The machine learning screening method for near-infrared II absorption organic photothermal eutectic as described in claim 1, characterized in that, In step 3, both the organic photothermal eutectic recognition sub-model and the NIR-II screening sub-model are XGBoost models.
8. The machine learning screening method for near-infrared II absorption organic photothermal eutectic as described in claim 1, characterized in that, In step 3, the key molecular descriptors obtained according to the weight ranking include GATS3c, ETA_5, NRB and TPSA, and the key molecular descriptors determined by experience are Kappa2, Kappa1, FCSP3, BCT and HAC.
9. The machine learning screening method for near-infrared II absorption organic photothermal eutectic as described in claim 1, characterized in that, In step 3, the hyperparameters include at least the number of trees, learning rate, maximum depth, subsampling ratio, and feature sampling ratio.
10. The machine learning screening method for near-infrared II absorption organic photothermal eutectic as described in claim 1, characterized in that, In step 3, 80% of the data in the first database and the second database are used for model training, and 20% of the data are used for model testing.
Citation Information
Patent Citations
Eutectic prediction method based on graph neural network and deep learning framework
CN111882044A
Vitamin K eutectic material for photodynamic therapy and preparation method thereof
CN119118814A