A method, device, storage medium and product for predicting an index of a light-emitting material

CN122551994APending Publication Date: 2026-08-11BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-25
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]然而,这种联合预测的方法实质上将每个指标视为相互独立的回归任务,忽略了指标之间固有的物理关联

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551994A_ABST
    Figure CN122551994A_ABST
Patent Text Reader

Abstract

This application provides a method, device, storage medium, and product for predicting the performance indicators of luminescent materials. The method includes: inputting the molecular structure of the luminescent material to be tested into a first model to obtain intermediate performance indicators predicted based on the molecular structure; mapping the intermediate performance indicators using the first model to obtain at least two associated performance indicators related to the intermediate performance indicators; and ensuring that the at least two associated performance indicators have physical consistency. This application can predict the intermediate performance indicators of luminescent materials through a model and map the intermediate performance indicators to obtain multiple associated performance indicators with physical consistency, thereby effectively avoiding the occurrence of unreasonable combinations of predicted indicators and improving the accuracy of virtual screening of luminescent materials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular to a method, device, storage medium and product for predicting the index of luminescent materials. Background Technology

[0002] In the development of luminescent materials for OLEDs (Organic Light-Emitting Diodes), it is often necessary to jointly predict multiple spectral indicators based on molecular structure, such as emission peak wavelength, full width at half maximum (FWHM), quantum yield, and luminescence lifetime.

[0003] Existing technologies typically employ multi-task graph neural networks to simultaneously predict these indicators, that is, by utilizing the shared molecular structure representation of these indicators to output the predicted values ​​of each indicator in order to achieve joint prediction.

[0004] However, this joint prediction method essentially treats each indicator as an independent regression task, ignoring the inherent physical correlation between the indicators. For example, quantum yield and luminescence lifetime have a definite coupling relationship in photophysics, but existing technologies may output unreasonable combinations of prediction results, such as high quantum yield and abnormally short luminescence lifetime, leading to a high false positive rate when virtually screening luminescent materials. Summary of the Invention

[0005] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide a method, device, storage medium and product for predicting the index of luminescent materials. The method predicts the intermediate performance index of luminescent materials through a model and maps the intermediate performance index to obtain multiple physically consistent correlated performance indexes, thereby effectively avoiding the occurrence of unreasonable combinations of prediction indexes and improving the accuracy of virtual screening of luminescent materials.

[0006] In a first aspect, this application provides a method for predicting the indicators of luminescent materials, the method comprising: The molecular structure of the luminescent material to be tested is input into the first model to obtain intermediate performance indicators based on molecular structure prediction. The intermediate performance index is mapped using the first model to obtain at least two associated performance indices; the at least two associated performance indices have physical consistency.

[0007] In conjunction with the first aspect, in one possible implementation, the training process of the first model includes: A sample set is obtained, which includes the molecular structure of various sample luminescent materials and the performance index labels of each sample luminescent material; the performance index labels are used to characterize the measured performance of the sample luminescent materials, and at least two performance index labels have physical consistency. Input the sample set into the initial model to obtain the predicted intermediate performance index and multiple predicted performance indices. Among them, the multiple predicted performance indices include at least two associated performance indices, which are obtained by mapping the predicted intermediate performance indices from the initial model. The initial model is iteratively trained based on the loss between the predicted performance metric and the corresponding performance metric label to obtain the first model.

[0008] In conjunction with the first aspect, one possible implementation method also includes: The sample set is input into the first model, and a first pseudo-label is added to the luminescent material of the sample based on the first model. The first pseudo-label is a label for the associated performance index. Update the first model based on the updated sample set; The updated sample set is input into the updated first model, and a second pseudo-label is added to the luminescent material of the sample based on the prediction performance index output by the updated first model. The first model is updated again based on the updated sample set.

[0009] In conjunction with the first aspect, in one possible implementation, the sample set is input into a first model, and a first pseudo-label is added to the luminescent material of the sample based on the first model, including: Input the sample set into the first model to obtain the correlation performance index output by the first model. The correlation performance index is obtained by mapping the intermediate performance index predicted by the first model. The first pseudo-label is added to the sample luminescent material based on the correlation performance index output by the first model. If there are abnormal samples with redundant labels, the redundant labels of the abnormal samples are deleted based on the differences between the labels of the abnormal samples.

[0010] In conjunction with the first aspect, in one possible implementation, a second pseudo-label is added to the sample luminescent material based on the predicted performance index of the updated first model output, including: If the predicted performance index output by the updated first model is an index other than the associated performance index, then the confidence level of the predicted performance index is determined. If the confidence level is higher than the confidence level threshold, then the predicted performance index is added as the second pseudo-label of the sample luminescent material. If the predicted performance index output by the updated first model is a correlated performance index, then the direct predicted value of the correlated performance index is determined, and the direct predicted value is directly output by the updated first model; and the predicted performance index that meets the first condition is added as the second pseudo-label of the sample luminescent material. The first condition is that the difference between the direct predicted value and the predicted performance index is less than the deviation threshold, the confidence level of the predicted performance index is higher than the confidence level threshold, and the predicted performance index is within the standard value range.

[0011] In conjunction with the first aspect, in one possible implementation, the loss when updating the first model again includes: the measured label loss term, the first pseudo-label loss term, the second pseudo-label loss term, and the consistency loss term, wherein the consistency loss term is determined based on the association performance index directly predicted by the first model and the association performance index obtained by mapping in the next update.

[0012] In conjunction with the first aspect, in one possible implementation, the weights of the measured label loss term, the first pseudo-label loss term, and the second pseudo-label loss term are all determined based on the label attributes, which include the measured label and the pseudo-label. The weight of the consistency loss term is determined based on a preset value range and the consistency error. The consistency error is determined based on the correlation performance index directly predicted by the first model and the correlation performance index obtained by mapping when the model is updated again.

[0013] In conjunction with the first aspect, one possible implementation method also includes: Construct a missing mask matrix for the sample set. The missing mask matrix is ​​used to label the missing labels in the sample set. During iterative training, gradient backpropagation at missing labels is masked based on the missing mask matrix.

[0014] In conjunction with the first aspect, in one possible implementation, the associated performance metrics include quantum yield and luminescence lifetime; Intermediate performance indicators include radiative attenuation intensity and non-radiative attenuation intensity.

[0015] In a second aspect, this application also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method of any one of the first aspects described above.

[0016] Thirdly, this application also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method of any one of the first aspects above.

[0017] Fourthly, this application also provides a computer program product containing instructions that, when executed, perform any of the methods described in the first aspect above. Attached Figure Description

[0018] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 Here is a flowchart of the method of this application in one embodiment; Figure 2As one embodiment, this application's first model training flowchart is shown. Figure 3 Here is a flowchart illustrating the process of adding the first and second pseudo-tags in one embodiment of this application; Figure 4 Here is a flowchart of the screening process for the first pseudo-label in this application, as shown in one embodiment. Figure 5 This is a structural diagram of the computer device of this application in one embodiment. Detailed Implementation

[0019] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0020] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present application will now be described in detail with reference to the accompanying drawings and embodiments. Furthermore, the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The terms "first" and "second," etc., in the specification and claims of the embodiments of this application are used to distinguish different objects, not to describe a specific order of objects.

[0021] In the research and development of OLED luminescent materials, it is often necessary to measure a number of key spectral indicators of the molecular structure of the luminescent materials, including emission peak wavelength, full width at half maximum (FWHM), quantum yield, and luminescence lifetime. These indicators are important bases for evaluating and screening high-performance luminescent materials.

[0022] Existing technologies typically employ multi-task graph neural networks, which output predicted values ​​for each indicator based on the shared molecular structure of these indicators, in order to achieve joint prediction.

[0023] However, existing technologies have the following problems: First, the joint prediction method essentially treats each indicator as an independent regression task, ignoring the inherent physical relationship between the indicators. For example, quantum yield and luminescence lifetime have a definite coupling relationship in photophysics, but existing technologies may output unreasonable combinations of prediction results such as high quantum yield and abnormally short luminescence lifetime, leading to a high false positive rate when virtually screening luminescent materials, and the prediction process is not transparent, making it impossible to explain the physical reasons behind the prediction results. Secondly, in actual R&D data, quantum yield and emission lifetime are measured only in a small number of samples in the database due to high measurement costs and long cycles. In contrast, emission peak wavelength and full width at half maximum (FWHM) are commonly measured, resulting in a structured missing distribution with significant differences in the number of labels among the indicators. Current multi-task model training methods typically only calculate prediction errors and update parameters for the measured indicator positions, excluding unmeasured positions. This makes the supervision signals for quantum yield and emission lifetime too sparse, making it difficult for the model to fully learn these two indicators, easily leading to underfitting and training instability. Simultaneously, dense prediction tasks (emission peak wavelength, FWHM) involve many training samples and frequent error terms, causing the model parameters to be adjusted primarily towards predicting the emission peak wavelength and FWHM, severely suppressing the predictive ability of quantum yield and emission lifetime.

[0024] Based on this, embodiments of this application provide a method, device, storage medium, and product for predicting the performance indicators of luminescent materials. This method can predict intermediate performance indicators of luminescent materials through a model and perform mapping processing on these intermediate performance indicators to obtain multiple physically consistent correlated performance indicators. This effectively avoids the occurrence of unreasonable combinations of predicted indicators and improves the accuracy of virtual screening of luminescent materials.

[0025] Firstly, such as Figure 1 As shown, this application provides a method for predicting the indicators of luminescent materials, the method comprising: S101. Input the molecular structure of the luminescent material to be tested into the first model to obtain intermediate performance indicators based on molecular structure prediction. S102. The intermediate performance index is mapped using the first model to obtain at least two related performance indices; the at least two related performance indices have physical consistency.

[0026] Specifically, molecular structure refers to the chemical structural information of OLED light-emitting materials, including the atomic composition of the molecule, the connection relationship between atoms, bond order, functional groups, and spatial configuration, which are chemical descriptions that can uniquely identify the molecule's identity and properties.

[0027] In this application, the first model can be a multi-task prediction model based on a graph neural network, whose input is the molecular structure of the luminescent material (e.g., a molecular graph with atoms as nodes and chemical bonds as edges), and whose output is the correlation performance index of the luminescent material.

[0028] In one possible implementation, the luminescent material of this application is primarily geared towards OLEDs. For example, quantum yield φ and luminescence lifetime τ are a pair of correlated performance indicators of the luminescent material. The intermediate performance indicators of the luminescent material include the radiative decay intensity variable. k r ( i ) and non-radiative attenuation intensity variablesk nr ( i It is worth mentioning that the radiation attenuation intensity variable k r ( i ) and non-radiative attenuation intensity variables k nr ( i All values ​​are positive. Quantum yield f ( i ) and luminescence lifetime τ( i The mapping relationships between these parameters and intermediate performance indicators are as follows:

[0029] In the formula, k r ( i ) represents the radiation attenuation intensity variable; k nr ( i ) represents a non-radiative attenuation intensity variable; i For the first i One luminescent material to be tested.

[0030] Specifically, physical consistency refers to the requirement that multiple related physical quantities obtained in scientific computing or machine learning predictions must satisfy known physical laws, conservation relationships, or empirical rules. This includes fulfilling specific mathematical equations, dimensional consistency, numerical range constraints, or causal logic. Simply put, the results cannot violate the fundamental rules of the objective physical world, such as the conservation of energy, the second law of thermodynamics, or kinematic equations. In the index predictions of this application, ensuring the physical consistency of multiple related performance indicators helps improve the reliability and interpretability of the predictions and avoids predictive combinations that are impossible in reality.

[0031] In this application, quantum yield f and luminescence lifetime t The physical consistency of is reflected in: f The increase will inevitably be accompanied by t The corresponding change (usually a decrease), and f Always between 0 and 1 t Since the value is always positive, the predicted combination of high quantum yield and abnormally short or long lifetimes violates the actual physical mechanism.

[0032] This embodiment proposes... f – t The physical consistency output mechanism, specifically, involves the first model predicting an interpretable intermediate performance metric—radiative attenuation intensity. krWith non-radiative attenuation intensity knr Then, derived from the physical relationship of light. f and t , so that the obtained f and t Naturally satisfying physical consistency and constraints provides a more reliable basis for subsequent screening of luminescent materials.

[0033] Furthermore, this application also makes the predicted results of quantum yield and luminescence lifetime physically interpretable, that is, researchers see... k r and k nr The specific values ​​can directly determine whether the low luminous efficiency is due to slow radiative decay or fast non-radiative decay, or which factor is dominant in determining the abnormal lifespan.

[0034] In one possible implementation, such as Figure 2 As shown, the training process of the first model includes: S201. Obtain a sample set, which includes the molecular structure of various sample luminescent materials and the performance index labels of each sample luminescent material; the performance index labels are used to characterize the measured performance of the sample luminescent materials, and at least two performance index labels have physical consistency. In one possible implementation, the performance indicator could be the emission peak wavelength. lem Half-peak width ( Full Width at Half Maximum , FWHM Quantum yield f and luminescence lifetime t .

[0035] Specifically, after obtaining the sample set, it is determined whether multiple performance indicator labels can be used to deduce intermediate performance indicators through physical formulas. If so, the multiple performance indicator labels satisfy physical consistency. For example, in the performance indicators of luminescent materials for OLEDs, quantum yield... f and luminescence lifetime t Can be determined by radiation attenuation intensity kr With non-radiative attenuation intensity knr The derivation yields the quantum yield. f and luminescence lifetime t It has physical consistency.

[0036] In one possible implementation, this application also performs unit normalization on the labels of each performance indicator.

[0037] In one possible implementation, this application also includes labels for performance metrics with physical consistency (e.g., lifetime). tThe variables were transformed to improve training stability. Preferably, the transformation can be a logarithmic transformation.

[0038] In one possible implementation, the sample set also includes test condition information to enable the first model to predict performance metrics under specified test conditions.

[0039] Optionally, the test conditions information can be at least one of the following: different solvents, different temperatures, different doping concentrations, or different excitation wavelengths.

[0040] It is worth mentioning that, due to limitations in actual measurement costs and experimental cycles, some samples have missing performance indicator labels.

[0041] S202. Input the sample set into the initial model to obtain the predicted intermediate performance index and multiple predicted performance indexes. Among them, the multiple predicted performance indexes include at least two related performance indexes, which are obtained by the initial model through mapping the predicted intermediate performance indexes. In other words, in the output of the initial model, the associated performance indicators need to be obtained through intermediate performance indicator mapping; while the prediction results of the other performance indicators are directly output by the initial model.

[0042] One possible implementation is an initial model based on a graph neural network, which extracts deep features of the molecular structure through a graph neural network encoder and outputs the prediction results of each performance index through different output branches.

[0043] S203. Iteratively train the initial model based on the loss between the predicted performance index and the corresponding performance index label to obtain the first model.

[0044] It is worth mentioning that during the training of the first model in this application, since multiple associated performance indicators are mapped to the same intermediate performance indicator, when a certain associated performance indicator lacks a measured label, the measured label of another associated performance indicator can still indirectly supervise the performance indicator prediction task (denoted as the sparse task) with missing labels through the intermediate performance indicator. This realizes the supervision transfer in the missing scenario, significantly improves the learnability of the sparse task and the overall data utilization efficiency, and ultimately provides physically reliable and mechanistically explainable prediction support for the research and development of OLED luminescent materials.

[0045] One possible implementation method also includes: Construct a missing mask matrix for the sample set. The missing mask matrix is ​​used to label the missing labels in the sample set. During iterative training, gradient backpropagation at missing labels is masked based on the missing mask matrix.

[0046] Specifically, this application first constructs a missing mask matrix M, which has the same dimension as the molecular structure-label pair matrix in the sample set. Each element M(i,k) in M ​​is used to mark whether the i-th sample has a measured label for the k-th performance metric: if M(i,k)=1, it means that the position has a measured label; if M(i,k)=0, it means that the label is missing at that position. During iterative training, the loss is calculated and backpropagation is performed only when M(i,k)=1, while positions where M(i,k)=0 are not involved in loss calculation or parameter updates at all. In this way, this application explicitly incorporates the label missing state into the iterative training process, enabling precise control of loss calculation and gradient scheduling based on the missing mask matrix.

[0047] One possible implementation is, such as Figure 3 As shown, the method also includes: S301. Input the sample set into the first model, and add a first pseudo-label to the sample luminescent material based on the first model. The first pseudo-label is a label for the associated performance index. like Figure 4 As shown, in one possible implementation, S301 includes: S401. Input the sample set into the first model to obtain the correlation performance index output by the first model. The correlation performance index is obtained by mapping the intermediate performance index predicted by the first model. One possible implementation involves the first model encoding the molecular structure of each luminescent material when inputting the sample set, resulting in a fixed-length numerical vector. This numerical vector captures the overall characteristics of the molecular structure; it is not generated separately for a specific indicator prediction task, but rather represents the fundamental features used by all indicator prediction tasks.

[0048] S402. Based on the performance index output by the first model, add a first pseudo-label to the sample luminescent material. If there are abnormal samples with redundant labels, delete the redundant labels of the abnormal samples based on the differences between the labels of the abnormal samples.

[0049] Specifically, when the sample set is input into the first model, the first model will not only predict the association performance index for missing labels, but will output the association performance index based on the predicted intermediate performance index, regardless of whether there are actual labels. This leads to label redundancy, which means that one or more association performance indices for anomaly samples simultaneously exist with both actual and predicted labels.

[0050] To address the issue of label redundancy, this application calculates the error value between multiple labels of abnormal samples. When the error value is less than a preset threshold, the first pseudo-label is retained as a label for key performance indicators, and the actual label is removed. When the error value is greater than the preset threshold, the actual label is retained, and the first pseudo-label is removed.

[0051] Optionally, when the error value is greater than the preset threshold, you can also choose to still remove the measured label and retain the first pseudo label, but when updating the first model in the future, the weight of the loss term corresponding to the first pseudo label will be preset to a lower value.

[0052] One possible implementation is that the error value is equal to the absolute value of the difference between the measured label and the first pseudo label.

[0053] In another possible implementation, redundant labels can be deleted simultaneously based on the confidence level and error value of the first pseudo-label. For example, when the confidence level of the first pseudo-label is less than the confidence threshold and the error value is greater than the deviation threshold, the redundant first pseudo-label is deleted, and the measured label is retained; otherwise, the measured label is deleted, and the first pseudo-label is retained. Each confidence level is output simultaneously by the first model when outputting each prediction result.

[0054] S302. Update the first model based on the updated sample set; For ease of description, S302 will be referred to as the first update phase. The loss in the first update phase is: L_stage1= L_real1+ μL_derived1+ ηL_phy1 In the formula: L_stage1 is the loss of the first update stage; L_real1 is the measured label loss value in the sample set; L_derived1 is the first pseudo-label loss value; L_phy1 is the physical consistency loss term; m and or The values ​​are the weights of the loss items, all of which are preset values ​​and can be set by those skilled in the art.

[0055] In one possible implementation, besides predicting intermediate performance indicators and then mapping them to obtain related performance indicators, the first model can also directly output the predicted values ​​of the related performance indicators. In this case, the physical consistency loss term is determined based on the related performance indicators directly predicted by the first model and the mapped related performance indicators in the first update phase. The specific determination method is as follows:

[0056] In the formula, Y_pred For the directly predicted first j One related performance metric; Y_phy The first obtained by mapping j One related performance metric; W j For the preset first j Consistency weights for each associated performance metric; m Indicates shared ownership mThere are several related performance indicators. It is worth mentioning that, since the dimensions and numerical ranges of different related performance indicators are significantly different, such as quantum yield, which is usually between 0 and 1, while luminescence lifetime may span multiple orders of magnitude (from nanoseconds to milliseconds), the numerical ranges of various related performance indicators need to be aligned first during the summation process.

[0057] The following uses OLED luminescent materials as an example to illustrate the calculation process for physical consistency loss: L_phy1=λ_φ · |φ_pred1- φ_phy1|+λ_τ · |log(τ_pred1+ ε) - log(τ_phy1+ ε)| In the formula, L_phy1 is the physical consistency loss term; λ_ f and λ_ t All are preset values, λ_ f λ_ represents the quantum yield consistency weight. t Lifetime consistency weight; f _pred1 represents the quantum yield directly predicted by the first model in the first update phase. f ; t _pred1 represents the lifetime τ directly predicted by the first model in the first update phase; f _phy1 represents the quantum yield obtained from the first model mapping during the first update phase. f ; t _phy1 represents the lifetime obtained from the first model mapping during the first update phase. t The log is used for numerical alignment.

[0058] The following are m and or The preset rules for the weights of each loss item are explained below: From the calculation formula of L_stage1, we know that the weight of the measured label loss is 1. m Determined according to the following rules: Actual tag weight > First pseudo tag weight m ; Example, m It can be 0.8 to 0.9.

[0059] about or The value range can be set to 0.05~0.1, and dynamically determined based on physical consistency. or When the physical consistency error is large, take or Set a small value (e.g., 0.05) to reduce the contribution of the physical consistency loss term to the total loss; when the physical consistency error is small, take... orA larger value (e.g., 0.1) is used to strengthen the constraint on the physical consistency loss term.

[0060] In one possible implementation, for the correlation performance index, the difference between the predicted value directly obtained by the first model and the predicted value after mapping is called the physical consistency error. The physical consistency error is calculated as follows:

[0061] In the formula, Y_pred For the directly predicted first j One related performance metric; Y_phy The first obtained by mapping j One related performance metric; m Indicates shared ownership m There are several related performance indicators. It is worth mentioning that, since the dimensions and numerical ranges of different related performance indicators are significantly different, such as quantum yield, which is usually between 0 and 1, while luminescence lifetime may span multiple orders of magnitude (from nanoseconds to milliseconds), the numerical ranges of various related performance indicators need to be aligned first during the summation process.

[0062] S303. Input the updated sample set into the updated first model, and add a second pseudo-label to the sample luminescent material based on the prediction performance index output by the updated first model. In one possible implementation, a second pseudo-label is added to the sample luminescent material based on the predicted performance metrics of the updated first model output, including: If the predicted performance index output by the updated first model is an index other than the associated performance index, then the confidence level of the predicted performance index is determined. If the confidence level is higher than the confidence level threshold, then the predicted performance index is added as the second pseudo-label of the sample luminescent material. If the predicted performance index output by the updated first model is a correlated performance index, then the direct predicted value of the correlated performance index is determined, and the direct predicted value is directly output by the updated first model; and the predicted performance index that meets the first condition is added as the second pseudo-label of the sample luminescent material. The first condition is that the difference between the direct predicted value and the predicted performance index is less than the deviation threshold, the confidence level of the predicted performance index is higher than the confidence level threshold, and the predicted performance index is within the standard value range.

[0063] In one possible implementation, the difference between the direct predicted value and the predicted performance index is the physical consistency error, which is the same as the meaning and determination principle of the physical consistency calculated in S302. The relevant values ​​can be substituted into the direct predicted value and predicted performance index of the associated performance index of the first model in this stage, without further explanation.

[0064] The following is based on luminescence lifetime t and quantum yield fAs a performance indicator related to the correlation, the first condition is explained as follows: When adding the second pseudo-label, for the updated first model, the intermediate performance metrics predicted by the first model are first mapped to obtain the associated performance metrics: f _phy = k_r _pred / ( kr_ pred + knr _pred + ε) t _phy = 1 / ( kr _pred + k_nr _pred + ε) In the formula, f _phy is obtained by mapping the first model. f The predicted value, for ease of description, is denoted as... f The mapping value; correspondingly, t _phy is t The mapping value; kr _pred is an intermediate performance metric kr The predicted value; knr _pred is an intermediate performance metric knr The predicted value.

[0065] Next, calculate the physical consistency error: E _phy =| f _pred - f _phy|+|log( t _pred + ε) - log( t _phy + ε)| in, E _phy represents the physical consistency error; ε is a small constant to prevent division by zero. f _pred represents the output of the first model without intermediate performance metric mapping. f The direct predicted value; similarly, t _pred is the direct prediction of τ; log is used to make t and f Align the numerical range.

[0066] The standard range of values ​​refers to: quantum yield f It cannot be less than 0, and usually cannot be greater than 1; t Lifespan must be positive; kr and knr It is the rate constant and cannot be negative.

[0067] The confidence score is output simultaneously by the updated first model when outputting each prediction performance metric.

[0068] This application transforms the "label missing position" from unsupervised to "controlled supervised" by adding a first pseudo-label and a second pseudo-label. Furthermore, through a confidence screening mechanism, the quality of the second pseudo-label is directly incorporated into the subsequent updates of the first model, thus preventing inferior pseudo-labels from deviating from the update direction of the first model.

[0069] S304. Update the first model again based on the updated sample set.

[0070] This application does not directly use the output of the first model to fill in all missing labels at once. This is because if the predictions of the first model were directly used as pseudo-labels for training, the inherent noise in the pseudo-labels would self-reinforce during training, causing systematic bias in the first model and ultimately reducing the reliability of the virtual screening. Therefore, this application first uses the measured labels and the screened first pseudo-labels to train the first model to a stable state, giving it reliable predictive capabilities. Then, based on this stable model, high-quality second pseudo-labels are generated and used together with the first pseudo-labels and measured labels to train the first model. This approach effectively improves the effective supervision density and convergence stability. Furthermore, because the pseudo-labels fill in data gaps, for indicator prediction tasks with severe label shortages (such as quantum yield and luminescence lifetime), the relative ranking of the first model's predicted values ​​is less affected by random sampling fluctuations, thereby improving the consistency of the screening ranking when screening luminescent materials multiple times.

[0071] In one possible implementation, the loss when updating the first model again includes: the measured label loss term, the first pseudo-label loss term, the second pseudo-label loss term, and the consistency loss term, wherein the consistency loss term is determined based on the association performance index directly predicted by the first model and the association performance index obtained by mapping in the next update.

[0072] Let the second update be denoted as the second update phase. The formula for calculating the loss in the second update phase is: L_stage2= L_real2+ μL_derived2+ νL_pseudo2+ ηL_phy2 In the formula: L_stage2 is the loss of the second update stage; L_real2 is the measured label loss term; L_derived2 is the first pseudo-label loss term; L_phy2 is the consistency loss term; L_pseudo is the second pseudo-label loss term; in the formula, m The weights of the first pseudo-label loss term; n The weights of the second pseudo-label loss term; or The weight of the consistency loss term.

[0073] The following is an example of the formula for calculating the consistency loss term: L_phy2=λ_φ · |φ_pred2- φ_phy2| +λ_τ · |log(τ_pred2+ ε) - log(τ_phy2+ ε)| In the formula, L_phy2 is the consistency loss term; λ_ f and λ_ t All values ​​are preset values, the same as those used in the first update phase, λ_ f λ_ represents the weight of the quantum yield consistency term. t Weights for lifetime consistency terms; f _pred2 represents the quantum yield directly predicted by the first model in the second update phase. f ; t _pred2 represents the lifetime directly predicted by the first model in the second update phase. t ; f _phy2 represents the quantum yield obtained from the first model mapping during the second update phase. f ; t _phy2 represents the lifetime obtained from the first model mapping during the second update phase. t The purpose of using log is to align the values ​​of multiple related performance metrics.

[0074] In one possible implementation, the weights of the measured label loss term, the first pseudo-label loss term, and the second pseudo-label loss term are all determined based on the label attributes, which include the measured label and the pseudo-label. The weight of the consistency loss term is determined based on a preset value range and the consistency error. The consistency error is determined based on the correlation performance index directly predicted by the first model and the correlation performance index obtained by mapping when the model is updated again.

[0075] The consistency error has the same meaning and determination formula as the physical consistency error in S302, except that the numerical value is the corresponding value output by the first model in the second update stage, which will not be elaborated further.

[0076] The following is about m , n and or The weights of each loss term are explained below. From the calculation formula for L_stage2, we know that the weight of the measured label loss term is 1. m and n Determined according to the following rules: The weight of the actual labeled loss term is greater than the weight of the first pseudo-label loss term. m The weights of the second pseudo-label loss term n ; Example, mIt can be 0.8~0.9; n It can be 0.7~0.8.

[0077] about or The value range can be set to 0.05~0.1, and dynamically determined based on the consistency error. or It is worth mentioning that, or It is usually negatively correlated with the magnitude of the consistency error; when the consistency error is large, take... or Set a small value (e.g., 0.05) to reduce the contribution of the consistency loss term to the total loss; when the consistency error is small, take... or A larger value (e.g., 0.1) is used to strengthen the constraint on the consistency loss term.

[0078] It is worth mentioning that, for the same performance metric, the consistency weight can take the same preset value when calculating the physical consistency loss in the first update phase and the consistency loss in the second update phase; correspondingly, for the same performance metric, the first pseudo-label weight in the first update phase... m Weight of the first pseudo-label in the second update phase m They can take the same value.

[0079] This application generates and adds a first pseudo-label based on the mapping relationship between correlation performance indicators and intermediate performance indicators, and then filters the first pseudo-label. Subsequently, a two-stage training strategy is adopted: in the first update stage, a first model is trained based on the measured labels and the filtered first pseudo-label; in the second update stage, a second pseudo-label, which has been filtered again, is introduced and trained together with the measured labels and the first pseudo-label to further improve the model performance. Through the above methods, this application significantly improves the utilization rate of the sample set, effectively reduces the interference of pseudo-label noise, and ensures that the prediction results always conform to the physical laws of OLED luminescent materials.

[0080] Furthermore, in the design of the loss function, this application assigns different weights to each training label based on the source type differences: the highest weight is set for labels measured in real experiments to maintain a faithful fit to the real data; a medium weight is set for the first pseudo-label calculated from intermediate performance metrics using physical formulas to strengthen physical consistency constraints; and a lower weight is set for the second pseudo-label generated by the first model and filtered by confidence to reduce the negative impact of residual noise in the pseudo-labels on model training. Through this source-weighted strategy, the first model can fully utilize various labels to increase supervision density during training, effectively suppress interference from unreliable labels, and highlight the constraining effect of the first pseudo-label on the model output.

[0081] In one possible implementation, the updated sample set includes three types of masks: M_real: the measured label mask; M_derived: the first pseudo-label mask; and M_pseudo: the second pseudo-label mask. This application introduces three masks, M_real, M_derived, and M_pseudo, to respectively label the measured label, the first pseudo-label, and the second pseudo-label in the sample set. The distribution of performance metrics. This classification mask design enables the retraining process to accurately distinguish supervision signals from different sources, thus allowing for independent weighting and participation strategies for each label class during loss calculation: the actual measured label has the highest weight, ensuring the guidance of real data; the first pseudo-label serves as intermediate supervision; the second pseudo-label, after confidence screening, participates with a lower weight to suppress noise interference. Simultaneously, the mask system provides a clear identifier for selective gradient backpropagation during training, significantly improving the controllability of multi-source label mixed training.

[0082] It should be noted that although the operations of the method of the present invention are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0083] On the other hand, this application also provides a computer-readable storage medium, which may be included in a computer device or exist independently without being assembled into the computer device. The aforementioned computer-readable storage medium stores one or more programs that, when used by one or more processors, execute the methods described in this application. For example, it may execute... Figure 1 The steps of the method shown.

[0084] In another aspect, this application also provides a computer device 50, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements... Figure 1 The steps of the method shown.

[0085] The following is for reference. Figure 5 , Figure 5 A schematic diagram of the structure of a computer device 50 suitable for implementing embodiments of this application is shown. like Figure 5As shown, the computer device 50 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 502 or programs loaded from storage section 508 into random access memory (RAM) 503. RAM 503 also stores various programs and data required for the system's operating instructions. CPU 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0086] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 510 as needed so that computer programs read from it can be installed into storage section 508 as needed.

[0087] Specifically, according to embodiments of this application, Figure 1 The described process can be implemented as a computer software program. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program contains program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit (CPU) 501, it performs the functions defined in the system of this application.

[0088] On another front, this application also provides a computer-readable storage medium, which may be included in an electronic device or exist independently without being assembled into the electronic device. The aforementioned computer-readable storage medium stores one or more programs that, when used by one or more processors, execute the methods of this application. For example, it may execute... Figure 1 The steps of the method shown.

[0089] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0090] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A method for predicting the indicators of luminescent materials, characterized in that, The method includes: The molecular structure of the luminescent material to be tested is input into the first model to obtain intermediate performance indicators predicted based on the molecular structure. The intermediate performance index is mapped using the first model to obtain at least two associated performance indices related to the intermediate performance index; the at least two associated performance indices have physical consistency with each other.

2. The method according to claim 1, characterized in that, The training process of the first model includes: A sample set is obtained, which includes the molecular structures of various sample luminescent materials and the performance index labels of each sample luminescent material; the performance index labels are used to characterize the measured performance indexes of the sample luminescent materials, and at least two of the performance index labels have physical consistency. The sample set is input into the initial model to obtain the predicted intermediate performance index and multiple predicted performance indexes. The multiple predicted performance indexes include at least two related performance indexes, which are obtained by the initial model through mapping the predicted intermediate performance indexes. The initial model is iteratively trained based on the loss between the predicted performance metric and the corresponding performance metric label to obtain the first model.

3. The method according to claim 2, characterized in that, The method further includes: The sample set is input into the first model, and a first pseudo-label is added to the luminescent material of the sample based on the first model. The first pseudo-label is the label of the correlation performance index. Update the first model based on the updated sample set; The updated sample set is input into the updated first model, and a second pseudo-label is added to the luminescent material of the sample based on the prediction performance index output by the updated first model. The first model is updated again based on the updated sample set.

4. The method according to claim 3, characterized in that, The step of inputting the sample set into the first model and adding a first pseudo-label to the luminescent material of the sample based on the first model includes: The sample set is input into the first model to obtain the correlation performance index output by the first model. The correlation performance index is obtained by mapping the intermediate performance index predicted by the first model. Based on the correlation performance index output by the first model, a first pseudo-label is added to the sample luminescent material. If there are abnormal samples with redundant labels, the redundant labels of the abnormal samples are deleted based on the differences between the labels of the abnormal samples.

5. The method according to claim 3, characterized in that, The method of adding a second pseudo-label to the sample luminescent material based on the predicted performance index output by the updated first model includes: If the predicted performance index output by the updated first model is an index other than the associated performance index, then the confidence level of the predicted performance index is determined. If the confidence level is higher than the confidence level threshold, then the predicted performance index is added as the second pseudo-label of the sample luminescent material. If the predicted performance index output by the updated first model is the associated performance index, then the direct predicted value of the associated performance index is determined, and the direct predicted value is directly output by the updated first model; and the predicted performance index that meets the first condition is added as the second pseudo-label of the sample luminescent material, where the first condition is that the difference between the direct predicted value and the predicted performance index is less than the deviation threshold, the confidence level of the predicted performance index is higher than the confidence level threshold, and the predicted performance index is within the standard value range.

6. The method according to claim 3, characterized in that, The loss when updating the first model again includes: measured label loss, first pseudo-label loss, second pseudo-label loss and consistency loss, wherein the consistency loss is determined based on the association performance index directly predicted by the first model and the mapping association performance index in the second update.

7. The method according to claim 6, characterized in that, The weights of the measured label loss term, the first pseudo-label loss term, and the second pseudo-label loss term are all determined based on label attributes, which include measured labels and pseudo-labels. The weight of the consistency loss term is determined based on a preset value range and consistency error, wherein the consistency error is determined based on the correlation performance index directly predicted by the first model and the mapped correlation performance index when the model is updated again.

8. The method according to claim 2, characterized in that, The method further includes: Construct a missing mask matrix for the sample set, the missing mask matrix being used to label the missing labels in the sample set; During the iterative training process, gradient backpropagation at the missing labels is masked based on the missing mask matrix.

9. The method according to any one of claims 1 to 8, characterized in that, The relevant performance indicators include quantum yield and luminescence lifetime; The intermediate performance indicators include radiation attenuation intensity and non-radiation attenuation intensity.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method of any one of claims 1 to 9.

11. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method of any one of claims 1 to 9.

12. A computer program product comprising instructions that, when the instructions are executed, perform the method of any one of claims 1 to 9.