Molecular attribute prediction method and related equipment

By constructing a gaze trajectory prediction sub-model and an attribute prediction sub-model, the judgment process of experts when observing molecular structure diagrams is simulated, which solves the problems of insufficient data and lack of chemical theoretical knowledge in the existing technology, and realizes more accurate and reliable prediction of molecular properties of organic materials.

CN121054136AActive Publication Date: 2025-12-02JIHUA LAB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511570868.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2025-12-02
Estimated Expiration
2045-10-30

AI Technical Summary

Technical Problem

Existing artificial intelligence models face problems such as insufficient data and lack of understanding of chemical theory when predicting the molecular properties of organic materials, resulting in inaccurate predictions and poor robustness.

Method used

By acquiring the gaze trajectory and judgment result labels of technicians when judging the molecular properties of organic materials, an attribute prediction model is constructed that includes a gaze trajectory prediction sub-model and an attribute prediction sub-model. This model simulates the judgment process of experts when observing molecular structure diagrams and makes predictions by combining molecular structure diagrams and gaze trajectory prediction information.

Benefits of technology

It significantly improves the accuracy and robustness of predicting the molecular properties of organic materials, enabling a deeper understanding of the relationship between molecular structure and properties, and providing prediction results that are closer to expert judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121054136A_ABST
    Figure CN121054136A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of molecular attribute prediction, and discloses a molecular attribute prediction method and related equipment, and the method comprises the steps: obtaining a fixation point track and a corresponding judgment result label when a technician judges the molecular attribute of an organic material, and forming a training data set; training an attribute prediction model comprising a fixation point trajectory prediction sub-model and an attribute prediction sub-model by using the data set; wherein the fixation point trajectory prediction sub-model is responsible for simulating an attention mode of an expert on a key region of a molecular structure diagram, and the attribute prediction sub-model is used for predicting molecular attributes by combining prediction information of the molecular structure diagram and the fixation point trajectory prediction sub-model; utilizing the trained attribute prediction model to predict whether the organic material molecule to be detected has the target attribute or not; therefore, key features in the molecular structure can be identified more accurately, deep chemical reasoning is carried out, and the accuracy and robustness of molecular attribute prediction of the complex organic material are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of molecular property prediction technology, and more specifically, to a molecular property prediction method and related equipment. Background Technology

[0002] In the field of materials science, accurate prediction of the molecular properties of organic materials is a crucial step in the development of new materials. Currently, experts (i.e., experienced technicians) often demonstrate stronger chemical intuition and better robustness than existing artificial intelligence models when judging the complex properties of organic material molecules, such as luminescence efficiency in optoelectronic properties. This is because experts have accumulated rich theoretical knowledge and experience through long-term practice, enabling them to identify key structural regions by observing the molecular structure diagrams of organic materials and make relatively accurate judgments about the potential properties of organic material molecules.

[0003] Specifically, existing artificial intelligence (AI) models generally face the challenge of insufficient data when predicting the molecular properties of organic materials. The high cost of acquiring high-quality organic material molecular property data results in limited datasets for model training. Furthermore, existing AI models are typically built upon mathematical and statistical principles, lacking an intrinsic understanding and integration of chemical theoretical knowledge. This means that, unlike experts, these models struggle to perform deep chemical reasoning by combining local features and overall configurations of organic material molecular structures. Consequently, existing AI models often fail to make accurate predictions when faced with complex organic material molecular property prediction tasks, and may even produce results inconsistent with chemical theory, thus limiting their reliability and effectiveness in practical applications.

[0004] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention

[0005] The purpose of this application is to provide a method and related equipment for predicting molecular properties, which aims to solve the problems faced by existing artificial intelligence models in predicting the molecular properties of organic materials, such as insufficient data, lack of understanding of chemical theory, and difficulty in making accurate predictions.

[0006] In a first aspect, this application provides a molecular property prediction method for predicting whether a molecule of an organic material to be tested possesses a target property, characterized in that the method includes the following steps: A1. Obtain the gaze trajectory of technicians in the process of judging whether an organic material molecule has the target attribute by using the molecular structure diagram of the organic material, as well as the corresponding judgment result label, to form a training dataset; A2. Train an attribute prediction model using a training dataset; the attribute prediction model includes a gaze trajectory prediction sub-model and an attribute prediction sub-model. The gaze trajectory prediction sub-model is used to predict the gaze trajectory based on the molecular structure diagram of the organic material, and the attribute prediction sub-model is used to predict whether the organic material molecule has the target attribute based on the molecular structure diagram of the organic material and the gaze trajectory prediction information of the gaze trajectory prediction sub-model. A3. Using the trained attribute prediction model, predict whether the organic material molecule to be tested has the target attribute.

[0007] Preferably, step A1 includes: A101. Obtain the true situation of whether organic material molecules have the target properties, and use it as the judgment result label of the organic material molecule structure diagram corresponding to the organic material molecule; A102. Obtain the gaze trajectory and corresponding judgment results of multiple technicians in judging whether an organic material molecule has the target attribute based on the organic material molecule's molecular structure diagram; A103. Remove the gaze point trajectories whose judgment results do not match the corresponding judgment result labels; A104. For the remaining fixation point trajectories, remove fixation point trajectories whose number of deviations from the fixation point exceeds a preset threshold; the deviation from the fixation point refers to a fixation point that does not fall on the molecular structure diagram of the organic material. A105. Cluster the remaining gaze point trajectories to obtain multiple gaze point trajectory clusters, and remove gaze point trajectories from each gaze point trajectory cluster that deviate too much from the center trajectory of that gaze point trajectory cluster; A106. Combine the remaining gaze point trajectory and the corresponding judgment result label into a sample and add it to the training dataset; A107. Perform steps A101-A106 for various organic material molecules to obtain the final training dataset.

[0008] Preferably, step A104 includes: Obtain the envelope of the molecular structure diagram of the organic material; The envelope is expanded outward according to a preset expansion distance; The number of gaze points whose trajectories fall outside the area enclosed by the expanded envelope is counted to obtain the number of gaze points deviating from the gaze point. If the number of deviations from the gaze point exceeds a preset threshold, the gaze point trajectory is discarded.

[0009] Preferably, step A105 includes: The remaining gaze point trajectories are clustered to obtain multiple gaze point trajectory clusters; For each gaze point trajectory cluster, obtain the center trajectory of the gaze point trajectory cluster based on the gaze point trajectory of the gaze point trajectory cluster; For each gaze point trajectory cluster, calculate the deviation between each gaze point trajectory in that gaze point trajectory cluster and the central trajectory; Remove gaze points whose deviation exceeds a preset deviation threshold.

[0010] Preferably, step A2 includes: A201. Based on the molecular structure diagrams of organic materials and the corresponding gaze point trajectories in the training dataset, train the gaze point trajectory prediction sub-model so that the gaze point trajectory prediction sub-model can predict the gaze point trajectory based on the molecular structure diagrams of organic materials. A202. Freeze the gaze trajectory prediction sub-model; A203. Based on the organic material molecular structure diagram and corresponding judgment result labels in the training dataset, and combined with the gaze trajectory prediction sub-model, train the attribute prediction sub-model so that the attribute prediction sub-model can predict whether the organic material molecule has the target attribute based on the organic material molecular structure diagram and the gaze trajectory prediction information of the gaze trajectory prediction sub-model.

[0011] Preferably, step A203 includes: The molecular structure diagrams of organic materials in the training dataset are input into the gaze trajectory prediction sub-model, and the prediction information of the gaze trajectory by the gaze trajectory prediction sub-model is extracted; the prediction information includes the gaze trajectory predicted by the gaze trajectory prediction sub-model and / or the hidden layer parameters generated by the gaze trajectory prediction sub-model during the prediction process. The molecular structure diagrams of organic materials in the training dataset and the extracted prediction information are input into the attribute prediction sub-model to obtain the actual judgment result output by the attribute prediction sub-model. The loss function is calculated based on the actual judgment result and the corresponding judgment result label, and the model parameters of the attribute prediction sub-model are optimized based on the loss function.

[0012] Preferably, step A3 includes: A301. Obtain the molecular structure diagram of the organic material to be tested; A302. Input the molecular structure diagram into the gaze trajectory prediction sub-model and extract the gaze trajectory prediction information of the gaze trajectory prediction sub-model; A303. Input the molecular structure diagram and the prediction information into the property prediction sub-model to obtain the prediction result of whether the organic material molecule to be tested has the target property.

[0013] Secondly, this application provides a molecular property prediction system for predicting whether the molecules of an organic material to be tested have the target property. The system includes a display, an eye tracker, and a host computer. The display is used to display the molecular structure diagram of organic materials under the control of the host computer; The eye tracker is used to collect the gaze trajectory of technicians as they determine whether an organic material molecule has the target attribute by viewing the molecular structure diagram of the organic material displayed on the monitor, and upload it to the host computer. The host computer is used to execute: The system obtains the judgment results of technicians on whether the organic material molecules corresponding to the organic material molecular structures displayed on the monitor have the target properties, and combines them with the gaze point trajectory uploaded by the eye tracker and the corresponding organic material molecular structure diagram to form a training dataset. An attribute prediction model is trained using a training dataset. The attribute prediction model includes a gaze trajectory prediction sub-model and an attribute prediction sub-model. The gaze trajectory prediction sub-model is used to predict the gaze trajectory based on the molecular structure diagram of the organic material. The attribute prediction model is used to predict whether the organic material molecule has the target attribute based on the molecular structure diagram of the organic material and the gaze trajectory prediction information of the gaze trajectory prediction sub-model. The trained property prediction model is used to predict whether the organic material molecule to be tested has the target property.

[0014] Thirdly, this application provides an electronic device including a processor and a memory, the memory storing a computer program executable by the processor, wherein when the processor executes the computer program, it performs the steps in the molecular property prediction method described above.

[0015] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the molecular property prediction method as described above.

[0016] Beneficial Effects: This application provides a molecular property prediction method and related equipment. By acquiring the gaze trajectory and judgment result labels of technicians when judging the molecular properties of organic materials, a training dataset incorporating the experience of technicians is constructed. Based on this, a property prediction model is trained, including a gaze trajectory prediction sub-model and a property prediction sub-model. The gaze trajectory prediction sub-model simulates the observation pattern of technicians, predicting the gaze trajectory based on the molecular structure diagram; the property prediction sub-model combines the molecular structure diagram and gaze trajectory prediction information to predict whether the molecule possesses the target property. This method effectively solves the problem of inaccurate predictions by existing artificial intelligence models when facing the molecular property prediction of complex organic materials due to insufficient data and lack of chemical intuition. By integrating the expert knowledge and judgment process of technicians into the model training, the method of this application can gain a deeper understanding of the relationship between molecular structure and properties, thereby significantly improving the accuracy and robustness of molecular property prediction, overcoming the shortcomings of existing models in performing deep chemical reasoning, and providing a more reliable prediction tool for new material development. Attached Figure Description

[0017] Figure 1 A flowchart of a molecular property prediction method provided in this application.

[0018] Figure 2 A schematic diagram of a molecular property prediction system provided in this application.

[0019] Figure 3 A schematic diagram of the structure of the electronic device provided in this application.

[0020] Figure 4 This is a schematic diagram of the molecular structure of organic materials and the corresponding gaze point trajectory.

[0021] Figure 5 This is a schematic diagram of an attribute prediction model.

[0022] Labeling explanations: 1. Display; 2. Eye tracker; 3. Host computer; 301. Processor; 302. Memory; 303. Communication bus. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0024] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0025] Please refer to Figure 1 A molecular property prediction method in some embodiments of this application is used to predict whether a molecule of an organic material to be tested has a target property. The method is characterized by comprising the following steps: A1. Obtain the gaze trajectory of technicians in the process of judging whether an organic material molecule has the target attribute by using the molecular structure diagram of the organic material, as well as the corresponding judgment result label, to form a training dataset; A2. Train an attribute prediction model using a training dataset; the attribute prediction model includes a gaze trajectory prediction sub-model and an attribute prediction sub-model. The gaze trajectory prediction sub-model is used to predict the gaze trajectory based on the molecular structure diagram of the organic material, and the attribute prediction sub-model is used to predict whether the organic material molecule has the target attribute based on the molecular structure diagram of the organic material and the gaze trajectory prediction information of the gaze trajectory prediction sub-model. A3. Using the trained attribute prediction model, predict whether the organic material molecule to be tested has the target attribute.

[0026] This application incorporates the gaze trajectory information of technicians to simulate the thought process of experts when judging the molecular properties of organic materials, thereby effectively making up for the shortcomings of existing models in terms of chemical intuition and robustness, and significantly improving the accuracy and reliability of predicting the molecular properties of complex organic materials.

[0027] "Molecular structure diagram of organic materials" refers to a graphical representation of the atomic connections and spatial arrangement in organic material molecules. It is commonly used in the field of chemistry to help technicians intuitively understand the composition of molecules.

[0028] "Target property" refers to the specific properties of the organic material molecules to be predicted, such as luminescence, conductivity, and stability.

[0029] "Gaze trajectory" refers to the record of the path of a technician's eyes moving and lingering on an image when observing the molecular structure diagram of organic materials. It reflects the key areas of focus and the technician's thought process. For example... Figure 4 This diagram shows the gaze point trajectory formed on the molecular structure diagram of an organic material. In the diagram, the horizontal coordinate u represents the horizontal coordinate of a pixel, and the vertical coordinate v represents the vertical coordinate of a pixel.

[0030] "Judgment result label" refers to the true judgment result given by technicians on whether organic material molecules have the target attribute, which is usually a binary (e.g., "yes" or "no") or multi-class label.

[0031] The "attribute prediction model" is a machine learning model designed to predict whether an organic material possesses a target property based on its molecular structure diagram. The model consists of two sub-models: a gaze trajectory prediction sub-model and an attribute prediction sub-model.

[0032] The “Gaze Trajectory Prediction Sub-model” is part of the attribute prediction model. Its function is to predict the gaze trajectory that a technician may generate when observing the molecular structure diagram of an organic material, based on the input diagram.

[0033] The "attribute prediction sub-model" is another part of the attribute prediction model. Its function is to combine the prediction information provided by the organic material molecular structure diagram and the gaze point trajectory prediction sub-model to ultimately predict whether the organic material molecule has the target attribute.

[0034] The implementation environment of this application typically includes a data processing module, a model training module, and a prediction module. These modules can run on one or more computer devices and exchange data via a network.

[0035] This application proposes a molecular property prediction method, the core of which lies in simulating the judgment process of technicians, thereby improving the accuracy of prediction.

[0036] Specifically, this method first requires acquiring the gaze trajectory and corresponding judgment label of technicians as they judge whether an organic material molecule possesses the target attribute based on its molecular structure diagram. This data forms the training dataset. In practice, eye-tracking devices can be used to record the eye movements of technicians while observing the organic material molecular structure diagrams, thus obtaining the gaze trajectory. For example, several experienced chemistry experts can be invited to view a series of organic material molecular structure diagrams on a monitor and judge whether each molecule possesses a specific target attribute. During this process, the eye tracker records the experts' gaze trajectories in real time. Simultaneously, the actual attribute label for each organic material molecule needs to be acquired as the judgment result label. This data collectively constitutes the training dataset for subsequent model training.

[0037] Furthermore, the resulting training dataset is used to train an attribute prediction model. This model is designed to include a gaze trajectory prediction sub-model and an attribute prediction sub-model. The gaze trajectory prediction sub-model is used to predict gaze trajectories based on the organic material molecular structure diagram. For example, a convolutional neural network (CNN)-based image processing model can be constructed as the gaze trajectory prediction sub-model, which can extract visual features from the input organic material molecular structure diagram and output the predicted gaze trajectory. The attribute prediction sub-model is then used to predict whether the organic material molecule possesses the target attribute based on the organic material molecular structure diagram and the gaze trajectory prediction information from the gaze trajectory prediction sub-model. For example, a multilayer perceptron (MLP) or another CNN model can be constructed as the attribute prediction sub-model, which receives the features of the organic material molecular structure diagram and the predicted gaze trajectory information as input and outputs the final attribute prediction result. During training, these two sub-models can be jointly trained or trained in stages to optimize the overall model performance.

[0038] Finally, the trained attribute prediction model is used to predict whether the tested organic material molecule possesses the target attribute. When predicting a new organic material molecule, its molecular structure diagram is first obtained and then input into the trained gaze trajectory prediction sub-model to obtain prediction information. Next, the molecular structure diagram and the prediction information are input together into the attribute prediction sub-model to finally obtain the prediction result of whether the tested organic material molecule possesses the target attribute.

[0039] The molecular property prediction method proposed in this application effectively solves the problem of lack of chemical intuition and robustness in existing artificial intelligence models by simulating the judgment process of technicians.

[0040] Specifically, this application incorporates the experiential knowledge of human experts when acquiring the gaze trajectories and judgment result labels of technicians to form the training dataset. Traditional methods often rely solely on molecular structure data and attribute labels for model training, neglecting the cognitive cues provided by experts during the judgment process. This application integrates this valuable cognitive information into the training data by collecting the gaze trajectories of experts when observing molecular structure diagrams, enabling the model to learn the key regions and reasoning paths that experts focus on.

[0041] When training the attribute prediction model using the training dataset, this application decomposes the attribute prediction model into a gaze trajectory prediction sub-model and an attribute prediction sub-model. The gaze trajectory prediction sub-model first predicts the gaze trajectory based on the molecular structure diagram of the organic material, which is equivalent to simulating the process by which an expert first identifies key structural regions when observing a molecular structure diagram. Subsequently, the attribute prediction sub-model combines the organic material molecular structure diagram and the prediction information to predict whether the organic material molecule has the target attribute. This staged prediction mechanism enables the model to not only extract features from the molecular structure diagram but also to use simulated expert attention information for deeper chemical reasoning. Compared with traditional models that directly predict attributes from molecular structure diagrams, the model in this application can better understand the intrinsic relationship between molecular structure and attributes, thereby improving the accuracy and robustness of the prediction.

[0042] Therefore, when using a trained attribute prediction model to predict whether a molecule of an organic material possesses a target attribute, the model in this application can provide prediction results that are closer to expert judgment. This method not only improves the accuracy of the prediction but also provides interpretability for the prediction results, because the predicted gaze trajectory can indicate the molecular structural region that the model focuses on when making its judgment. This has important guiding significance for the research and optimization of new materials.

[0043] The molecular property prediction method proposed in this application has significant advantages and innovations compared to existing technologies. Traditional molecular property prediction methods typically input molecular structure diagrams directly into deep learning models for end-to-end prediction. While this method can achieve property prediction to some extent, it often exhibits limitations when faced with complex properties due to a lack of intrinsic understanding of chemical theory and the integration of expert experience. This is especially true when the amount of data is insufficient, resulting in poor generalization ability and robustness of the model.

[0044] The core innovation of this application lies in incorporating the gaze trajectory information of technicians and constructing a two-stage prediction architecture that includes a gaze trajectory prediction sub-model and an attribute prediction sub-model. By acquiring the gaze trajectory of technicians when judging molecular attributes, this application integrates the cognitive processes and attention mechanisms of human experts into the model training. This enables the gaze trajectory prediction sub-model to learn and simulate the key areas that experts focus on when observing molecular structure diagrams, thereby providing more instructive feature information for the attribute prediction sub-model.

[0045] For example, when predicting the luminescence efficiency of organic material molecules, traditional models may learn solely from global features, while the model in this application can focus on specific functional groups or conjugated systems in the molecule that are closely related to luminescence performance through predicted gaze trajectories. This focus on key regions enables the property prediction sub-model to perform more accurate chemical reasoning, thereby significantly improving prediction accuracy. Furthermore, this method provides interpretability for the prediction results, as the predicted gaze trajectories can visually demonstrate the molecular structural regions upon which the model bases its judgments. This is of great value for chemists to understand the model's decision-making process and to design new materials. Therefore, this application not only improves prediction performance but also enhances the model's credibility and practicality, providing a more intelligent and efficient solution for predicting the properties of organic material molecules.

[0046] In some implementations, step A1 includes: A101. Obtain the true situation of whether organic material molecules have the target properties, and use it as the judgment result label of the organic material molecule structure diagram corresponding to the organic material molecule; A102. Obtain the gaze trajectory and corresponding judgment results of multiple technicians in judging whether an organic material molecule has the target attribute based on the organic material molecule's molecular structure diagram; A103. Remove the gaze point trajectories whose judgment results do not match the corresponding judgment result labels; A104. For the remaining fixation point trajectories, remove fixation point trajectories whose number of deviations from the fixation point exceeds a preset threshold; the deviation from the fixation point refers to a fixation point that does not fall on the molecular structure diagram of the organic material. A105. Cluster the remaining gaze point trajectories to obtain multiple gaze point trajectory clusters, and remove gaze point trajectories from each gaze point trajectory cluster that deviate too much from the center trajectory of that gaze point trajectory cluster; A106. Combine the remaining gaze point trajectory and the corresponding judgment result label into a sample and add it to the training dataset; A107. Perform steps A101-A106 for various organic material molecules to obtain the final training dataset.

[0047] Specifically, step A101 aims to obtain the true state of whether the organic material molecule possesses the target property. This is typically obtained through experimental verification or querying authoritative databases, serving as a benchmark for subsequent data cleaning and model training. This true state is used as a label for the judgment result corresponding to the molecular structure diagram of the organic material.

[0048] In step A102, eye-tracking devices are used to acquire the gaze trajectories of multiple technicians as they observe and judge the molecular structure diagrams of organic materials, along with their individual judgment results. This step aims to collect multi-source, multi-angle human judgment data.

[0049] Step A103 is used to perform the first layer of filtering on the initially acquired data. Specifically, if the technician's judgment result is inconsistent with the actual situation obtained in step A101 (i.e., the judgment result label), it is considered that the technician's gaze trajectory may be misleading or an error in judgment, and therefore the inconsistent gaze trajectory is removed. The purpose is to ensure that the gaze trajectories in the training dataset correspond to the correct judgment result labels, thereby improving the reliability of the data.

[0050] Furthermore, step A104 aims to eliminate trajectories containing a large amount of noise or invalid gaze points. A gaze point deviating from the molecular structure diagram of an organic material refers to a gaze point that does not fall on the diagram, such as when a technician briefly shifts their gaze away from the molecular structure diagram during observation. By counting the number of deviating gaze points in the gaze point trajectory and comparing it with a preset threshold, low-quality gaze point trajectories resulting from inattention or misoperation can be effectively identified and eliminated. The purpose is to remove gaze point information irrelevant to the molecular structure diagram itself, allowing the remaining gaze point trajectories to focus more on the key regions of the molecular structure.

[0051] Building upon this, step A105 performs clustering on the initially screened gaze points trajectories. Clustering groups trajectories with similar gaze patterns into a single category, forming multiple gaze point trajectory clusters. For each gaze point trajectory cluster, its central trajectory is calculated; this central trajectory can be understood as the representative pattern of the gaze point trajectories within that cluster. Subsequently, gaze point trajectories that deviate excessively from this central trajectory are removed from the cluster. The purpose of this step is to further remove outliers or isolated trajectories within the cluster, ensuring high consistency and representativeness of the trajectories within each gaze point trajectory cluster, thereby extracting more common expert gaze patterns.

[0052] In step A106, the remaining gaze point trajectories after the above multiple screening and optimization processes are combined with the corresponding judgment result labels to form a high-quality sample, which is then added to the training dataset. Finally, step A107 ensures the richness and diversity of the training dataset. By repeatedly performing steps A101 to A106 for various different organic material molecules, a comprehensive training dataset containing judgment results for multiple molecule types and properties can be constructed.

[0053] The proposed solution effectively addresses the potential noise, inconsistencies, and low quality issues of the original data through meticulous management of the training dataset construction process. Specifically, step A103 compares the judgments of technicians with actual conditions, eliminating erroneous gaze trajectories and preventing the model from learning from incorrect human judgments. Step A104 identifies and removes numerous gaze points that deviate from the molecular structure diagram, ensuring the effectiveness of the gaze trajectories and enabling the model to focus more on the key regions of the molecular structure itself. Furthermore, step A105 further refines the commonalities of expert gaze patterns through clustering and outlier removal within clusters, reducing biases caused by individual differences or chance factors. This allows the training dataset to more accurately reflect the core concerns and decision-making paths of technicians when judging molecular properties. It is precisely these data cleaning and optimization steps that result in a higher quality and representativeness of the training dataset.

[0054] Through the above technical solution, this application can significantly improve the quality and reliability of the training dataset. By employing multiple screening and optimization processes, inaccurate, inconsistent, or noisy gaze trajectories are eliminated, enabling the training dataset to more accurately reflect the true and effective gaze patterns and decision-making logic of technicians when judging the molecular properties of organic materials. Consequently, the property prediction model trained using this high-quality training dataset will exhibit significantly improved learning efficiency and prediction accuracy, as well as enhanced generalization ability, thus enabling more reliable prediction of whether the tested organic material molecule possesses the target property.

[0055] In some possible implementations, step A104 may include: Obtain the envelope of the molecular structure diagram of the organic material; The envelope is expanded outward according to a preset expansion distance; The number of gaze points whose trajectories fall outside the area enclosed by the expanded envelope is counted to obtain the number of gaze points deviating from the gaze point. If the number of deviations from the gaze point exceeds a preset threshold, the gaze point trajectory is discarded.

[0056] Specifically, the envelope of the organic material molecular structure diagram can be understood as the smallest convex polygon or smallest rectangular region that can completely enclose the molecular structure diagram. Its purpose is to provide a clear boundary for subsequent expansion operations.

[0057] Furthermore, the outward expansion of the envelope refers to the uniform expansion of the envelope in all directions by a preset expansion distance, forming a region larger than the original envelope. For example, the expansion distance can be set according to the accuracy of the eye tracker, the screen resolution, or empirical values. Its purpose is to provide a certain margin of error for the fixation point, so as to avoid misjudging deviations from the fixation point due to minor deviations.

[0058] The counting of gaze points falling outside the area enclosed by the expanded envelope involves determining the position of each gaze point in the gaze point trajectory. If its coordinates are outside the area enclosed by the expanded envelope (including the envelope itself), it is counted as a gaze point deviating from the gaze point. This yields the total number of gaze points deviating from the gaze point trajectory. In practical applications, if the number of gaze points deviating from the gaze point exceeds a preset threshold (which can be set according to actual needs), the gaze point trajectory is considered to contain excessive noise or invalid information and should be removed from the training dataset. The preset threshold can be adjusted according to actual needs and data characteristics; for example, it can be set as a percentage of the total number of gaze point trajectories, with the aim of ensuring the quality of the training dataset.

[0059] This application's solution constructs an effective gaze region with a certain tolerance by introducing and expanding the envelope of the molecular structure diagram of organic materials. When the gaze trajectory of a technician is collected, even if there are slight eye movement errors or the gaze point slightly exceeds the actual boundary of the molecular structure diagram, it will not be immediately judged as a deviation from the gaze point as long as it still falls within the area enclosed by the expanded envelope. Only when the gaze point significantly deviates from the expanded region will it be accurately identified as a deviation from the gaze point. Therefore, by counting the number of gaze points falling outside the expanded region and comparing them with a preset threshold, gaze trajectories that truly contain a large amount of noise or invalid information can be identified and eliminated more accurately and robustly. This avoids the mistaken deletion of valid data due to overly strict boundary judgments or the retention of too much noisy data due to overly lenient boundary judgments.

[0060] The above technical solution enables more accurate identification and removal of gaze point trajectories containing excessive noise in the training dataset. Compared to simply determining whether a gaze point falls on the molecular structure map, using the expanded envelope as the criterion provides a reasonable margin of error for the gaze point, effectively reducing the misjudgment rate caused by limitations in eye tracker accuracy or differences in the gaze habits of technicians. This significantly improves the quality of the training dataset, thereby helping to train a more stable and accurate property prediction model and enhancing the reliability of the model in predicting the molecular properties of organic materials.

[0061] Further, step A105 may include: The remaining gaze point trajectories are clustered to obtain multiple gaze point trajectory clusters; For each gaze point trajectory cluster, obtain the center trajectory of the gaze point trajectory cluster based on the gaze point trajectory of the gaze point trajectory cluster; For each gaze point trajectory cluster, calculate the deviation between each gaze point trajectory in that gaze point trajectory cluster and the central trajectory; Remove gaze points whose deviation exceeds a preset deviation threshold.

[0062] Clustering the remaining gaze points can be understood as grouping gaze points with similar features or patterns. For example, algorithms such as K-means, DBSCAN, and hierarchical clustering can be used to cluster gaze points. The purpose is to group trajectories with similar areas of interest or scanning paths formed by experts during the judgment process into one category, so as to facilitate subsequent refined screening.

[0063] Furthermore, for each gaze point trajectory cluster, a central trajectory for that cluster is obtained based on the gaze point trajectories within that cluster. This central trajectory can be considered a representative trajectory of the cluster, reflecting the common patterns or average behavior of the gaze point trajectories within the cluster. For example, the central trajectory can be obtained by calculating the average trajectory, median trajectory, or by selecting the trajectory within the cluster that has the smallest sum of distances to all other trajectories. The purpose is to provide a benchmark for subsequent deviation calculations.

[0064] In practical applications, for each gaze point trajectory cluster, the deviation between each gaze point trajectory in that cluster and the central trajectory is calculated. This deviation is a quantitative indicator measuring the degree of difference between an individual gaze point trajectory within a cluster and the cluster's central trajectory. For example, the deviation can be calculated using methods such as Dynamic Time Warping (DTW), Euclidean distance, and Fréchet distance. The purpose is to quantify the degree of conformity between each trajectory and the representative pattern of the cluster.

[0065] Therefore, gaze point trajectories with a deviation greater than a preset deviation threshold are removed. The preset deviation threshold is a pre-defined value used to define whether a gaze point trajectory is considered an outlier or noise within the cluster. When the deviation of a gaze point trajectory exceeds this threshold, it is considered that the trajectory differs too significantly from the overall pattern of the cluster and should be removed. This threshold can be determined based on experience, statistical analysis, or through cross-validation. The purpose is to remove unrepresentative outlier trajectories within the cluster, further purifying the training dataset.

[0066] This application's solution employs a refined screening of clustered gaze point trajectories by introducing a center trajectory and deviation calculation. Specifically, after classifying similar gaze point trajectories into different gaze point trajectory clusters, a representative center trajectory is determined for each cluster, providing a clear reference for subsequent quality assessment. Subsequently, by quantifying the deviation between each gaze point trajectory within a cluster and the center trajectory, the "normality" of each trajectory can be objectively evaluated. It is precisely this quantified deviation assessment that enables the accurate identification and removal of abnormal trajectories that significantly deviate from the mainstream pattern within the cluster, based on a preset deviation threshold. This mechanism effectively avoids errors that may result from subjective judgment or rough screening, ensuring a high degree of consistency and representativeness within each gaze point trajectory cluster.

[0067] Through the above technical solution, this application can perform more refined and accurate screening of gaze points in the training dataset. By calculating the deviation between the gaze point trajectory and the center trajectory and setting a threshold, abnormal trajectories with excessive deviation within the cluster can be effectively identified and eliminated, thereby significantly improving the purity and quality of the training dataset. This helps ensure that the attribute prediction model learns more representative and reliable expert gaze patterns during training, thereby improving the accuracy and robustness of the model in predicting the molecular properties of organic materials and avoiding prediction bias caused by noise or outliers in the training data.

[0068] Specifically, step A2 includes: A201. Based on the molecular structure diagrams of organic materials and the corresponding gaze point trajectories in the training dataset, train the gaze point trajectory prediction sub-model so that the gaze point trajectory prediction sub-model can predict the gaze point trajectory based on the molecular structure diagrams of organic materials. A202. Freeze the gaze trajectory prediction sub-model; A203. Based on the organic material molecular structure diagram and corresponding judgment result labels in the training dataset, and combined with the gaze trajectory prediction sub-model, train the attribute prediction sub-model so that the attribute prediction sub-model can predict whether the organic material molecule has the target attribute based on the organic material molecular structure diagram and the gaze trajectory prediction information of the gaze trajectory prediction sub-model.

[0069] Specifically, step A201 aims to first independently train the gaze trajectory prediction sub-model. In this stage, the sub-model is input into the molecular structure diagrams of organic materials in the training dataset and learned using the corresponding real gaze trajectories as supervision signals. The goal is to enable the gaze trajectory prediction sub-model to accurately extract and predict the technician's gaze trajectory from the molecular structure diagrams.

[0070] Step A202 refers to fixing the model parameters of the gaze trajectory prediction sub-model after it has completed training and achieved the expected prediction performance, and no longer allowing it to be updated in subsequent training processes. This ensures that the gaze trajectory prediction sub-model provides stable and high-quality gaze trajectory prediction information during the training phase of the attribute prediction sub-model.

[0071] In practical applications, step A203 is performed after the gaze trajectory prediction sub-model is frozen. At this stage, the attribute prediction sub-model is trained, with inputs including the molecular structure diagrams of organic materials from the training dataset and gaze trajectory prediction information provided by the frozen gaze trajectory prediction sub-model. The attribute prediction sub-model uses the corresponding judgment result labels as supervision signals to learn how to combine the molecular structure diagrams and gaze trajectory prediction information to accurately predict whether organic material molecules possess the target attribute.

[0072] The proposed solution effectively addresses the aforementioned training stability issue by decomposing the attribute prediction model training process into two sequential stages and introducing a mechanism to freeze the gaze trajectory prediction sub-model. First, the gaze trajectory prediction sub-model is trained independently, allowing it to focus on learning the ability to accurately predict gaze trajectories from the molecular structure diagram of organic materials. This training stage ensures that the gaze trajectory prediction sub-model can provide high-quality intermediate features or prediction results. Subsequently, during the training of the attribute prediction sub-model, the gaze trajectory prediction sub-model is frozen, and its parameters are no longer updated. Thus, the attribute prediction sub-model can be trained on a stable upstream model that has already learned effective gaze information, thereby avoiding potential mutual interference and optimization difficulties in joint training. The attribute prediction sub-model can then focus more on learning how to combine molecular structure diagram information with stable gaze trajectory prediction information to accurately determine the final attribute.

[0073] Through the above technical solution, the training process of the attribute prediction model is optimized, significantly improving the stability and efficiency of training. Since the gaze trajectory prediction sub-model is fully trained and frozen before the attribute prediction sub-model is trained, the attribute prediction sub-model can learn on a more reliable feature base, thus avoiding training fluctuations caused by the unstable output of the gaze trajectory prediction sub-model. This allows the attribute prediction sub-model to converge more effectively, ultimately improving the accuracy and robustness of the entire attribute prediction model in predicting the target properties of organic material molecules.

[0074] In some preferred embodiments, this application is implemented as follows: As a specific implementation, in step A201, the gaze trajectory prediction sub-model can be constructed as a sequence-to-sequence model based on a convolutional neural network (CNN) and a recurrent neural network (RNN). The CNN part is used to extract spatial features from the molecular structure diagram of organic materials, while the RNN part (e.g., a long short-term memory network LSTM or a gated recurrent unit GRU) is used to predict the gaze trajectory sequence based on the extracted feature sequence. During training, mean squared error (MSE) can be used as the loss function to minimize the difference between the predicted gaze trajectory and the true gaze trajectory. Specifically, in step A202, after the gaze trajectory prediction sub-model is trained, all its trainable parameters (e.g., convolutional kernel weights, RNN unit weights, etc.) are set to a non-updateable state. This can be achieved by setting the `requires_grad` attribute of the corresponding layer to `False` in a deep learning framework (such as TensorFlow or PyTorch). Further, in step A203, the attribute prediction sub-model can be designed as a multimodal fusion network. The network receives two inputs: the original molecular structure diagram of the organic material and prediction information from a frozen gaze trajectory prediction sub-model. The prediction information may include the predicted gaze trajectory sequence itself, or hidden layer parameters (e.g., the final hidden state of the RNN or attention weights) generated by the gaze trajectory prediction sub-model during the prediction process. The attribute prediction sub-model may consist of a separate CNN branch processing the molecular structure diagram and a separate network branch processing the gaze trajectory prediction information. The output features of these two branches are then concatenated or fused via an attention mechanism and input into a fully connected layer for binary classification (e.g., having the target attribute or not having the target attribute). During training, a cross-entropy loss function can be used, and a gradient descent optimizer (such as Adam) can be employed to update the parameters of the attribute prediction sub-model.

[0075] Preferably, step A203 may include: The molecular structure diagrams of organic materials in the training dataset are input into the gaze trajectory prediction sub-model, and the prediction information of the gaze trajectory by the gaze trajectory prediction sub-model is extracted; the prediction information includes the gaze trajectory predicted by the gaze trajectory prediction sub-model and / or the hidden layer parameters generated by the gaze trajectory prediction sub-model during the prediction process. The molecular structure diagrams of organic materials in the training dataset and the extracted prediction information are input into the attribute prediction sub-model to obtain the actual judgment result output by the attribute prediction sub-model. The loss function is calculated based on the actual judgment result and the corresponding judgment result label, and the model parameters of the attribute prediction sub-model are optimized based on the loss function.

[0076] Specifically, the predicted information can be understood as any information generated by the gaze trajectory prediction sub-model that is helpful for property prediction when processing molecular structure diagrams of organic materials. For example, it may include the gaze trajectory predicted by the gaze trajectory prediction sub-model based on the molecular structure diagram of the organic material. This trajectory can be represented in the form of a heatmap, coordinate sequence, or probability distribution, indicating the regions or paths that human experts might focus on when judging molecular properties. Furthermore, the predicted information may also include hidden layer parameters generated by the gaze trajectory prediction sub-model during the prediction process. These hidden layer parameters are typically intermediate representations of the input features abstracted and encoded within the model. They may contain richer and deeper semantic information than the final predicted trajectory, reflecting the model's focus on the molecular structure diagram and the feature extraction process.

[0077] In practical applications, the molecular structure diagrams of organic materials and the extracted prediction information from the training dataset are input into the property prediction sub-model to provide it with more comprehensive input. After receiving these inputs, the property prediction sub-model outputs a practical judgment result regarding whether the organic material molecule possesses the target property, based on its internal logic and parameters. This practical judgment result is typically a classification label.

[0078] Furthermore, to optimize the performance of the attribute prediction sub-model, a loss function needs to be calculated based on the actual judgment result and the corresponding judgment result label. The loss function quantifies the difference or error between the model's prediction result and the true label. For example, for binary classification problems, a binary cross-entropy loss function can be used; for multi-class classification problems, a cross-entropy loss function can be used. Based on the loss function, the model parameters of the attribute prediction sub-model are iteratively adjusted and optimized using a backpropagation algorithm and an optimizer (e.g., gradient descent, Adam, SGD, etc.) to minimize the loss function value, thereby making the prediction result of the attribute prediction sub-model closer to the true judgment result label.

[0079] The solution in this application explicitly extracts the prediction information from the gaze trajectory prediction sub-model and combines it with the original organic material molecular structure. Figure 1 As input to the attribute prediction sub-model, the sub-model can fully utilize the "attention" or "reasoning path" captured by the gaze trajectory prediction sub-model when human experts judge molecular attributes during the learning process. This mechanism effectively transfers the experiential knowledge of human experts to the attribute prediction sub-model in a computable form, thereby bridging the semantic gap that may exist when predicting attributes solely based on molecular structure diagrams. By calculating the loss function between the actual judgment result and the true label, and optimizing the model parameters of the attribute prediction sub-model accordingly, it is ensured that the attribute prediction sub-model can continuously learn and improve, thereby enhancing its predictive ability and making its internal decision-making mechanism closer to the judgment logic of human experts.

[0080] Through the aforementioned technical solutions, the attribute prediction sub-model can not only learn features from the molecular structure diagram of organic materials itself, but also integrate prediction information provided by the gaze trajectory prediction sub-model, which simulates the focus of human experts. This significantly enhances the learning efficiency and prediction accuracy of the attribute prediction sub-model. This combination allows the model to better understand key regions or features in the molecular structure related to the target attribute, thereby improving the reliability and interpretability of the prediction results. Furthermore, by iteratively optimizing the model parameters based on the loss function, the attribute prediction sub-model can continuously adapt and learn, further improving its generalization ability and robustness. This enables it to make more accurate attribute predictions that better align with expert judgment when faced with new organic material molecules.

[0081] Specifically, step A3 includes: A301. Obtain the molecular structure diagram of the organic material to be tested; A302. Input the molecular structure diagram into the gaze trajectory prediction sub-model and extract the gaze trajectory prediction information of the gaze trajectory prediction sub-model; A303. Input the molecular structure diagram and the prediction information into the property prediction sub-model to obtain the prediction result of whether the organic material molecule to be tested has the target property.

[0082] Specifically, in step A301, the molecular structure diagram of the organic material to be tested can be obtained in various forms, such as by searching a chemical structure database, manual input by the user, or extraction from literature using image recognition technology. This molecular structure diagram is the basic input information for subsequent predictions.

[0083] In step A302, after the molecular structure map is input into the gaze trajectory prediction sub-model, the sub-model generates a series of data representing potential gaze regions or paths, i.e., prediction information, based on the knowledge it has acquired during training. The prediction information may include the predicted gaze trajectory itself, or it may be the hidden layer parameters generated by the gaze trajectory prediction sub-model during the prediction process. These parameters can reflect the sub-model's deep understanding of the molecular structure map and the results of feature extraction.

[0084] Further, in step A303, the attribute prediction sub-model receives the raw information from the molecular structure map and the predicted information obtained from the gaze trajectory prediction sub-model. By combining these two types of information, the attribute prediction sub-model can more comprehensively and accurately understand the features of the molecular structure map and its correlation with the target attribute. Finally, the attribute prediction sub-model outputs a judgment result, such as a binary classification result (yes / no target attribute), indicating the probability that the tested organic material molecule possesses the target attribute.

[0085] The proposed solution first inputs the molecular structure diagram of the organic material to be tested into a gaze trajectory prediction sub-model to simulate the visual attention points of human experts during the judgment process, thereby extracting predictive information related to human cognitive processes. Subsequently, the original information from the molecular structure diagram and the extracted predictive information are jointly input into an attribute prediction sub-model. This dual-input mechanism allows the attribute prediction sub-model to utilize not only the features of the molecular structure diagram itself but also the simulated human expert gaze points, thus incorporating the "interpretive" or "attentional" mechanisms of expert experience into the prediction process. Therefore, the attribute prediction model can more comprehensively and deeply understand the complex relationship between molecular structure and target attributes, improving the accuracy and interpretability of predictions.

[0086] The above technical solution fully utilizes the trained gaze trajectory prediction sub-model and attribute prediction sub-model when predicting whether a molecule of an organic material possesses a target attribute. Specifically, by acquiring the molecular structure diagram of the molecule and inputting it into the gaze trajectory prediction sub-model to obtain prediction information, and then inputting both the molecular structure diagram and the prediction information into the attribute prediction sub-model, accurate prediction of molecular attributes can be achieved. This method not only improves the accuracy of prediction but also, by introducing the gaze trajectory prediction sub-model, makes the prediction process more interpretable, helping technicians understand the basis for the model's judgments, thereby enhancing the model's practicality and reliability.

[0087] refer to Figure 2 This application provides a molecular property prediction system for predicting whether the molecules of an organic material to be tested have the target property. The system includes a display 1, an eye tracker 2, and a host computer 3. The display 1 is used to display the molecular structure diagram of organic materials under the control of the host computer 3; The eye tracker 2 is used to collect the gaze trajectory of the technician in the process of judging whether the organic material molecule has the target attribute by means of the organic material molecular structure diagram displayed on the display 1, and upload it to the host computer 3; The host computer 3 is used to execute: The technician's judgment on whether the organic material molecules corresponding to the organic material molecular structure displayed on the monitor 1 have the target properties is obtained. Combined with the gaze point trajectory uploaded by the eye tracker 2 and the corresponding organic material molecular structure diagram, a training dataset is formed (the specific process can be referred to step A1 above). An attribute prediction model is trained using a training dataset. The attribute prediction model includes a gaze trajectory prediction sub-model and an attribute prediction sub-model. The gaze trajectory prediction sub-model is used to predict the gaze trajectory based on the molecular structure diagram of the organic material. The attribute prediction sub-model is used to predict whether the organic material molecule has the target attribute based on the molecular structure diagram of the organic material and the gaze trajectory prediction information of the gaze trajectory prediction sub-model (for details, please refer to step A2 above). Using the trained property prediction model, predict whether the organic material molecule to be tested has the target property (for details, refer to step A3 above).

[0088] Display 1 can be any display device capable of clearly displaying images, such as a liquid crystal display (LCD), an organic light-emitting diode display (OLED), or a projector. Its main function is to provide technicians with a visual interface to observe and analyze the molecular structure diagrams of organic materials.

[0089] The eye tracker 2 can be an eye-tracking device based on the principle of infrared light reflection or an eye-tracking device based on video image processing. It is placed at a suitable position when the technician is observing the monitor 1 to accurately capture the technician's eye movement data.

[0090] Technicians can directly input their judgment results (such as "yes" or "no") on the host computer 3 via a keyboard, mouse, or other input devices, or convert verbal judgments into text data through a voice recognition system.

[0091] Please refer to Figure 3This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device includes a processor 301 and a memory 302. The processor 301 and the memory 302 are interconnected and communicate with each other via a communication bus 303 and / or other connection mechanisms (not shown). The memory 302 stores a computer program executable by the processor 301. When the electronic device is running, the processor 301 executes the computer program to perform the molecular property prediction method in any optional implementation of the above embodiments, to achieve the following functions: acquiring the gaze trajectory and corresponding judgment result labels of a technician judging whether an organic material molecule has a target property based on an organic material molecular structure diagram, forming a training dataset; training an attribute prediction model using the training dataset; the attribute prediction model includes a gaze trajectory prediction sub-model and an attribute prediction sub-model. The gaze trajectory prediction sub-model is used to predict the gaze trajectory based on the organic material molecular structure diagram, and the attribute prediction sub-model is used to predict whether an organic material molecule has the target property based on the organic material molecular structure diagram and the gaze trajectory prediction information of the gaze trajectory; and using the trained attribute prediction model, predicting whether the tested organic material molecule has the target property.

[0092] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it performs the molecular property prediction method in any optional implementation of the above embodiments to achieve the following functions: acquiring the gaze trajectory and corresponding judgment result labels of the process by which a technician judges whether an organic material molecule has a target property based on an organic material molecular structure diagram, forming a training dataset; training an attribute prediction model using the training dataset; the attribute prediction model includes a gaze trajectory prediction sub-model and an attribute prediction sub-model, wherein the gaze trajectory prediction sub-model is used to predict the gaze trajectory based on the organic material molecular structure diagram, and the attribute prediction sub-model is used to predict whether the organic material molecule has the target property based on the organic material molecular structure diagram and the gaze trajectory prediction information of the gaze trajectory prediction sub-model; and using the trained attribute prediction model, predicting whether the organic material molecule to be tested has the target property. The computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0093] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A molecular property prediction method for predicting whether a molecule of an organic material to be tested possesses a target property, characterized in that, The method includes the following steps: A1. Obtain the gaze trajectory of technicians in the process of judging whether an organic material molecule has the target attribute by using the molecular structure diagram of the organic material, as well as the corresponding judgment result label, to form a training dataset; A2. Train an attribute prediction model using a training dataset; the attribute prediction model includes a gaze trajectory prediction sub-model and an attribute prediction sub-model. The gaze trajectory prediction sub-model is used to predict the gaze trajectory based on the molecular structure diagram of the organic material, and the attribute prediction sub-model is used to predict whether the organic material molecule has the target attribute based on the molecular structure diagram of the organic material and the gaze trajectory prediction information of the gaze trajectory prediction sub-model. A3. Using the trained attribute prediction model, predict whether the organic material molecule to be tested has the target attribute.

2. The molecular property prediction method according to claim 1, characterized in that, Step A1 includes: A101. Obtain the true situation of whether organic material molecules have the target properties, and use it as the judgment result label of the organic material molecule structure diagram corresponding to the organic material molecule; A102. Obtain the gaze trajectory and corresponding judgment results of multiple technicians in judging whether an organic material molecule has the target attribute based on the organic material molecule's molecular structure diagram; A103. Remove the gaze point trajectories whose judgment results do not match the corresponding judgment result labels; A104. For the remaining fixation point trajectories, remove fixation point trajectories whose number of deviations from the fixation point exceeds a preset threshold; the deviation from the fixation point refers to a fixation point that does not fall on the molecular structure diagram of the organic material. A105. Cluster the remaining gaze point trajectories to obtain multiple gaze point trajectory clusters, and remove gaze point trajectories from each gaze point trajectory cluster that deviate too much from the center trajectory of that gaze point trajectory cluster; A106. Combine the remaining gaze point trajectory and the corresponding judgment result label into a sample and add it to the training dataset; A107. Perform steps A101-A106 for various organic material molecules to obtain the final training dataset.

3. The molecular property prediction method according to claim 2, characterized in that, Step A104 includes: Obtain the envelope of the molecular structure diagram of the organic material; The envelope is expanded outward according to a preset expansion distance; The number of gaze points whose trajectories fall outside the area enclosed by the expanded envelope is counted to obtain the number of gaze points deviating from the gaze point. If the number of deviations from the gaze point exceeds a preset threshold, the gaze point trajectory is discarded.

4. The molecular property prediction method according to claim 2, characterized in that, Step A105 includes: The remaining gaze point trajectories are clustered to obtain multiple gaze point trajectory clusters; For each gaze point trajectory cluster, obtain the center trajectory of the gaze point trajectory cluster based on the gaze point trajectory of the gaze point trajectory cluster; For each gaze point trajectory cluster, calculate the deviation between each gaze point trajectory in that gaze point trajectory cluster and the central trajectory; Remove gaze points whose deviation exceeds a preset deviation threshold.

5. The molecular property prediction method according to claim 1, characterized in that, Step A2 includes: A201. Based on the molecular structure diagrams of organic materials and the corresponding gaze point trajectories in the training dataset, train the gaze point trajectory prediction sub-model so that the gaze point trajectory prediction sub-model can predict the gaze point trajectory based on the molecular structure diagrams of organic materials. A202. Freeze the gaze trajectory prediction sub-model; A203. Based on the organic material molecular structure diagram and corresponding judgment result labels in the training dataset, and combined with the gaze trajectory prediction sub-model, train the attribute prediction sub-model so that the attribute prediction sub-model can predict whether the organic material molecule has the target attribute based on the organic material molecular structure diagram and the gaze trajectory prediction information of the gaze trajectory prediction sub-model.

6. The molecular property prediction method according to claim 5, characterized in that, Step A203 includes: The molecular structure diagrams of organic materials in the training dataset are input into the gaze trajectory prediction sub-model, and the prediction information of the gaze trajectory by the gaze trajectory prediction sub-model is extracted; the prediction information includes the gaze trajectory predicted by the gaze trajectory prediction sub-model and / or the hidden layer parameters generated by the gaze trajectory prediction sub-model during the prediction process. The molecular structure diagrams of organic materials in the training dataset and the extracted prediction information are input into the attribute prediction sub-model to obtain the actual judgment result output by the attribute prediction sub-model. The loss function is calculated based on the actual judgment result and the corresponding judgment result label, and the model parameters of the attribute prediction sub-model are optimized based on the loss function.

7. The molecular property prediction method according to claim 1, characterized in that, Step A3 includes: A301. Obtain the molecular structure diagram of the organic material to be tested; A302. Input the molecular structure diagram into the gaze trajectory prediction sub-model and extract the gaze trajectory prediction information of the gaze trajectory prediction sub-model; A303. Input the molecular structure diagram and the prediction information into the property prediction sub-model to obtain the prediction result of whether the organic material molecule to be tested has the target property.

8. A molecular property prediction system for predicting whether a molecule of an organic material to be tested possesses a target property, characterized in that, The system includes a display, an eye tracker, and a host computer; The display is used to display the molecular structure diagram of organic materials under the control of the host computer; The eye tracker is used to collect the gaze trajectory of technicians as they determine whether an organic material molecule has the target attribute by viewing the molecular structure diagram of the organic material displayed on the monitor, and upload it to the host computer. The host computer is used to execute: The system obtains the judgment results of technicians on whether the organic material molecules corresponding to the organic material molecular structures displayed on the monitor have the target properties, and combines them with the gaze point trajectory uploaded by the eye tracker and the corresponding organic material molecular structure diagram to form a training dataset. An attribute prediction model is trained using a training dataset. The attribute prediction model includes a gaze trajectory prediction sub-model and an attribute prediction sub-model. The gaze trajectory prediction sub-model is used to predict the gaze trajectory based on the molecular structure diagram of the organic material. The attribute prediction model is used to predict whether the organic material molecule has the target attribute based on the molecular structure diagram of the organic material and the gaze trajectory prediction information of the gaze trajectory prediction sub-model. The trained property prediction model is used to predict whether the organic material molecule to be tested has the target property.

9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a computer program executable by the processor, which, when executed by the processor, performs the steps of the molecular property prediction method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it performs the steps of the molecular property prediction method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method and device for training fixation point prediction model and electronic equipment

    CN115359092A

  • Training method and device of molecular attribute prediction model, equipment and storage medium

    CN116959611A

  • Point of fixation position prediction model pre-training method and super-division model pre-training method

    CN117218490A

  • Two-channel comparison model for predicting molecular properties

    CN118155746A

  • Eye movement tracking system, method and equipment integrating electroencephalogram signals and non-contact eye tracker

    CN120406749A