A system and method for predicting the absorption spectrum of probe molecules targeting pathogenic proteins

By designing and constructing probe molecules targeting pathogenic proteins, using quantum mechanics/molecular mechanics and machine learning methods, the problem of low absorption spectrum acquisition efficiency of targeted pathogenic protein probe molecules is solved, and efficient and accurate spectral simulation and disease detection are achieved.

CN115662527BActive Publication Date: 2025-08-19QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211328725.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-26
Publication Date
2025-08-19
Estimated Expiration
2042-10-26

AI Technical Summary

Technical Problem

In the prior art, the acquisition efficiency of probe molecular absorption spectra targeting pathogenic proteins is low, experimental operation requirements are high, costly, and there are calculation bottlenecks in theoretical simulation calculations, which limits the detection of disease biomarkers and the full display of pathological mechanisms.

Method used

Design and construct probe molecules targeting pathogenic proteins, optimize the configuration of probe molecules-pathogenic protein complexes through quantum mechanics/molecular mechanics methods and ONIOM models, combine with fully connected neural networks for machine learning, and predict the absorption spectrum of probe molecules in pathogenic proteins.

Benefits of technology

It realizes efficient and accurate simulation of molecular absorption spectra of targeted pathogenic protein probes, reduces measurement and calculation costs, shortens the R&D cycle, and improves the efficiency of biomarker detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115662527B_ABST
    Figure CN115662527B_ABST
Patent Text Reader

Abstract

The present invention discloses a system and method for predicting the absorption spectrum of probe molecules targeting pathogenic proteins, belonging to the field of protein detection. The present invention obtains multiple groups of probe molecule-pathogenic protein complex conformations; divides each group of conformations into regions, calculates and obtains the absorption wavelength and intensity of different excited states of the probe molecules; counts the internal coordinates of the probe molecules after removing hydrogen atoms in the conformation, uses bond length, bond angle, and dihedral angle as molecular descriptors and also as characteristic variables, and uses absorption wavelength and intensity as output, divides the data into training sets and test sets, and builds a fully connected neural network for machine learning; uses the absorption wavelength and intensity of different excited states as the horizontal and vertical coordinates, respectively, to draw the absorption spectrum. The present invention helps researchers accurately and efficiently predict the absorption spectrum of biomarker probe molecules, reduces the time cost, labor cost, and testing cost of measuring the absorption spectrum of probe molecules targeting pathogenic proteins, and is expected to be applied to medical care, biopharmaceuticals and other fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of protein detection, and in particular relates to a system and method for predicting the absorption spectrum of probe molecules targeting pathogenic proteins. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Proteins are the building blocks of life and essential components of cells. They play a crucial role in maintaining normal metabolism and transporting various substances throughout the body. The occurrence of several common diseases is often closely linked to specific proteins. For example, amyloid β (Aβ) and tau are common proteins that contribute to Alzheimer's disease. In the early stages of the disease, amyloid deposits in patients trigger tau phosphorylation, forming neurofibrillary tangles, impaired synaptic function, neuronal loss, and neuroinflammation. Another example is α-synuclein, a key pathogenic protein in Parkinson's disease. It accumulates in neurons, forming Lewy bodies that cannot be degraded by cells, leading to pathological features such as progressive loss of dopaminergic neurons in the substantia nigra pars compacta. Detecting pathogenic proteins in vivo can assess disease progression at the molecular level, providing further insights into disease onset and progression. Small molecule fluorescent probes, due to their advantages such as structural tunability, sensitive response, high selectivity, and visual analysis, show great potential for application in environmental analysis, biomarkers, cell and tissue imaging, clinical diagnosis, and treatment. However, due to the complex and diverse nature of biological systems, the development of small-molecule fluorescent probes with excellent optical properties, sensitive response, and good selectivity for biological sample analysis remains a hot topic and a challenge. In particular, the self-absorption and autofluorescence of probe molecules can interfere with detection, significantly affecting the imaging and detection of proteins in vivo. Therefore, studying the absorption spectra of protein probe molecules is particularly important for achieving clear imaging of proteins in vivo.

[0004] Currently, the main experimental methods for measuring molecular absorption spectra include spectrophotometry, colorimetry, and transmission methods. However, these methods require high instrumentation and technical expertise, are relatively time-consuming, and require high in vivo measurement costs. This limits the acquisition of absorption spectra of probe molecules targeting pathogenic proteins, hindering the detection of relevant disease biomarkers and the comprehensive understanding of pathology and pathogenesis. Traditional quantum chemical calculations can theoretically characterize the optical absorption spectra of fluorescent probes at the molecular level, but the large, complex, and highly variable structure of proteins presents a serious computational bottleneck in calculating the absorption spectra of probe molecules targeting proteins.

[0005] Therefore, in view of the current low efficiency of obtaining the absorption spectra of early disease marker probe molecules, there is an urgent need to develop a method to predict the absorption spectra of protein molecular probes to solve the problems of high requirements for instruments and technicians in experimental operations, relatively long cycles, high costs of in vivo measurements, and computational bottlenecks in theoretical simulation calculations. Summary of the Invention

[0006] To overcome the above-mentioned deficiencies of the prior art, the present invention provides a system and method for predicting the absorption spectrum of probe molecules targeting pathogenic proteins.

[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0008] In a first aspect, a method for predicting the absorption spectrum of a probe molecule targeting a pathogenic protein is disclosed, comprising:

[0009] Design and construct probe molecules targeting pathogenic proteins and preserve the most stable probe molecular structure;

[0010] Extract the corresponding pathogenic protein structure file from the protein database, perform docking between the probe molecule and the pathogenic protein, and obtain multiple sets of probe molecule-pathogenic protein complex conformations;

[0011] Each conformational group is divided into regions, with the probe molecule portion as the quantum mechanics region and the protein portion as the molecular mechanics region. The configuration of the probe molecule-pathogenic protein complex is optimized using quantum mechanics / molecular mechanics methods and the ONIOM model. The excited state of the probe molecule in the pathogenic protein is calculated using time-dependent density functional theory to obtain the absorption wavelength and intensity of different excited states of the probe molecule.

[0012] The internal coordinates of the probe molecules in each conformation after removing hydrogen atoms are calculated. Bond lengths, bond angles, and dihedral angles are used as molecular descriptors and characteristic variables, and the absorption wavelength and intensity of different excited states of the probe molecules are used as output. All data are divided into training and test sets, and a fully connected neural network is built for machine learning. The prediction results of the test set and the loss function curve are plotted.

[0013] The absorption wavelength and intensity of different excited states of the probe molecules in the test set are used as the horizontal and vertical coordinates respectively to draw the absorption spectrum of the probe molecules in the pathogenic protein.

[0014] A further technical solution is to design and construct probe molecules targeting pathogenic proteins and preserve the most stable probe molecular structure. Specifically, the following are: collect or design probe molecules targeting pathogenic proteins, use the molecular configuration visualization software ChemDraw to build the probe molecular structure, minimize the molecular energy through Chem3D, and take the molecular structure with the highest absolute energy value as the most stable probe molecular structure and save it.

[0015] Further technical solutions, docking specifically includes: using Autodock software to perform semi-flexible docking of the probe molecule and the pathogenic protein, the docking box is set to cover the entire pathogenic protein to perform blind docking, and the docking parameters are set as follows: the number of grid points in the docking box in the x, y, and z directions is 120, 100, and 50 respectively, and the length of each grid point is The coordinates of the box center are (15, 25, 5), and the number of docking conformations is set to 25,000.

[0016] A further technical solution is to obtain multiple groups of probe molecule-pathogenic protein complex conformations after docking, specifically: sorting the absolute values of the binding energy of the probe molecule-pathogenic protein complexes obtained by docking in order from high to low, and taking the multiple groups of probe molecule-pathogenic protein complex conformations with the largest energy values as the more stable conformations; preferably, taking the first 12,500 groups of probe molecule-pathogenic protein complex conformations.

[0017] A further technical solution is to divide each group of conformations into regions, using the probe molecule part as the quantum mechanics region and the protein part as the molecular mechanics region. The configuration of the probe molecule-pathogenic protein complex is optimized using quantum mechanics / molecular mechanics methods and the ONIOM model, and the excited state of the probe molecule in the protein is calculated using the time-dependent density functional method. The absorption wavelength and intensity of different excited states of the probe molecule are obtained as follows:

[0018] GaussView6 software was used to divide each set of conformations into a quantum mechanical region and a molecular mechanical region, corresponding to the fluorescent probe molecule and the pathogenic protein, respectively. Quantum chemistry methods were used to calculate the properties of the quantum mechanical region, while molecular mechanics methods were used to describe the influence of the surrounding environment. The two regions were coupled together through electron embedding.

[0019] Then, each set of conformations was optimized using the command ONIOM(b3lyp / 6-31+g(d):UFF=QEq)=Embedcharge opt geom=connectivity. During the optimization, the pathogenic protein portion was fixed, and the probe molecule portion was described using the b3lyp functional and the 6-31+g(d) basis set to optimize the most stable probe molecule-pathogenic protein configuration.

[0020] Based on the most stable probe molecule-pathogenic protein configuration obtained, ONIOM (b3lyp / 6-31+g(d)TD=(NStates=8):UFF=QEq)=Embedcharge geom=connectivity was used to calculate the absorption wavelengths and intensities of the first eight excited states of the probe molecule in the probe molecule-pathogenic protein complex.

[0021] Further technical solutions, specifically machine learning, include:

[0022] Dataset division and normalization: Use the train_test_split() function to divide the total sample data into training set and test set according to the ratio of 4:1, and normalize the values;

[0023] Define the network architecture: Use a fully connected neural network with one input layer, four hidden layers, and one output layer. The number of neurons in each hidden layer is 128, 64, 32, and 16 respectively, and the ReLU activation function is used for activation.

[0024] Model training and testing: The Adam optimizer was used, with MSE as the loss function and MAE as the network performance metric for evaluating the model during training and testing. During model training, the number of samples in each batch was set to 16, and the number of iterations was set to 150.

[0025] Plot the prediction results of the test set and the loss function curve.

[0026] In a second aspect, a system for predicting the absorption spectrum of a probe molecule targeting a pathogenic protein is disclosed, comprising:

[0027] The probe molecule building module is configured to: design and construct probe molecules targeting pathogenic proteins and preserve the most stable probe molecule structure;

[0028] The complex conformation acquisition module is configured to: extract the corresponding pathogenic protein structure file from the protein database, perform docking of the probe molecule and the pathogenic protein, and obtain multiple sets of probe molecule-pathogenic protein complex conformations;

[0029] The absorption wavelength and intensity acquisition module is configured to: divide each set of conformations into regions, using the probe molecule portion as the quantum mechanics region and the protein portion as the molecular mechanics region; optimize the configuration of the probe molecule-pathogenic protein complex using quantum mechanics / molecular mechanics methods and the ONIOM model; and calculate the excited state of the probe molecule in the protein using time-dependent density functional theory to obtain the absorption wavelength and intensity of different excited states of the probe molecule;

[0030] The machine learning module is configured to: calculate the internal coordinates of the probe molecules in each set of conformations after removing hydrogen atoms, use bond lengths, bond angles, and dihedral angles as molecular descriptors and feature variables, and output the absorption wavelength and intensity of different excited states of the probe molecules. It then divides all data into training and test sets, builds a fully connected neural network for machine learning, and plots the prediction results of the test set and the loss function curve.

[0031] The absorption spectrum drawing module is configured to use the absorption wavelength and intensity of different excited states of the probe molecules in the test set as the horizontal axis and the vertical axis respectively, and draw the absorption spectrum of the probe molecules in the pathogenic protein.

[0032] A further technical solution is to divide each group of conformations into regions, using the probe molecule part as the quantum mechanics region and the protein part as the molecular mechanics region. The configuration of the probe molecule-pathogenic protein complex is optimized using quantum mechanics / molecular mechanics methods and the ONIOM model, and the excited state of the probe molecule in the protein is calculated using the time-dependent density functional method. The absorption wavelength and intensity of different excited states of the probe molecule are obtained as follows:

[0033] GaussView6 software was used to divide each set of conformations into a quantum mechanical region and a molecular mechanical region, corresponding to the fluorescent probe molecule and the pathogenic protein, respectively. Quantum chemistry methods were used to calculate the properties of the quantum mechanical region, while molecular mechanics methods were used to describe the influence of the surrounding environment. The two regions were coupled together through electron embedding.

[0034] Each set of conformations was then optimized using the command ONIOM(b3lyp / 6-31+g(d):UFF=QEq)=Embedcharge opt geom=connectivity. During optimization, the pathogenic protein portion was fixed, and the probe molecule portion was described using the b3lyp functional and the 6-31+g(d) basis set. The most stable probe molecule-pathogenic protein configuration was obtained.

[0035] Based on the most stable probe molecule-pathogenic protein configuration obtained, ONIOM (b3lyp / 6-31+g(d)TD=(NStates=8):UFF=QEq)=Embedcharge geom=connectivity was used to calculate the absorption wavelengths and intensities of the first eight excited states of the probe molecule in the probe molecule-pathogenic protein complex;

[0036] Alternatively, the machine learning specifically includes:

[0037] Dataset division and normalization: Use the train_test_split() function to divide the total sample data into training set and test set according to the ratio of 4:1, and normalize the values;

[0038] Define the network architecture: Use a fully connected neural network with one input layer, four hidden layers, and one output layer. The number of neurons in each hidden layer is 128, 64, 32, and 16 respectively, and the ReLU activation function is used for activation.

[0039] Model training and testing: The Adam optimizer was used, with MSE as the loss function and MAE as the network performance metric for evaluating the model during training and testing. During model training, the number of samples in each batch was set to 16, and the number of iterations was set to 150.

[0040] Plot the prediction results of the test set and the loss function curve.

[0041] In a third aspect, a computer device is disclosed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.

[0042] In a fourth aspect, a computer-readable storage medium stores a computer program, which performs the steps of the above method when executed by a processor.

[0043] One or more of the above technical solutions have the following beneficial effects:

[0044] Compared with existing experimental measurement and theoretical calculation techniques, the method of the present invention for predicting the absorption spectrum of probe molecules targeting pathogenic proteins helps researchers accurately and efficiently predict the absorption spectrum of biomarker probe molecules. The method of the present invention reduces the time cost, labor cost and testing cost of measuring the absorption spectrum of probe molecules targeting pathogenic proteins, shortens the R&D cycle, and reduces testing and calculation costs. At the same time, it achieves efficient and accurate simulation of the absorption spectrum of probe molecules targeting pathogenic proteins, and is expected to be applied to the fields of medical care, biopharmaceuticals, etc.

[0045] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0047] Figure 1 This is a schematic diagram of the technical route of Example 1 of the present invention;

[0048] Figure 2 Schematic diagram of the probe molecular structure selected in Example 1 of the present invention;

[0049] Figure 3 These are the four sites with the largest absolute values of binding energy obtained by molecular docking in Example 1 of the present invention. DETAILED DESCRIPTION

[0050] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0051] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.

[0052] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0053] Explanation of terms: Pathogenic protein refers to a protein associated with a disease, and probe molecule refers to a molecular system that can specifically bind to the target and produce a significant change in the fluorescence signal.

[0054] Example 1

[0055] This embodiment discloses a method for predicting the absorption spectrum of a probe molecule targeting a pathogenic protein, comprising:

[0056] (1) Probe molecular model construction and preliminary optimization

[0057] choose Figure 2 The fluorescent probe molecule targeting the Alzheimer's disease-causing protein Aβ is used as the research object. The probe molecular structure is built using ChemDraw, and the molecular energy is minimized using Chem3D to save the molecular structure.

[0058] (2) From the Protein Data Bank (PDB) https: / / www.pdbus.org / ) Download the structure PDB file of the Alzheimer's disease-causing protein Aβ (PDB ID: 5OQV).

[0059] (3) Autodock software was used to perform semi-flexible docking of the probe molecule and the pathogenic protein. Since the binding site of the probe molecule on Aβ was uncertain, the docking box was set to cover the entire protein to perform blind docking. The docking parameters were set as follows: the number of grid points in the docking box in the x, y, and z directions was 120, 100, and 50, respectively, and the length of each grid point was The center coordinates of the box are (15, 25, 5), and the number of docking conformations is set to 25000. The absolute values of the binding energies obtained by docking are sorted from high to low (the four docking sites with the largest absolute values of binding energy in the molecular docking results are shown in Figure 2). Figure 3 ), taking the first 12,500 sets of probe molecule-pathogenic protein complex conformations.

[0060] (4) GaussView 6.0 was used to divide each probe molecule-pathogenic protein complex system into regions: the probe molecule was set as the quantum mechanics region, the protein was set as the molecular mechanics region, and the two parts were coupled together by electron embedding.

[0061] (5) The configuration of the probe molecule in the protein was optimized in the Gaussian16 software package using the following command. During the optimization, the protein portion was fixed, and the probe molecule portion was described using the b3lyp functional and the 6-31+g(d) basis set:

[0062] ONIOM(b3lyp / 6-31+g(d):UFF=QEq)=Embedcharge opt geom=connectivity.

[0063] (6) Based on the optimized configuration, the absorption wavelengths of the first eight excited states of the probe molecule in the protein were calculated using the following command in the Gaussian16 software package:

[0064] ONIOM(b3lyp / 6-31+g(d)TD=(NStates=8):UFF=QEq)=Embedcharge geom=connectivity.

[0065] (7) The internal coordinates of each group of probe molecules after removing hydrogen atoms were counted, and 16 molecular descriptors were screened as characteristic variables (see Table 1).

[0066] Table 1 Molecular descriptors

[0067]

[0068]

[0069] (8) Using absorption wavelength and intensity as output, a fully connected neural network is constructed for machine learning, specifically including:

[0070] Dataset division and normalization: The total sample data is 12,500. The train_test_split() function is used to divide the dataset into training and test sets in a ratio of 4:1, with 10,000 and 2,500 sets of data respectively, and the values are normalized.

[0071] Define the network architecture: Use a fully connected neural network with one input layer, four hidden layers, and one output layer. The number of neurons in each hidden layer is 128, 64, 32, and 16 respectively, and the relu activation function is used for activation.

[0072] Model training and testing: The Adam optimizer was used, with MSE as the loss function and MAE as the performance metric for evaluating the model during training and testing. During model training, the number of samples per batch was set to 16, and the number of iterations was set to 150.

[0073] Plot the prediction results of the test set and the loss function curve.

[0074] (9) The absorption wavelength and intensity of the outputted different excited states are used as the horizontal and vertical coordinates, respectively, to plot the absorption spectrum of the probe molecule in Aβ.

[0075] Example 2

[0076] A system for predicting the absorption spectrum of a probe molecule targeting a pathogenic protein, comprising:

[0077] The probe molecule building module is configured to: design and construct probe molecules targeting pathogenic proteins and preserve the most stable probe molecule structure.

[0078] The complex conformation acquisition module is configured to: extract the corresponding pathogenic protein structure file from the protein database, perform docking of the probe molecule and the pathogenic protein, and obtain multiple sets of probe molecule-pathogenic protein complex conformations.

[0079] The absorption wavelength and intensity acquisition module is configured to: divide each set of conformations into regions, using the probe molecule portion as the quantum mechanics region and the protein portion as the molecular mechanics region; optimize the configuration of the probe molecule-pathogenic protein complex using quantum mechanics / molecular mechanics methods and the ONIOM model; and calculate the excited state of the probe molecule in the protein using time-dependent density functional theory to obtain the absorption wavelength and intensity of different excited states of the probe molecule;

[0080] GaussView6 software was used to divide each set of conformations into a quantum mechanical region and a molecular mechanical region, corresponding to the fluorescent probe molecule and the pathogenic protein, respectively. Quantum chemistry methods were used to calculate the properties of the quantum mechanical region, while molecular mechanics methods were used to describe the influence of the surrounding environment. The two regions were coupled together through electron embedding.

[0081] Each set of conformations was then optimized using the command ONIOM(b3lyp / 6-31+g(d):UFF=QEq)=Embedcharge opt geom=connectivity. During optimization, the pathogenic protein portion was fixed, and the probe molecule portion was described using the b3lyp functional and the 6-31+g(d) basis set. The most stable probe molecule-pathogenic protein configuration was obtained.

[0082] Based on the most stable probe molecule-pathogenic protein configuration obtained, ONIOM (b3lyp / 6-31+g(d)TD=(NStates=8):UFF=QEq)=Embedcharge geom=connectivity was used to calculate the absorption wavelengths and intensities of the first eight excited states of the probe molecule in the probe molecule-pathogenic protein complex.

[0083] The machine learning module is configured to: calculate the internal coordinates of the probe molecules in each set of conformations after removing hydrogen atoms, use bond lengths, bond angles, and dihedral angles as molecular descriptors and feature variables, and output the absorption wavelength and intensity of different excited states of the probe molecules. It then divides all data into training and test sets, builds a fully connected neural network for machine learning, and plots the prediction results of the test set and the loss function curve.

[0084] The machine learning specifically includes:

[0085] Dataset division and normalization: Use the train_test_split() function to divide the total sample data into training set and test set according to the ratio of 4:1, and normalize the values;

[0086] Define the network architecture: Use a fully connected neural network with one input layer, four hidden layers, and one output layer. The number of neurons in each hidden layer is 128, 64, 32, and 16 respectively, and the ReLU activation function is used for activation.

[0087] Model training and testing: The Adam optimizer was used, with MSE as the loss function and MAE as the network performance metric for evaluating the model during training and testing. During model training, the number of samples in each batch was set to 16, and the number of iterations was set to 150.

[0088] Plot the prediction results of the test set and the loss function curve.

[0089] The absorption spectrum drawing module is configured to use the absorption wavelength and intensity of different excited states of the probe molecules in the test set as the horizontal axis and the vertical axis respectively, and draw the absorption spectrum of the probe molecules in the pathogenic protein.

[0090] Example 3

[0091] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.

[0092] Example 4

[0093] The purpose of this embodiment is to provide a computer-readable storage medium.

[0094] A computer-readable storage medium stores a computer program, which executes the steps of the above method when executed by a processor.

[0095] The steps involved in the above embodiments 3 and 4 correspond to those in the method embodiment 1. For detailed implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media that includes one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to perform any method of the present invention.

[0096] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0097] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.

Claims

1. A method for predicting the absorption spectrum of a probe molecule targeting a pathogenic protein, characterized in that: include: Design and construct probe molecules targeting pathogenic proteins and preserve the most stable probe molecular structure; Extract the corresponding pathogenic protein structure file from the protein database, perform docking between the probe molecule and the pathogenic protein, and obtain multiple sets of probe molecule-pathogenic protein complex conformations; Each conformational group is divided into regions, with the probe molecule portion as the quantum mechanics region and the protein portion as the molecular mechanics region. The configuration of the probe molecule-pathogenic protein complex is optimized using quantum mechanics / molecular mechanics methods and the ONIOM model. The excited state of the probe molecule in the pathogenic protein is calculated using time-dependent density functional theory to obtain the absorption wavelength and intensity of different excited states of the probe molecule. The internal coordinates of the probe molecules in each conformation after removing hydrogen atoms are calculated. Bond lengths, bond angles, and dihedral angles are used as molecular descriptors and characteristic variables, and the absorption wavelength and intensity of different excited states of the probe molecules are used as output. All data are divided into training and test sets, and a fully connected neural network is built for machine learning. The prediction results of the test set and the loss function curve are plotted. The absorption wavelength and intensity of different excited states of the probe molecules in the test set are used as the horizontal and vertical coordinates respectively to draw the absorption spectrum of the probe molecules in the pathogenic protein.

2. The method for predicting the absorption spectrum of a probe molecule targeting a pathogenic protein according to claim 1, wherein: The design and construction of probe molecules targeting pathogenic proteins and preservation of the most stable probe molecular structure are specifically as follows: collecting or designing probe molecules targeting pathogenic proteins, building the probe molecular structure using molecular configuration visualization software ChemDraw, minimizing molecular energy through Chem3D, and taking the molecular structure with the highest absolute energy value as the most stable probe molecular structure, and preserving it.

3. The method for predicting the absorption spectrum of a probe molecule targeting a pathogenic protein according to claim 1, wherein: The docking is specifically as follows: using Autodock software to perform semi-flexible docking of the probe molecule and the pathogenic protein, the docking box is set to cover the entire pathogenic protein to implement blind docking, and the docking parameters are set as follows: the number of grid points in the docking box in the x, y, and z directions are 120, 100, and 50, respectively, the length of each grid point is 0.375 Å, the coordinates of the box center are (15, 25, 5), and the number of docking conformations is set to 25,000.

4. The method for predicting the absorption spectrum of a probe molecule targeting a pathogenic protein according to claim 1, wherein: After docking, multiple sets of probe molecule-pathogenic protein complex conformations are obtained as follows: the absolute values of the binding energies of the docked probe molecule-pathogenic protein complexes are sorted in descending order, and the multiple sets of probe molecule-pathogenic protein complex conformations with the largest energy values are taken as the more stable conformations; the top 12,500 sets of probe molecule-pathogenic protein complex conformations are taken.

5. The method for predicting the absorption spectrum of a probe molecule targeting a pathogenic protein according to claim 1, wherein: Each group of conformations is divided into regions, with the probe molecule portion as the quantum mechanics region and the protein portion as the molecular mechanics region. The configuration of the probe molecule-pathogenic protein complex is optimized using quantum mechanics / molecular mechanics methods and the ONIOM model, and the excited state of the probe molecule in the protein is calculated using a time-dependent density functional method to obtain the absorption wavelengths and intensities of different excited states of the probe molecule as follows: GaussView6 software was used to divide each set of conformations into a quantum mechanical region and a molecular mechanical region, corresponding to the fluorescent probe molecule and the pathogenic protein, respectively. Quantum chemistry methods were used to calculate the properties of the quantum mechanical region, while molecular mechanics methods were used to describe the influence of the surrounding environment. The two regions were coupled together through electron embedding. Each set of conformations was then optimized using the command ONIOM(b3lyp / 6-31+g(d):UFF=QEq)=Embedcharge opt geom=connectivity. During the optimization, the pathogenic protein portion was fixed, and the probe molecule portion was described using the b3lyp functional and the 6-31+g(d) basis set. The most stable probe molecule-pathogenic protein configuration was obtained through optimization. Based on the most stable probe molecule-pathogenic protein configuration obtained, the absorption wavelengths and intensities of the first eight excited states of the probe molecule in the probe molecule-pathogenic protein complex were calculated using ONIOM (b3lyp / 6-31+g(d) TD=(NStates=8):UFF=QEq)=Embedcharge geom=connectivity.

6. The method for predicting the absorption spectrum of a probe molecule targeting a pathogenic protein according to claim 1, wherein: The machine learning specifically includes: Dataset division and normalization: Use the train_test_split() function to divide the total sample data into training and test sets in a 4:1 ratio, and normalize the values; Define the network architecture: Use a fully connected neural network with one input layer, four hidden layers, and one output layer. The number of neurons in each hidden layer is 128, 64, 32, and 16 respectively, and the ReLU activation function is used for activation. Model training and testing: The Adam optimizer was used, with MSE as the loss function and MAE as the network performance metric for evaluating the model during training and testing. During model training, the number of samples in each batch was set to 16, and the number of iterations was set to 150. Plot the prediction results of the test set and the loss function curve.

7. A system for predicting the absorption spectrum of probe molecules targeting pathogenic proteins, characterized in that: include: The probe molecule building module is configured to: design and construct probe molecules targeting pathogenic proteins and preserve the most stable probe molecule structure; The complex conformation acquisition module is configured to: extract the corresponding pathogenic protein structure file from the protein database, perform docking of the probe molecule and the pathogenic protein, and obtain multiple sets of probe molecule-pathogenic protein complex conformations; The absorption wavelength and intensity acquisition module is configured to: divide each set of conformations into regions, using the probe molecule portion as the quantum mechanics region and the protein portion as the molecular mechanics region; optimize the configuration of the probe molecule-pathogenic protein complex using quantum mechanics / molecular mechanics methods and the ONIOM model; and calculate the excited state of the probe molecule in the protein using time-dependent density functional theory to obtain the absorption wavelength and intensity of different excited states of the probe molecule; The machine learning module is configured to: calculate the internal coordinates of the probe molecules in each set of conformations after removing hydrogen atoms, use bond lengths, bond angles, and dihedral angles as molecular descriptors and feature variables, and output the absorption wavelength and intensity of different excited states of the probe molecules. It then divides all data into training and test sets, builds a fully connected neural network for machine learning, and plots the prediction results of the test set and the loss function curve. The absorption spectrum drawing module is configured to use the absorption wavelength and intensity of different excited states of the probe molecules in the test set as the horizontal axis and the vertical axis respectively, and draw the absorption spectrum of the probe molecules in the pathogenic protein.

8. The system for predicting the absorption spectrum of probe molecules targeting pathogenic proteins according to claim 7, wherein: Each group of conformations is divided into regions, with the probe molecule portion as the quantum mechanics region and the protein portion as the molecular mechanics region. The configuration of the probe molecule-pathogenic protein complex is optimized using quantum mechanics / molecular mechanics methods and the ONIOM model, and the excited state of the probe molecule in the protein is calculated using a time-dependent density functional method to obtain the absorption wavelengths and intensities of different excited states of the probe molecule as follows: GaussView6 software was used to divide each set of conformations into a quantum mechanical region and a molecular mechanical region, corresponding to the fluorescent probe molecule and the pathogenic protein, respectively. Quantum chemistry methods were used to calculate the properties of the quantum mechanical region, while molecular mechanics methods were used to describe the influence of the surrounding environment. The two regions were coupled together through electron embedding. Each set of conformations was then optimized using the command ONIOM(b3lyp / 6-31+g(d):UFF=QEq)=Embedcharge opt geom=connectivity. During the optimization, the pathogenic protein portion was fixed, and the probe molecule portion was described using the b3lyp functional and the 6-31+g(d) basis set. The most stable probe molecule-pathogenic protein configuration was obtained through optimization. Based on the most stable probe molecule-pathogenic protein configuration obtained, ONIOM (b3lyp / 6-31+g(d) TD=(NStates=8):UFF=QEq)=Embedcharge geom=connectivity was used to calculate the absorption wavelengths and intensities of the first eight excited states of the probe molecule in the probe molecule-pathogenic protein complex; Alternatively, the machine learning specifically includes: Dataset division and normalization: Use the train_test_split() function to divide the total sample data into training and test sets in a 4:1 ratio, and normalize the values; Define the network architecture: Use a fully connected neural network with one input layer, four hidden layers, and one output layer. The number of neurons in each hidden layer is 128, 64, 32, and 16 respectively, and the ReLU activation function is used for activation. Model training and testing: The Adam optimizer was used, with MSE as the loss function and MAE as the network performance metric for evaluating the model during training and testing. During model training, the number of samples in each batch was set to 16, and the number of iterations was set to 150. Plot the prediction results of the test set and the loss function curve.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are performed.

Citation Information

Patent Citations

  • Method for predicting properties of tetraphenyl porphyrin compounds substituted by different substituents

    CN107563121A

  • Rapid design method of fluorescent probe for exploring and detecting preliminary tumor screening indexes

    CN114496220A