Method for evaluating activity of multi-target candidate compounds for cerebral ischemia based on in vitro detection

CN122545765APending Publication Date: 2026-08-11KUNMING MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]然而现有的脑缺血多靶点药物筛选方法在实际应用中,多依赖逐项、离散的体外实验进行化合物评估,缺乏面向脑缺血多靶点场景的标准化、系统化检测流程,各检测环节相互独立,检测指标覆盖不全,难以在同一体系内同步完成多靶点活性、入脑能力与安全毒性的联合测定,导致检测效率低、周期长、数据可比性差,同时在计算辅助层面,现有方法多依赖通用分子结构模型进行虚拟筛选,其预测结果缺乏以实验检测数据为基准的校准与验证机制,未建立基于实测数据的参照比对体系,难以对检测结果进行有效的预测引导与数据回溯,且计算预测与实验检测相互脱节,无法形成检测反馈优化的闭环迭代流程,导致筛选结果的可靠性不足、假阳性率高,难以满足脑缺血多靶点药物临床前研发对检测通量与数据精度的双重需求

Benefits of technology

1、本发明通过搭建标注完备的药物靶点关联检测基准数据库,整合分子结构、靶点通路及病理药效多维度实验数据,形成基准数据驱动的标准化筛选逻辑,实现从经验式筛选到可重复基准驱动筛选的升级,并以体外实验实测结果为最终判定依据,深度学习作为特征处理与数据分析的辅助工具,提升多靶点药物筛选的可靠性与筛选通量,为脑缺血治疗药物的候选化合物富集提供了稳定、可追溯的技术路径,同时采用预测初筛和实验复筛的两级血脑屏障筛选机制,先依托基准数据库中已验证的实验数据,通过预测模型完成大规模化合物的入脑能力快速排序初筛,大幅缩小待验证样本范围,提升高通量筛选效率;再通过体外Transwell血脑屏障模型进行实体实验复筛,以实测表观渗透系数作为淘汰判定的硬标准,在提升筛选通量的同时保障入脑性能判定的准确性,避免纯计算预测带来的假阳性偏差。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122545765A_ABST
    Figure CN122545765A_ABST
Patent Text Reader

Abstract

This invention discloses a method for evaluating the activity of multi-target candidate compounds in cerebral ischemia based on in vitro detection, belonging to the field of drug activity detection and screening technology, including the following steps: S1, constructing a benchmark database; S2, completing the initial screening and ranking of compounds based on the benchmark database to achieve blood-brain barrier penetration, and extracting the measured feature vectors of the retained compounds; S3, comparing the measured feature vectors with the benchmark database to predict the multi-target efficacy, verifying the binding activity in vitro using surface plasmon resonance technology, and obtaining a sample set of candidate compounds; S4, collecting in vitro efficacy, toxicity, and blood-brain barrier penetration experimental data of the candidate compounds, and integrating them into a core feature parameter set; S5, combining preset screening rules to output a comprehensive rating, target compatibility results, and safety risk identification. This invention uses experimental verification as the core basis, reduces the false positive rate of screening, and improves screening efficiency and drug-likeness guidance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drug activity detection and screening technology, and more particularly to a method for evaluating the activity of multi-target candidate compounds in cerebral ischemia based on in vitro detection. Background Technology

[0002] Multi-target drug screening for cerebral ischemia is a core step in the development of new drugs for the central nervous system. It is a key bridge connecting the compound library with in vivo efficacy evaluation. Its screening accuracy, ability to identify multi-target synergistic effects, precision of comprehensive judgment on drugability, and efficiency of large-scale compound screening directly determine the development cycle, development cost, and clinical translation success rate of drugs for the treatment of cerebral ischemia.

[0003] The pathological mechanism of cerebral ischemia is complex. The therapeutic potential of candidate compounds is constrained by multiple factors such as target activity, blood-brain barrier penetration ability, metabolic stability, and safety and toxicity. These factors exhibit nonlinear interactions, making it difficult for single-dimensional detection data to effectively support drug evaluation. Objectively, it is necessary to establish a comprehensive detection system covering multiple target binding activity, brain penetration performance, and safety indicators.

[0004] However, existing methods for screening multi-target drugs for cerebral ischemia rely heavily on discrete, item-by-item in vitro experiments for compound evaluation in practical applications. They lack standardized and systematic testing procedures for multi-target scenarios in cerebral ischemia. Each testing step is independent, and the coverage of testing indicators is incomplete. It is difficult to simultaneously determine the combined activity, brain penetration ability, and safety and toxicity of multiple targets within the same system, resulting in low testing efficiency, long cycles, and poor data comparability. At the computational level, existing methods rely heavily on general molecular structure models for virtual screening. Their prediction results lack calibration and verification mechanisms based on experimental testing data. They have not established a reference comparison system based on measured data, making it difficult to effectively guide predictions and backtrack data. Furthermore, computational predictions are disconnected from experimental testing, failing to form a closed-loop iterative process for testing feedback optimization. This leads to insufficient reliability of screening results and a high false positive rate, making it difficult to meet the dual requirements of high throughput and data accuracy for preclinical development of multi-target drugs for cerebral ischemia.

[0005] There are currently no effective solutions to the problems in the relevant technologies. Summary of the Invention

[0006] To address the problems in related technologies, this invention proposes a method for evaluating the activity of multi-target candidate compounds in cerebral ischemia based on in vitro detection, thereby overcoming the aforementioned technical problems in existing related technologies.

[0007] To achieve the above objectives, the specific technical solution adopted by the present invention is as follows: A method for evaluating the activity of multi-target candidate compounds for cerebral ischemia based on in vitro detection includes the following steps: S1. Obtain the compound samples to be screened and preset the drug target association detection benchmark database; S2. Based on the drug target association detection benchmark database, the compounds to be screened are initially ranked according to their blood-brain barrier penetration. The compounds ranked first are then subjected to apparent permeability coefficient determination using the in vitro Transwell blood-brain barrier model. Compounds with unsatisfactory measured values ​​are removed. The remaining compounds are then subjected to molecular structure characterization to obtain measured feature vectors. As a preferred approach, based on a drug target association detection benchmark database, the compounds to be screened are initially ranked according to their blood-brain barrier penetration. The top-ranked compounds are then subjected to apparent permeability coefficient determination using an in vitro Transwell blood-brain barrier model. Compounds with substandard measured values ​​are removed. The remaining compounds are then subjected to molecular structure characterization to obtain measured feature vectors, including the following steps: S21. Input the molecular structure information of the compound sample into the pre-built blood-brain barrier penetration prediction model in the drug target association detection benchmark database, calculate the blood-brain barrier penetration prediction value of each compound and sort it. As a preferred approach, the molecular structure information of the compound samples is input into a pre-built blood-brain barrier penetration prediction model in a drug target association detection benchmark database. The blood-brain barrier penetration prediction values ​​of each compound are calculated and ranked, including the following steps: S211. Using the blood-brain barrier penetration experimental data stored in the drug target association detection benchmark database as the training set, and training a blood-brain barrier penetration prediction model based on deep learning algorithm. S212. Extract the molecular fingerprint features of the compounds to be screened, and input the molecular structure features into the blood-brain barrier penetration prediction model to calculate the blood-brain barrier penetration prediction value of each compound. S213. Sort the blood-brain barrier penetration prediction values ​​from high to low, and select compounds whose prediction values ​​exceed the set threshold as the initial screening compounds.

[0008] S22. For compounds that ranked high in the initial screening, the apparent permeability coefficient of each compound was determined using an in vitro Transwell blood-brain barrier model. As a preferred approach, for compounds that rank highly in the initial screening, the apparent permeability coefficient of each compound is determined using an in vitro Transwell blood-brain barrier model, including the following steps: S221. Brain microvascular endothelial cells were seeded in Transwell chambers and cultured until a tight monolayer was formed to establish an in vitro blood-brain barrier model. S222. Add the compounds that passed the initial screening to the upper chamber of the Transwell chamber and incubate them under culture conditions for a predetermined time. S223. Take samples from the lower chamber, determine the concentration of the compounds, and calculate the apparent permeability coefficient of each compound.

[0009] S23. Remove compounds with an apparent permeability coefficient lower than a preset threshold to obtain the remaining compounds after secondary screening; As a preferred approach, removing compounds with an apparent permeability coefficient lower than a preset threshold to obtain the remaining compounds after secondary screening includes the following steps: S231. Compare the preset lower limit of penetration with the apparent permeability coefficient of the compound; S232. Based on the comparison results, compounds with apparent permeability coefficients below the lower limit threshold are removed, and the remaining compounds are recorded as rescreened retained compounds.

[0010] S24. Perform molecular structure feature processing on the retained compounds after rescreening. Convert the SMILES string of the compound into a molecular graph structure containing atomic features and chemical bond features. Then, use a graph neural network to encode the molecular graph structure to obtain a preliminary molecular representation vector. Then, input the preliminary molecular representation vector into a deep learning representation network for feature extraction and dimension mapping, and output a fixed-dimensional measured feature vector.

[0011] As a preferred approach, the retained compounds from the secondary screening undergo molecular structure characterization. The SMILES string of the compound is converted into a molecular graph structure containing atomic and chemical bond features. A graph neural network is then used to encode the molecular graph structure to obtain a preliminary molecular representation vector. This preliminary molecular representation vector is then input into a deep learning representation network for feature extraction and dimension mapping, outputting a fixed-dimensional measured feature vector. The steps include: S241. Perform molecular structure standardization preprocessing on the compounds retained after repeated screening, convert their SMILES strings into molecular graph structures with atoms as nodes and chemical bonds as edges, and assign chemical attribute features to the nodes and edges. S242. Input the molecular graph structure into a pre-trained graph neural network and encode it through message passing and neighborhood aggregation to obtain the initial molecular representation vector. S243. Input the preliminary molecular representation vector into the deep learning representation network for feature dimensionality reduction and mapping, and output a fixed-dimensional measured feature vector.

[0012] S3. The measured feature vectors are compared with the drug target correlation detection benchmark database in multiple dimensions to obtain compounds with potential multi-target efficacy. For the initially screened compounds, the equilibrium dissociation constants of their interactions with multiple brain ischemia-related targets are determined one by one using surface plasmon resonance technology. Then, a sample set of candidate compounds is obtained based on the measured results and activity screening. As a preferred approach, the measured feature vectors are compared with a drug target correlation detection benchmark database in multiple dimensions to obtain compounds with potential multi-target efficacy. For the initially screened compounds, surface plasmon resonance technology is used to determine their equilibrium dissociation constants with multiple cerebral ischemia-related targets. Then, based on the measured values ​​combined with activity screening, a candidate compound sample set is obtained, including the following steps: S31. Perform feature alignment matching between the measured feature vector and the pre-stored multimodal fusion feature vector in the drug target association detection benchmark database, and calculate the molecular structure feature similarity between the compound sample and the benchmark database sample. As a preferred approach, the measured feature vectors are matched with the pre-stored multimodal fusion feature vectors in the drug target association detection benchmark database. The calculation of the molecular structure feature similarity between the compound sample and the benchmark database sample includes the following steps: S311. Normalize the measured feature vector and the multimodal fusion feature vector pre-stored in the benchmark database to make them the same scale. S312. Calculate the cosine similarity between the normalized measured feature vector and each multimodal fused feature vector, and use it as the molecular structure feature similarity. S313. Based on the similarity of molecular structure features and the target pathway association information in the benchmark database, the multi-target efficacy prediction score of the compound samples is calculated by weighting, and compounds with scores higher than the threshold are selected as compounds with potential multi-target efficacy.

[0013] S32. Based on the target pathway experimental correlation information in the drug target correlation detection benchmark database, infer the target pathway coupling degree corresponding to the compound sample, and combine the molecular structure feature similarity with the target pathway coupling degree to screen out compounds with potential multi-target efficacy. S33. For the screened compounds, surface plasmon resonance technology is used to conduct in vitro multi-target binding experiments, determine the equilibrium dissociation constant of the compound with each target, determine the actual binding activity based on the equilibrium dissociation constant, and screen to obtain a sample set of candidate compounds.

[0014] S4. For the compounds in the candidate compound sample set, in vitro multi-target binding inhibition experiments were performed to determine the half-maximal inhibitory concentration (IC50). The cell protection rate was determined using an oxygen-glucose deprivation-reperfusion neural cell model. Liver microsomal metabolic stability, cardiomyocyte toxicity, gene mutation risk, and blood-brain barrier penetration were also tested. The experimental data were collected, standardized, and integrated to obtain the core feature parameter set. As a preferred approach, for the compounds in the candidate compound sample set, in vitro multi-target binding inhibition experiments were performed to determine the half-maximal inhibitory concentration (IC50). Cell protection rates were measured using an oxygen-glucose deprivation-reperfusion neural cell model. Liver microsomal metabolic stability, cardiomyocyte toxicity, gene mutation risk, and blood-brain barrier penetration verification experiments were also conducted. The data from these experiments were then collected, standardized, and integrated to obtain the core feature parameter set, including the following steps: S41. For the compounds in the candidate compound sample set, in vitro multi-target binding inhibition experiments were performed to detect the half-maximal inhibitory concentrations of the compounds on multiple targets related to cerebral ischemia. Cell activity was detected using an oxygen-glucose deprivation-reperfusion neural cell model to obtain neural cell protection rate data, which were then integrated to obtain multi-dimensional pharmacodynamic data. S42. For the compounds in the candidate compound sample set, liver microsomal metabolic stability test, cardiomyocyte toxicity test and gene mutation risk test are performed respectively. The corresponding metabolic stability parameters, cardiotoxicity risk parameters and gene mutation risk parameters are obtained as toxicity data. At the same time, the blood-brain barrier penetration verification experimental data of the compounds are collected. S43. Standardize the multi-dimensional efficacy data, toxicity data, and blood-brain barrier penetration test data respectively, and integrate them according to the compound sample identifier to generate a core feature parameter set.

[0015] S5. It presets efficacy threshold rules, toxicity rejection rules, and penetration lower limit rules, and constructs a multi-index comprehensive decision evaluation model to comprehensively evaluate the core feature parameter set and output the comprehensive efficacy rating, target adaptation results, and safety risk identification of each candidate compound.

[0016] As a preferred approach, pre-defined efficacy threshold rules, toxicity rejection rules, and penetration lower limit rules are established. A multi-index comprehensive decision-making evaluation model is constructed to comprehensively evaluate the core feature parameter set and output the comprehensive efficacy rating, target adaptation results, and safety risk identification of each candidate compound, including the following steps: S51. Combining the multi-target half-maximal inhibitory concentration, neuronal protection rate, metabolic stability parameters, cardiotoxicity risk parameters, gene mutation risk parameters, and blood-brain barrier penetration verification data from the core feature parameter set, establish efficacy threshold rules, toxicity rejection rules, and penetration lower limit rules, and construct a multi-index evaluation matrix that includes efficacy indicators, toxicity indicators, and penetration indicators. S52. The multi-index evaluation matrix is ​​standardized, and a comprehensive scoring algorithm based on weighted distance is used to calculate the comprehensive performance score of each candidate compound. Then, based on the target equilibrium dissociation constant and the preset ideal solution of the target, the target fit is calculated. S53. The toxicity rejection rule and the lower limit of penetration rule are used to screen candidate compounds with a single vote. For compounds that are not rejected, the comprehensive rating score and target fit ranking are output based on the comprehensive efficacy score and target fit. S54. Match the comprehensive rating score with the efficacy threshold rule to determine the comprehensive efficacy rating level, and review it in conjunction with the results of in vitro validation experiments to generate the target adaptation results and safety risk labels for each candidate compound.

[0017] The beneficial effects of this invention are as follows: 1. This invention establishes a comprehensive drug target association detection benchmark database, integrating multi-dimensional experimental data on molecular structure, target pathways, and pathological efficacy to form a standardized screening logic driven by benchmark data. This upgrades the screening process from experience-based to repeatable benchmark-driven screening, using in vitro experimental results as the final criterion. Deep learning serves as an auxiliary tool for feature processing and data analysis, improving the reliability and throughput of multi-target drug screening. This provides a stable and traceable technical path for enriching candidate compounds for the treatment of cerebral ischemia. Simultaneously, a two-stage blood-brain barrier screening mechanism is employed: predictive primary screening and experimental secondary screening. First, based on validated experimental data in the benchmark database, a predictive model rapidly ranks compounds for brain penetration, significantly narrowing the range of samples to be validated and improving high-throughput screening efficiency. Second, an in vitro Transwell blood-brain barrier model is used for physical experimental secondary screening, with the measured apparent permeability coefficient as the hard standard for elimination. This improves screening throughput while ensuring the accuracy of brain penetration performance determination, avoiding false positive biases caused by purely computational predictions.

[0018] 2. This invention establishes a closed-loop screening logic that combines multi-target efficacy prediction with in vitro experimental verification. It compares molecular structural features with a benchmark database from multiple dimensions, combining molecular structural similarity with target pathway coupling to preliminarily predict multi-target efficacy, rapidly enriching potentially effective compounds. Then, surface plasmon resonance (SPR) technology is used for in vitro multi-target binding experiments to verify the efficacy. The measured equilibrium dissociation constant is used as the criterion for determining binding activity, reducing the false positive rate of multi-target screening and improving the target matching accuracy of candidate compounds. Simultaneously, the in vitro experimental system collects data on efficacy, toxicity, and blood-brain barrier penetration. The core parameters comprehensively cover the therapeutic potential, safety risks, and brain penetration ability of compounds, providing a solid experimental data foundation for drugability assessment. Then, three types of screening rules—efficacy threshold, toxicity rejection, and lower limit of penetration—are integrated into a multi-task deep learning analysis network as constraints to simultaneously complete comprehensive drugability rating, target suitability determination, and safety risk identification. This achieves integrated joint evaluation of multi-dimensional indicators, which not only improves the efficiency and consistency of drugability assessment but also strengthens the drugability orientation of screening results. The final output rating and label can be directly used for prioritizing compounds in subsequent in vivo efficacy experiments. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart of a method for evaluating the activity of multi-target candidate compounds in cerebral ischemia based on in vitro detection, according to an embodiment of the present invention. Detailed Implementation

[0021] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0022] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0023] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, the deep learning-based multi-target drug screening method for cerebral ischemia according to an embodiment of the present invention includes the following steps: S1. Obtain the compound samples to be screened and preset the drug target association detection benchmark database; Specifically, after initially selecting a batch of candidate compounds from natural product isolated components or commercial small molecule libraries, the purity of each compound is determined one by one using high performance liquid chromatography-mass spectrometry (HPLC-MS). Only compounds with a purity ≥95% are retained. At the same time, the maximum absorption wavelength is determined by ultraviolet-visible spectrophotometry to confirm solubility, and dynamic light scattering technology is used to confirm that they do not aggregate in buffer solution. After passing the above analysis and testing, the compounds are prepared into stock solutions of uniform concentration and coded and registered, thereby obtaining "compound samples to be screened" with clear composition and controllable quality.

[0024] The apparent permeability coefficients of hundreds of known cerebral ischemia-related compounds were determined using an in vitro Transwell blood-brain barrier model. The equilibrium dissociation constants of these compounds with key cerebral ischemia targets (such as TNF-α, Caspase-3, and NR2B subunits) were determined using surface plasmon resonance (SPR) technology. The half-maximal inhibitory concentration (ICC) was determined using enzyme activity inhibition experiments, and the metabolic half-life was determined using a liver microsomal incubation system. These experimental data were then correlated with the molecular structure information of the compounds, and a graph neural network was used to extract fixed-dimensional molecular feature vectors, which were then stored in a database.

[0025] Simultaneously, based on authoritative signaling pathway databases such as KEGG and Reactome, as well as the STRING protein interaction database, core signaling pathways and protein interaction network data related to the pathological process of cerebral ischemia were extracted. Combined with published experimental validation results of cerebral ischemia pathway cascades, a structured construction of target-pathway association information was completed: First, using the UniProt target number as a unique identifier, the upstream and downstream regulatory relationships of each target in the pathway were sorted out, and a target-pathway mapping relationship table was constructed; then, based on the protein interaction network topology, the node connectivity of each target was calculated, and the signal cascade amplification coefficient of each target was calibrated and normalized to the 0~1 range based on the experimental quantitative data of pathway cascade amplification effect; finally, the target-pathway mapping relationship, node connectivity, cascade amplification coefficient, and pathway topology information were associated with the aforementioned compound experimental data and molecular feature vectors using the target identifier as the association key, forming a queryable, comparable, and multi-target efficacy extrapolation drug target association detection benchmark database. S2. Based on the drug target association detection benchmark database, the compounds to be screened are initially ranked according to their blood-brain barrier penetration. The compounds ranked first are then subjected to apparent permeability coefficient determination using the in vitro Transwell blood-brain barrier model. Compounds with unsatisfactory measured values ​​are removed. The remaining compounds are then subjected to molecular structure characterization to obtain measured feature vectors. In this embodiment of the invention, based on a drug target association detection benchmark database, the compounds to be screened are initially ranked according to their blood-brain barrier penetration. The top-ranked compounds are then subjected to apparent permeability coefficient determination using an in vitro Transwell blood-brain barrier model. Compounds with substandard measured values ​​are removed. The remaining compounds are then subjected to molecular structure characterization to obtain measured feature vectors, including the following steps: S21. Input the molecular structure information of the compound sample into the pre-built blood-brain barrier penetration prediction model in the drug target association detection benchmark database, calculate the blood-brain barrier penetration prediction value of each compound and sort it. In this embodiment of the invention, the steps of inputting the molecular structure information of the compound sample into a pre-built blood-brain barrier penetration prediction model in a drug target association detection benchmark database, calculating and sorting the blood-brain barrier penetration prediction values ​​of each compound include: S211. Using the blood-brain barrier penetration experimental data stored in the drug target association detection benchmark database as the training set, and training a blood-brain barrier penetration prediction model based on deep learning algorithm. Specifically, from the drug target association detection benchmark database, the apparent permeability coefficient data of hundreds of compounds obtained in advance through in vitro Transwell blood-brain barrier model are retrieved as labels. The molecular fingerprint features corresponding to each compound are extracted as input. The molecular fingerprint features include, but are not limited to, the atomic composition, functional group distribution, topological structure descriptor and electrical parameters of the compound. High performance liquid chromatography-mass spectrometry is used to verify the purity of some training compounds, and abnormal samples with purity below 95% are removed. The remaining data are divided into training set and validation set according to a preset ratio.

[0026] Subsequently, an initial prediction model was constructed based on a deep neural network architecture. The molecular fingerprint features of the compound were used as input layer nodes, and the predicted value of the apparent permeability coefficient was used as output layer nodes. The network parameters were iteratively optimized through the backpropagation algorithm until the root mean square error between the predicted value and the measured value on the validation set was less than 0.15 log units. The network weight parameters were then fixed to obtain the blood-brain barrier penetration prediction model.

[0027] S212. Extract the molecular fingerprint features of the compounds to be screened, and input the molecular structure features into the blood-brain barrier penetration prediction model to calculate the blood-brain barrier penetration prediction value of each compound. Specifically, for each compound in the sample to be screened, an appropriate amount of powder was weighed and dissolved in deuterated dimethyl sulfoxide. One-dimensional proton and carbon spectra were measured using nuclear magnetic resonance spectroscopy to confirm the correct molecular structure. At the same time, high performance liquid chromatography-mass spectrometry was used to determine the retention time and mass-to-charge ratio of each compound to ensure that the purity of the compound to be tested was ≥95%. After confirmation of compliance, molecular fingerprint features of each compound are extracted using cheminformatics tools, including at least one of atom-pair fingerprints, topological torsion fingerprints, and Morgan fingerprints. After standardizing and encoding the molecular fingerprint features, they are input into the trained blood-brain barrier penetration prediction model in batches. The model calculates based on the input molecular fingerprint features through forward propagation and outputs the blood-brain barrier penetration prediction value corresponding to each compound. The batch results are output in descending order of the prediction value. For batches with abnormally low prediction values, high-performance liquid chromatography-mass spectrometry is repeated to eliminate data deviations caused by sample degradation or preparation errors.

[0028] S213. Sort the blood-brain barrier penetration prediction values ​​from high to low, and select compounds whose prediction values ​​exceed the set threshold as the initial screening compounds.

[0029] Specifically, the predicted values ​​of each compound calculated by the blood-brain barrier penetration prediction model were sorted in descending order of numerical value, and the penetration threshold was set to 1×10⁻⁶. -6 cm / s, this threshold was determined with reference to the lower limit of the apparent permeability coefficient of positive control compounds (such as caffeine and diazepam) that have been experimentally confirmed to have good penetration ability by the Transwell blood-brain barrier model in the drug target association detection benchmark database; a predicted value ≥1×10 cm / s was selected. -6 For compounds with a permeability of cm / s, at least 10% of the samples are randomly selected from the top-ranked compounds and their apparent permeability coefficients are measured and verified using an in vitro Transwell blood-brain barrier model. If the deviation between the measured value and the predicted value exceeds ±0.2 log units, the prediction model parameters are adjusted or the screening threshold is increased, and the ranking and selection operations are repeated. After successful verification, compounds with predicted values ​​exceeding the threshold are registered as initially screened compounds, and their predicted values ​​and ranking positions are recorded. The purity of each initially screened compound is verified using high performance liquid chromatography-mass spectrometry (HPLC-MS). Compounds with a purity below 95% due to degradation during storage or operation are removed, and the remaining compounds are officially initially screened compounds and proceed to the subsequent in vitro assay stage.

[0030] S22. For compounds that ranked high in the initial screening, the apparent permeability coefficient of each compound was determined using an in vitro Transwell blood-brain barrier model. In this embodiment of the invention, the apparent permeability coefficient of each compound ranked high in the initial screening is determined using an in vitro Transwell blood-brain barrier model, including the following steps: S221. Brain microvascular endothelial cells were seeded in Transwell chambers and cultured until a tight monolayer was formed to establish an in vitro blood-brain barrier model. Specifically, human brain microvascular endothelial cells passaged to the 3rd to 5th generation were used at a concentration of 1×10⁻⁶. 5 Seeds were inoculated at a density of cells / well onto Transwell chamber polyester membranes pre-coated with type I rat tail collagen. Endothelial cell culture medium containing 10% fetal bovine serum, 1% endothelial cell growth supplement and penicillin-streptomycin antibiotics was added to the upper chamber, and DMEM high-glucose medium containing 10% fetal bovine serum was added to the lower chamber. The chambers were incubated statically at 37°C and 5% CO2. The culture medium in the upper and lower chambers was changed every 48 hours. Starting from day 5, the transendothelial resistance values ​​on both sides of the polyester membrane of the chambers were measured daily using a transepithelial resistance meter.

[0031] Simultaneously, fluorescein isothiocyanate-labeled dextran was added to the lower chamber culture medium. After 1 hour of incubation, samples were taken from the lower chamber, and the fluorescence intensity was measured using a fluorescence spectrophotometer at an excitation wavelength of 490 nm and an emission wavelength of 520 nm. The fluorescein transmittance was calculated. When the transendothelial resistance value stably reached 200 Ω·cm2 The above and the fluorescein transmittance is less than 1×10 -6 When the flow rate reaches cm / s, it is determined that a monolayer tight junction has been formed, that is, the in vitro blood-brain barrier model has been established. Before use, high performance liquid chromatography-mass spectrometry is used to detect the lactate dehydrogenase activity in the culture medium supernatant to confirm the cell membrane integrity and remove Transwell chambers with cell damage caused by operation.

[0032] S222. Add the compounds that passed the initial screening to the upper chamber of the Transwell chamber and incubate them under culture conditions for a predetermined time. Specifically, compounds that passed the initial screening were prepared into a 10 μmol / L test solution using a Hanks equilibrium salt solution containing 0.5% dimethyl sulfoxide. The actual concentration of the compounds in the test solution was determined using high-performance liquid chromatography (HPLC) to ensure that the preparation error did not exceed ±5%. A transendothelial resistance value ≥200 Ω·cm was used. 2 Furthermore, the transmittance of fluorescein is less than 1×10⁻⁶. -6 A qualified Transwell chamber with a speed of cm / s was prepared. The old culture medium in both the upper and lower chambers was aspirated, and the chambers were gently rinsed three times with 1 mL of pre-warmed Hanks' balanced salt solution (37°C) each time to remove interference from serum components. 500 μL of the test solution was added to the upper chamber, and 1.5 mL of Hanks' balanced salt solution containing 4% bovine serum albumin was added to the lower chamber as the receiving solution. The chamber was placed in an incubator at 37°C, 5% CO2, and 95% relative humidity, and incubated on a 60 rpm steady-state shaker. During incubation, the resistance value was monitored every 30 minutes using a transepithelial resistivity meter. If the resistance value drops by more than 20% of the initial value, the integrity of the barrier in the chamber is deemed to have been compromised during the experiment, the labeled data is invalid, and the test is terminated. After incubation for the preset 120 minutes, 100 μL samples are taken from the upper and lower chambers respectively, and an equal volume of acetonitrile containing an internal standard (deuterated counterpart or structural analog) is added to precipitate the protein. The mixture is vortexed for 3 minutes and centrifuged at 4°C and 14,000 rpm for 15 minutes. The supernatant is then used to determine the concentration of the compound using liquid chromatography-tandem mass spectrometry. The apparent permeability coefficient is calculated based on the initial concentration in the upper chamber and the final concentration in the lower chamber.

[0033] S223. Take samples from the lower chamber, determine the concentration of the compounds, and calculate the apparent permeability coefficient of each compound.

[0034] Specifically, the supernatant from the lower chamber after protein precipitation and centrifugation was used for quantitative analysis using liquid chromatography-tandem mass spectrometry (LC-MS / MS). The LC conditions were set as follows: C18 reversed-phase column, column temperature 40℃; mobile phase A was an aqueous solution containing 0.1% formic acid; mobile phase B was an acetonitrile solution containing 0.1% formic acid; the gradient elution program was as follows: 0-1 min to maintain 5% B phase, 1-3 min to linearly increase from 5% B phase to 95% B phase, 3-5 min to maintain 95% B phase, 5-5.1 min to linearly decrease from 95% B phase to 5% B phase, and 5.1-7 min to maintain 5% B phase for column equilibration; flow rate 0.4 mL / min. The injection volume was 5 μL. Mass spectrometry conditions were set as follows: electrospray ionization source in positive or negative ion mode, multiple reaction monitoring (MRM) mode. Characteristic precursor ion-daughter ion pairs were selected as quantitative and qualitative ion pairs for each analyte and internal standard, respectively. Quantification was performed using the internal standard curve method. The standard curve was prepared using a blank receiving solution without the compound as the matrix, with a concentration range of 0.1 nmol / L to 10 μmol / L. Linear regression was performed with the ratio of the compound peak area to the internal standard peak area in each standard solution as the ordinate and the compound concentration as the abscissa. The correlation coefficient R0 of the standard curve was calculated. 2 Only samples with a concentration ≥0.995 can be used. During each batch of testing, blank samples, samples with the lowest limit of quantitation, and quality control samples with three concentrations (high, medium, and low) should be run simultaneously. The deviation between the measured concentration and the labeled concentration of the quality control samples should be within ±15%. Record the concentration of the compound in the lower chamber at each time point, and set up three parallel Transwell chambers for testing the same compound. Calculate the mean and relative standard deviation of the three parallel test results. If the relative standard deviation exceeds 20%, add two parallel chambers and retest. After removing outliers, take the mean as the final apparent permeability coefficient.

[0035] S23. Remove compounds with an apparent permeability coefficient lower than a preset threshold to obtain the remaining compounds after secondary screening; As a preferred approach, removing compounds with an apparent permeability coefficient lower than a preset threshold to obtain the remaining compounds after secondary screening includes the following steps: S231. Compare the preset lower limit of penetration with the apparent permeability coefficient of the compound; Specifically, apparent permeability coefficient data of positive control compounds (such as caffeine and diazepam) and negative control compounds (such as sucrose and mannitol) measured in vitro using the Transwell blood-brain barrier model were retrieved from the drug target association detection benchmark database. The lowest value of the measured apparent permeability coefficient of the positive control compound was used as the upper limit of reference, and the highest value of the measured apparent permeability coefficient of the negative control compound was used as the lower limit of reference. Combined with the clinical blood-brain barrier penetration requirements of drugs for treating cerebral ischemia, the lower limit of penetration threshold was preset to 5 × 10⁻⁶. -7cm / s, this threshold is lower than the lower limit of the positive control to ensure that no potentially effective compounds are missed, and higher than the upper limit of the negative control to effectively exclude compounds with insufficient penetration ability.

[0036] Next, the apparent permeability coefficients of each compound calculated above are compared with the lower limit threshold value. A dual independent calculation verification method is used, with two experimenters recalculating the apparent permeability coefficient based on the integrated peak area of ​​the original chromatogram obtained from liquid chromatography-tandem mass spectrometry. If the deviation between the two calculations exceeds ±10%, the lower chamber sample of that compound is re-injected and measured until the deviation is within ±10%. The verified apparent permeability coefficient is then compared with the lower limit threshold of 5 × 10⁻⁶. -7 The apparent permeability coefficient of each compound was compared one by one with the cm / s value, and the difference between the apparent permeability coefficient and the threshold value was recorded, along with the comparison conclusion, to form a permeability evaluation record table.

[0037] S232. Based on the comparison results, compounds with apparent permeability coefficients below the lower limit threshold are removed, and the remaining compounds are recorded as rescreened retained compounds.

[0038] Specifically, based on the apparent permeability coefficient and lower limit threshold (5×10) of each compound in the generated permeability assessment record table... -7 The comparison of cm / s concludes that those with an apparent permeability coefficient below 5×10 -7 Compounds with a penetration rate of cm / s are marked as "non-compliant" and added to the rejection list one by one. Before rejection, for each batch of compounds below the threshold, at least 20% of the samples are randomly selected, and the concentration of the compound in the lower chamber sample is re-determined using liquid chromatography-tandem mass spectrometry. At the same time, the initial concentration of the test solution in the upper chamber is re-measured. The apparent penetration coefficient is recalculated based on the re-measurement data. If the re-measurement result crosses the threshold (i.e., the re-measurement value ≥ 5 × 10⁻⁶), the apparent penetration coefficient is rejected. -7 If the apparent permeability is ≥5×10 cm / s, the compound is removed from the exclusion list, and the retesting rate for other excluded compounds in the same batch is increased to 50% to rule out false negatives caused by injection errors or pretreatment losses. After the retest confirms that there are no errors, the test solution in the upper chamber and the receiving solution in the lower chamber of the Transwell chamber corresponding to the compound in the exclusion list are frozen at -80°C for a period of not less than 6 months for traceability verification. -7For compounds retained at cm / s, the actual concentration of each compound in the stock solution was determined using high-performance liquid chromatography (HPLC) to ensure that the concentration change after long-term storage does not exceed ±10%. At the same time, dynamic light scattering technology was used to confirm that no aggregates formed in the test solution. After passing the test, the compound number, apparent permeability coefficient value, relative standard deviation of parallel determination, and storage conditions were recorded one by one. After being uniformly coded, they were officially recorded in the list of compounds to be retained for rescreening. The complete set of data for each compound, including the original Transwell measured spectrum data, peak area integral records, standard curve equations, and correlation coefficients, was archived as a traceable file.

[0039] S24. Perform molecular structure feature processing on the retained compounds after rescreening. Convert the SMILES string of the compound into a molecular graph structure containing atomic features and chemical bond features. Then, use a graph neural network to encode the molecular graph structure to obtain a preliminary molecular representation vector. Then, input the preliminary molecular representation vector into a deep learning representation network for feature extraction and dimension mapping, and output a fixed-dimensional measured feature vector.

[0040] In this embodiment of the invention, molecular structure characterization processing is performed on the retained compounds after repeated screening. The SMILES string of the compound is converted into a molecular graph structure containing atomic features and chemical bond features. A graph neural network is used to encode the molecular graph structure to obtain a preliminary molecular representation vector. The preliminary molecular representation vector is then input into a deep learning representation network for feature extraction and dimension mapping, and a fixed-dimensional measured feature vector is output. The process includes the following steps: S241. Perform molecular structure standardization preprocessing on the compounds retained after repeated screening, convert their SMILES strings into molecular graph structures with atoms as nodes and chemical bonds as edges, and assign chemical attribute features to the nodes and edges. Specifically, the structures of each compound retained from the secondary screening were confirmed and standardized: approximately 2 mg of each compound powder was dissolved in 0.6 mL of deuterated dimethyl sulfoxide, and one-dimensional nuclear magnetic resonance spectroscopy was performed using a 600 MHz nuclear magnetic resonance spectrometer. 1 H spectrum and 13 C spectrum, additional measurements if necessary. 1 H- 1 Two-dimensional spectra of HCOSY, HSQC, and HMBC were used to compare the measured chemical shift values ​​with literature data or theoretical predictions to confirm that the molecular linkage sequence and stereoconfiguration were correct. At the same time, a trace amount of compound solution was taken for high-resolution mass spectrometry, and the positive and negative ion switching mode of the electrospray ionization source was used to record the accurate mass number. The deviation between the measured mass-to-charge ratio and the theoretical value was required to be less than 3 ppm. The molecular formula was verified by isotope abundance distribution.

[0041] After structural verification, the molecular structure is read in using cheminformatics tools to generate a standardized SMILES string. The SMILES string is then standardized, including removing redundant stereochemical identifiers, unifying tautomer representations, hydrogenation, and charge neutralization. The standardized SMILES string is then parsed into a molecular graph structure, with each non-hydrogen atom as a node and the chemical bonds between atoms as edges. Each atomic node is assigned chemical attributes, including atom type, hybridization state, formal charge, chiral label, ring assignment, and values ​​calculated using the Gasteiger method. The local charge distribution is calculated; each edge is assigned a bond type (single, double, triple, or aromatic) and conjugation attribute; for some compounds with special structures, the retention factor is determined by reversed-phase high-performance liquid chromatography on a C18 column with methanol-water gradient elution, and then the measured value of the lipid-water partition coefficient is calculated. This measured lipid-water partition coefficient is added to the graph representation as a global feature of the molecular graph; the confirmed molecular structure spectrum data, SMILES string, molecular graph node-edge attribute matrix, and associated measured physicochemical parameters of each compound are stored in the database to form a structure confirmation and characterization file.

[0042] S242. Input the molecular graph structure into a pre-trained graph neural network and encode it through message passing and neighborhood aggregation to obtain the initial molecular representation vector. Specifically, the molecular graph structure and node-edge attribute matrix of each of the re-screened and retained compounds are used as inputs and loaded one by one into a pre-trained graph neural network model. This graph neural network adopts a graph isomorphic network architecture, containing four message passing layers, with each layer having a hidden dimension of 128, and the ReLU function is used as the activation function. In the message passing stage, for each atom node in the molecular graph, the feature vectors of its neighboring nodes are concatenated with the bond type features of the connected edges. A learnable message function is used to generate neighbor passing messages, and a summation aggregation function is used to aggregate the node's own features with all neighbor passing messages to update the hidden layer representation of the node. After each layer of message passing is completed, a batch normalization layer is used to normalize the node features to prevent gradient vanishing or exploding.

[0043] After four layers of message passing and neighborhood aggregation iterations, a readout layer performs global summation pooling on all node features, outputting a preliminary 128-dimensional molecular characterization vector for each molecule. During the encoding process, high-performance liquid chromatography-mass spectrometry (HPLC-MS / MS) is used to randomly select at least 5% of the samples from each batch of input compounds for purity verification. If the purity of any sampled compound is lower than 95%, the encoding for that batch is paused, and the purity of the test solutions for all compounds in that batch is re-determined. After removing the substandard compounds, the encoding process is restarted. The preliminary molecular characterization vector... After output, the vectors generated by the same coding process are compared with those of standard compounds in the drug target-associated detection benchmark database that have a clear blood-brain barrier penetration category (strong / weak penetration). The cosine similarity with the vectors of strong-penetrating standards is calculated. If the similarity of a compound vector with all strong-penetrating standards is less than 0.5, the molecular conformation of the compound is verified by one-dimensional nuclear magnetic resonance hydrogen spectroscopy to confirm that it has not undergone structural isomerization or degradation in solution. Only then can the coded vector continue to be used. Otherwise, it is marked as an abnormal characterization, and the sample is re-prepared and re-coded.

[0044] S243. Input the preliminary molecular representation vector into the deep learning representation network for feature dimensionality reduction and mapping, and output a fixed-dimensional measured feature vector.

[0045] Specifically, the 128-dimensional preliminary molecular characterization vectors of each retained compound from the secondary screening were loaded batch by batch into a deep learning characterization network for feature dimensionality reduction and mapping. This deep learning characterization network adopted an autoencoder architecture with three fully connected layers. The encoder part had dimensions of 128-64-32, and the decoder part had dimensions of 32-64-128. The intermediate bottleneck layer was set to 32 dimensions. Each fully connected layer was followed by a batch normalization layer and a LeakyReLU activation function. The network training aimed to minimize the mean square error between the input vector and the reconstructed output vector. The Adam optimizer was used for iterative training until the reconstruction error converged to below 0.01. After the input vector was compressed layer by layer by the encoder, the 32-dimensional vector output by the bottleneck layer was taken as the measured feature vector after dimensionality reduction. Before each batch of feature dimensionality reduction processing, the high performance liquid chromatography-tandem mass spectrometry standard quality control data pre-stored in the drug target association detection benchmark database was used as a reference.

[0046] Next, the absorbance ratio of the batch-processed samples at 260 nm and 280 nm was determined using UV-Vis spectrophotometry. Sample vectors with abnormal ratios (outside the range of 1.8-2.0) were discarded to eliminate characteristic deviations caused by test solution contamination or protein residue. After each batch of characteristic vectors was output, no fewer than three compounds were randomly selected from that batch, and their one-dimensional proton NMR spectra were re-determined using proton NMR spectroscopy. The fingerprint region of these spectra was compared with the stored confirmatory spectral data, and the peak shape correlation coefficient between the two in the δ 6.0-8.5 ppm range was calculated. The correlation coefficient was required to be no less than 0.95 to confirm the compound's multi-phase absorption. During the second processing, the molecular structure did not undergo degradation or isomerization. After passing the comparison, the 32-dimensional measured feature vector of this batch was officially written into the database of compounds retained for rescreening. For compounds whose vectors after feature dimensionality reduction had abnormal distributions (the variance of a single dimension exceeded 3 times the variance of the entire batch of vectors), the non-specific binding response value of the compound to serum albumin was determined using surface plasmon resonance technology. If the response value exceeded 50RU, the compound was determined to have a non-specific aggregation tendency. This information was added to the measured feature vector as an additional feature dimension to form the final 32+1-dimensional measured feature vector and marked with an aggregation warning mark.

[0047] S3. The measured feature vectors are compared with the drug target correlation detection benchmark database in multiple dimensions to obtain compounds with potential multi-target efficacy. For the initially screened compounds, the equilibrium dissociation constants of their interactions with multiple brain ischemia-related targets are determined one by one using surface plasmon resonance technology. Then, a sample set of candidate compounds is obtained based on the measured results and activity screening. In this embodiment of the invention, the measured feature vectors are compared with a drug target correlation detection benchmark database in multiple dimensions to obtain compounds with potential multi-target efficacy. For the initially screened compounds, surface plasmon resonance technology is used to determine their equilibrium dissociation constants with multiple cerebral ischemia-related targets. Then, based on the measured values ​​combined with activity screening, a candidate compound sample set is obtained, including the following steps: S31. Perform feature alignment matching between the measured feature vector and the pre-stored multimodal fusion feature vector in the drug target association detection benchmark database, and calculate the molecular structure feature similarity between the compound sample and the benchmark database sample. In this embodiment of the invention, the process of performing feature alignment matching between the measured feature vector and the pre-stored multimodal fusion feature vector in the drug target association detection benchmark database, and calculating the molecular structure feature similarity between the compound sample and the benchmark database sample includes the following steps: S311. Normalize the measured feature vector and the multimodal fusion feature vector pre-stored in the benchmark database to make them the same scale. Specifically, before normalization, the integrity of the measured feature vectors of each retained compound from the secondary screening is verified: the absorbance of the compound test solution is measured at 260 nm and 280 nm using ultraviolet-visible spectrophotometry, the A260 / A280 ratio is calculated, and the corresponding vectors of samples with ratios exceeding the range of 1.8-2.0 are removed to eliminate interference from nucleic acid or protein residues on the feature vectors; at the same time, dynamic light scattering technology is used to determine the hydrated particle size of the compound test solution, and the corresponding vectors of samples with hydrated particle sizes exceeding 100 nm are removed to eliminate false positive signals caused by aggregate formation; the verified measured feature vectors are temporarily retained for later use.

[0048] Pre-stored multimodal fusion feature vectors were retrieved from the drug target association detection benchmark database. The source of each vector was traced and verified. High-performance liquid chromatography-mass spectrometry was used to verify the purity of the stock solution of the vector source compound. Only the multimodal fusion feature vectors corresponding to compounds with a purity ≥95% were retained. At the same time, surface plasmon resonance technology was used to determine the equilibrium dissociation constant between the vector source compound and the corresponding target. The measured value was compared with the value recorded in the database, and vectors with a deviation of more than ±0.3 log units were removed to ensure that the multimodal fusion feature vectors used as normalization references are consistent with the actual binding activity.

[0049] The verified measured feature vectors and the verified multimodal fusion feature vectors were then processed using the Min-Max normalization method. The minimum and maximum values ​​of all dimensions of each vector were used as scaling boundaries to linearly map the values ​​of each dimension to the interval [0, 1]. To verify the scale consistency after normalization, at least 5% of the samples were randomly selected from each of the two types of vectors, and the L2 norm of each sample vector after normalization was calculated. If the difference between the mean L2 norms of the two types of vectors exceeded 0.05, the normalization boundary parameters were recalibrated and the normalization operation was re-executed. After normalization, the absorbance of each compound test solution retained after normalization was measured at 280 nm using a UV spectrophotometer to confirm that the compound had not undergone degradation or adsorption loss during the normalization process. Finally, the normalized measured feature vectors and multimodal fusion feature vectors were stored separately, and the verification results of the consistency between the normalization boundary parameters and the L2 norm were recorded.

[0050] S312. Calculate the cosine similarity between the normalized measured feature vector and each multimodal fused feature vector, and use it as the molecular structure feature similarity. Specifically, the normalized measured feature vectors and multimodal fused feature vectors are retrieved from the normalized data storage area. Before calculating the cosine similarity, all vectors involved in the calculation are pre-checked: the absorbance of the corresponding compound test solution for each vector is measured at 260 nm and 280 nm using ultraviolet-visible spectrophotometry, the A260 / A280 ratio is calculated, and vectors with ratios exceeding the range of 1.8-2.0 are removed to eliminate interference from nucleic acid or protein contamination on the accuracy of vector characterization; simultaneously, dynamic light scattering technology is used to measure the water content of the test solution. Vectors with a hydrated particle size exceeding 100 nm are removed to eliminate false positive signals caused by light scattering due to compound aggregation. Vectors that pass the pre-detection enter the formal calculation process, and the cosine similarity between the measured feature vector and each multimodal fusion feature vector is calculated one by one using the following formula: Cosine similarity = (A·B) / (||A||×||B||), where A is the normalized measured feature vector, B is the normalized multimodal fusion feature vector, A·B is the inner product of the two vectors, and ||A|| and ||B|| are the L2 norms of the two vectors, respectively.

[0051] The calculation process was set up with two independent calculations, with two experimenters writing and executing the calculation scripts independently. If the difference between the two calculation results exceeded 0.01, the original normalized data for that pair of vectors was retrieved again for manual verification until the difference between the two calculation results was within 0.01. After the similarity calculation of each pair of vectors was completed, surface plasmon resonance (SPR) technology was used simultaneously at low flow rate to determine the non-specific adsorption response values ​​of the compounds corresponding to the measured feature vectors and the target proteins. The calculated cosine similarity was correlated and compared with the SPR non-specific adsorption response values. If the cosine similarity was higher than 0.7 but the SPR non-specific adsorption response value also exceeded 50 RU, the similarity value was marked as "requires verification", indicating that the high similarity may be partly due to non-specific binding interference rather than true molecular structural similarity. After the cosine similarity calculation was completed, the similarity between each compound and each multimodal fusion feature vector in the database was sorted from high to low. The highest similarity value and its corresponding benchmark compound number and target association information were recorded. At the same time, the independent calculation records of the two experimenters and the SPR non-specific adsorption detection results were saved to form a molecular structure feature similarity matrix and archived.

[0052] S313. Based on the similarity of molecular structure features and the target pathway association information in the benchmark database, the multi-target efficacy prediction score of the compound samples is calculated by weighting, and compounds with scores higher than the threshold are selected as compounds with potential multi-target efficacy.

[0053] Specifically, before performing weighted calculations, for each compound in the generated molecular structure similarity matrix, the equilibrium dissociation constants of the highest similarity benchmark compounds with key targets of cerebral ischemia (TNF-α, Caspase-3, NR2B subunit, NMDA receptor, and P2X7 receptor) were determined one by one using surface plasmon resonance (SPR) technology. The measured values ​​were compared with the values ​​recorded in the database, and the similarity data corresponding to the benchmark compounds with deviations exceeding ±0.3 log units were removed to ensure that the similarity used for weighted calculations were based on experimentally verified binding activity data. At the same time, high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS) was used to determine the actual purity of each candidate compound. If the purity was lower than 95%, the compound was removed from the weighted calculation queue to avoid impurities interfering with efficacy prediction. The similarity values ​​of molecular structure features of each compound after verification (denoted as S, with a value range of 0-1) are weighted and fused with the target pathway coupling degree parameters (denoted as C, with a value range of 0-1, obtained by normalizing the node connectivity and cascade amplification coefficient of the target in the cerebral ischemia signaling pathway) retrieved from the drug target association detection benchmark database. The weighted calculation adopts the following formula: Multi-target efficacy prediction score = S×α + C×β, where α and β are weight coefficients, which are 0.6 and 0.4 respectively. The weight coefficients are set according to the optimal discrimination threshold determined by the subject operating characteristic curve analysis after the known positive multi-target compounds (such as curcumin and resveratrol) and known negative single-target compounds in the database have undergone the same weighted calculation process.

[0054] The calculation process involves two people independently programming and cross-validating the results. If the difference between the two calculations exceeds 0.05, the original data is traced back and the calculation is repeated. After calculation, the predicted scores of the multi-target efficacy of each compound are sorted from high to low, with a preset score threshold of 0.65. This threshold is determined based on the lower quartile of the known qualified compound score distribution in the database. Compounds with a predicted score ≥ 0.65 are selected as preliminary qualified compounds. For the preliminary qualified compounds, surface plasmon resonance (SPR) technology is used to measure their binding response values ​​with the five key targets of cerebral ischemia in batches. If a compound has a response value higher than 30 RU at three or more targets, it is officially registered as a potential qualified compound for multi-target efficacy. For boundary compounds with predicted scores between 0.60 and 0.64, SPR multi-target measurement screening is also performed. Boundary compounds with actual response values ​​exceeding 30 RU at three or more targets are added to the qualified list. Finally, a list of potential qualified compounds for multi-target efficacy is compiled, and the predicted score, weighted calculation parameters, SPR multi-target measurement response value, and inclusion basis of each compound are recorded and archived.

[0055] S32. Based on the target pathway experimental correlation information in the drug target correlation detection benchmark database, infer the target pathway coupling degree corresponding to the compound sample, and combine the molecular structure feature similarity with the target pathway coupling degree to screen out compounds with potential multi-target efficacy. Specifically, experimental association information of brain ischemia-related target pathways is retrieved from the drug target association detection benchmark database. This information includes the upstream and downstream connections of each target in the signaling pathway, the amplification coefficient of the phosphorylation cascade, and the functional weight data verified by gene knockout / overexpression experiments. Before calling the target pathway information, the activity verification of pathway node proteins stored in the database is performed: using surface plasmon resonance technology, each pathway node protein immobilized on the CM5 chip is used as a ligand, and a positive control compound with known high affinity is introduced as an analyte. The equilibrium dissociation constant is measured. If the measured value deviates from the value recorded in the database by more than ±0.3 log units, the node protein is re-identified by mass spectrometry and its activity is calibrated. The corresponding record in the database is updated before it can be called, ensuring that the binding activity of each node in the pathway association information is confirmed by experimental testing.

[0056] For each compound to be screened, firstly, based on the calculated molecular structure similarity, a benchmark compound with the highest similarity was identified. The target regulatory spectrum of this benchmark compound, verified by in vitro enzyme activity inhibition experiments and cell signaling pathway reporter gene experiments, was retrieved from the database. Based on this, the potential target set of the compounds to be screened was initially inferred. Subsequently, the purity of each candidate compound was verified using liquid chromatography-tandem mass spectrometry (LC-MS / MS), and compounds with a purity below 95% were eliminated. After confirmation, each compound was individually tested in an in vitro kinase inhibition experiment at a concentration of 10 μmol / L to determine its inhibition rate against each target in the inferred target set. The experiment used time-resolved fluorescence resonance energy transfer (TRFR) or homogeneous time-resolved fluorescence spectrometry (RTFS) for detection, and the inhibition rate data for each target were recorded. Targets with a measured inhibition rate ≥ 50% were identified as effective targets. Based on the pathway node connectivity and cascade amplification factor of the effective targets in the database, the target pathway coupling degree C was calculated using the formula C = Σ(connectivity of each effective target × cascade amplification factor) / theoretical maximum coupling value of the pathway, and the C value was normalized to the range of 0-1. The calculated target pathway coupling degree C and the calculated molecular structure feature similarity S are integrated according to the weighted fusion formula (efficiency score = S × 0.6 + C × 0.4). The weighting coefficient is set based on the optimal discrimination threshold determined by the analysis of the receiver operating characteristic curve after the known multi-target positive and negative compounds in the database have undergone the same calculation process. The calculation process adopts independent programming and cross-validation by two people. If the difference between the two calculation results exceeds 0.05, the original inhibition rate data is traced back to recalculate. The compounds were ranked from highest to lowest based on their overall efficacy scores, with a preset score threshold of 0.65. Compounds with scores ≥0.65 were selected as initial screening compounds. For compounds with scores ≥0.65, surface plasmon resonance (SPR) was used to determine their equilibrium dissociation constants with each effective target site. Compounds with equilibrium dissociation constants ≤1 μmol / L at three or more target sites were retained. For boundary compounds with scores between 0.60 and 0.64, simultaneous SPR multi-target site measurements were performed. Compounds that met the standard of equilibrium dissociation constants ≤1 μmol / L at three target sites were added to the list and finally registered as potential qualified compounds for multi-target efficacy.

[0057] S33. For the screened compounds, surface plasmon resonance technology is used to conduct in vitro multi-target binding experiments, determine the equilibrium dissociation constant of the compound with each target, determine the actual binding activity based on the equilibrium dissociation constant, and screen to obtain a sample set of candidate compounds.

[0058] Specifically, the selected compounds with potential multi-target efficacy were subjected to in vitro multi-target binding experiments using surface plasmon resonance one batch at a time. Before the experiment, the purity of each compound's stock solution was verified using high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS), and compounds with a purity below 95% were removed. At the same time, dynamic light scattering was used to determine the hydrated particle size of the compound in the running buffer (phosphate buffer containing 5% dimethyl sulfoxide), and compounds with a hydrated particle size exceeding 100 nm were removed to eliminate interference from aggregates on the binding signal. After the purity and particle size were found to be qualified, the actual concentration was determined by ultraviolet-visible spectrophotometry at the maximum absorption wavelength of the compound. Each compound was then serially diluted with the running buffer to at least six concentration points in the range of 0.1 nmol / L to 100 μmol / L. In terms of target protein preparation, key target proteins of cerebral ischemia (TNF-α, Caspase-3, NR2B subunit, NMDA receptor extracellular domain and P2X7 receptor extracellular domain) were immobilized on the CM5 sensor chip using amino coupling method, with the immobilization level controlled in the range of 800-1200 RU.

[0059] After fixation, a positive control compound with a known equilibrium dissociation constant was introduced for activity verification. A 1:1 Langmuir binding model was used to fit the sensor map. If the measured equilibrium dissociation constant of the positive control compound deviated from the known value by more than ±0.3 log units, the chip was re-prepared and the protein was re-fixed until the deviation was within ±0.3 log units before proceeding to the formal experiment. In the formal experiment, compounds of different concentrations were sequentially introduced into flow cells immobilized with different target proteins at a flow rate of 30 μL / min, with a binding time of 120 seconds and a dissociation time of 300 seconds. After dissociation, a 10 mmol / L glycine-hydrochloric acid solution was introduced for regeneration for 30 seconds. A blank cycle of running buffer was introduced before and after each concentration injection to subtract background signals in real time.

[0060] After the experiment, Biacore analysis software was used to perform dual-reference subtraction on the sensor maps. The blank flow cell signal was used as the reference channel, and the running buffer blank cycle was used as the reference cycle for subtraction. The steady-state response value was plotted against the compound concentration, and the equilibrium dissociation constant of each compound with each target was calculated using a steady-state affinity model. Simultaneously, at least 20% of the sensor map data were randomly selected, and two independent experimenters performed blind fitting analysis using different fitting software. If the deviation between the two equilibrium dissociation constants exceeded ±0.15 log units, the compound-target combination was re-measured. For each compound, its correlation with the target concentration was summarized. Compounds with equilibrium dissociation constants ≤1 μmol / L at three or more target sites are considered to have met the multi-target binding activity standard. For boundary compounds that meet the standard at only two target sites but have equilibrium dissociation constants ≤5 μmol / L at the remaining target sites, orthogonal verification is performed using isothermal titration calorimetry to determine the binding enthalpy change and stoichiometry of the compound with the boundary target site. If the isothermal titration calorimetry experiment confirms the existence of specific binding (absolute value of binding enthalpy change ≥5 kcal / mol and stoichiometry between 0.8 and 1.2), the boundary target site is included in the compliance count. Compounds meeting the compliance condition of three or more target sites are also retained. Compounds determined to meet the multi-target binding activity standard are registered as a candidate compound sample set.

[0061] S4. For the compounds in the candidate compound sample set, in vitro multi-target binding inhibition experiments were performed to determine the half-maximal inhibitory concentration (IC50). The cell protection rate was determined using an oxygen-glucose deprivation-reperfusion neural cell model. Liver microsomal metabolic stability, cardiomyocyte toxicity, gene mutation risk, and blood-brain barrier penetration were also tested. The experimental data were collected, standardized, and integrated to obtain the core feature parameter set. In this embodiment of the invention, for compounds in the candidate compound sample set, in vitro multi-target binding inhibition experiments were performed to determine the half-maximal inhibitory concentration (IC50). Cell protection rate was determined using an oxygen-glucose deprivation-reperfusion neural cell model. Liver microsomal metabolic stability, cardiomyocyte toxicity, gene mutation risk, and blood-brain barrier penetration verification experiments were also conducted. The data from these experiments were collected, standardized, and then integrated to obtain a core feature parameter set, including the following steps: S41. For the compounds in the candidate compound sample set, in vitro multi-target binding inhibition experiments were performed to detect the half-maximal inhibitory concentrations of the compounds on multiple targets related to cerebral ischemia. Cell activity was detected using an oxygen-glucose deprivation-reperfusion neural cell model to obtain neural cell protection rate data, which were then integrated to obtain multi-dimensional pharmacodynamic data. Specifically, for each compound in the candidate compound sample set, the purity was first verified using high performance liquid chromatography-tandem mass spectrometry, and compounds with a purity ≥95% were retained. At the same time, the concentration of each compound in the stock solution was determined by ultraviolet-visible spectrophotometry at the maximum absorption wavelength. The compounds were diluted to 10 μmol / L with a buffer solution containing 0.1% dimethyl sulfoxide, and dynamic light scattering technology was used to confirm that the hydrated particle size in the diluted solution was ≤100 nm to eliminate aggregate interference. In vitro multi-target binding inhibition assays were conducted targeting key targets of cerebral ischemia: TNF-α, Caspase-3, NR2B subunit, NMDA receptor, and P2X7 receptor. Homogeneous time-resolved fluorescence or time-resolved fluorescence resonance energy transfer was used for detection. Each target protein and fluorescently labeled probe were incubated for 60 minutes in buffer solutions containing or without a series of compound concentrations (0.1 nmol / L to 100 μmol / L, 8 concentration gradients). A solvent control group and a known positive inhibitor control group (such as TNF-α antibody, Caspase-3 inhibitor Z-DEVD-FMK, and fenfenadil) were included. Fluorescence signals were measured using a multi-mode microplate reader, and the inhibition rate at each concentration point was calculated. The half-maximal inhibitory concentration (IC50) was obtained by fitting the four-parameter logistic equation. 50 Each experiment should be conducted with three replicates, each independently repeated three times. The coefficient of variation between replicates should be ≤15%; otherwise, the experiment should be repeated.

[0062] Simultaneously, Z′ factor quality control was run on each target plate. A Z′ factor ≥ 0.5 was acceptable. The oxygen-glucose deprivation-reperfusion neural cell model used human neuroblastoma SH-SY5Y cells seeded in 96-well plates. After reaching 80% confluence, the medium was replaced with glucose-free DMEM and the cells were deprived for 3 hours in a 37℃, 95% N2 / 5% CO2 hypoxic incubator. Then, reoxygenation and glucose resuscitation were performed, followed by treatment with 0.1, 1, and 10 μmol / L candidate compounds for 24 hours. A positive control group (edaravone) and a solvent control group were set up. After treatment, each well was incubated with CCK-8 reagent for 2 hours, and the absorbance at 450 nm was measured using a microplate reader to calculate cell viability. The protection rate of each compound was calculated with the undeprived normal group as 100%. The experiment was performed in 6 replicates, independently three times. A protection rate ≥ 60% in the positive control group was required as the batch validity standard; otherwise, the entire plate was re-tested. All IC50 values ​​were collected. 50 After obtaining cell protection rate data, the results were normalized using the Z-score normalization method, and the multi-target IC50 values ​​were sorted by compound number. 50 Values, protection rates, and corresponding coefficients of variation are integrated into a multi-dimensional efficacy data matrix, and quality control parameters are recorded simultaneously, and then integrated to obtain multi-dimensional efficacy data.

[0063] S42. For the compounds in the candidate compound sample set, liver microsomal metabolic stability test, cardiomyocyte toxicity test and gene mutation risk test are performed respectively. The corresponding metabolic stability parameters, cardiotoxicity risk parameters and gene mutation risk parameters are obtained as toxicity data. At the same time, the blood-brain barrier penetration verification experimental data of the compounds are collected. Specifically, for each compound in the candidate compound sample set, high performance liquid chromatography-tandem mass spectrometry is first used to verify its purity, retaining compounds with a purity ≥95%. At the same time, dynamic light scattering technology is used to confirm that the hydrated particle size of the compound in each subsequent experimental buffer is ≤100nm. Only after eliminating aggregate interference can the compound proceed to the toxicity detection process.

[0064] For the determination of hepatic microsomal metabolic stability, human liver microsomal proteins and various compounds were co-incubated in potassium phosphate buffer containing an NADPH regeneration system. The final concentration of the compounds was 1 μmol / L, and the organic solvent content was controlled below 0.1%. After preheating the incubation system in a 37°C water bath for 5 minutes, NADPH was added to start the reaction. 100 μL samples were taken at 0, 5, 15, 30, and 60 minutes, and 200 μL of ice-cold acetonitrile containing internal standard was immediately added to terminate the reaction and precipitate the proteins. After vortexing for 3 minutes, the mixture was centrifuged at 14,000 rpm for 15 minutes at 4°C. The supernatant was collected, and the residual concentration of the compounds at each time point was determined by liquid chromatography-tandem mass spectrometry. The in vitro metabolic half-life and intrinsic clearance rate were calculated by fitting a first-order elimination kinetic equation with the zero-time-point concentration as a reference. Testosterone was set as a positive control in the experiment. The deviation between the measured value and the literature value of its metabolic half-life should not exceed ±20%, otherwise the entire batch should be repeated. Three replicates were set for each batch, and the coefficient of variation between replicates was ≤15%. Compounds with a metabolic half-life ≥ 60 minutes are labeled as having good metabolic stability, those with a half-life of 30-60 minutes are labeled as having moderate stability, and those with a half-life < 30 minutes are labeled as having unstable metabolism.

[0065] For the detection of cardiomyocyte toxicity, human ventricular myocytes were used in 96-well plates at a concentration of 1×10⁻⁶ cells / well. 4 Cells were cultured at a density of 80% confluence in each well. A series of concentrations (0.1, 1, 10, 30, 100 μmol / L) of candidate compounds were added for treatment for 48 hours. A solvent control group and a positive control group (doxorubicin) were included. After treatment, each well was incubated with CCK-8 reagent for 2 hours. The absorbance at 450 nm was measured using a microplate reader, and cell viability at each concentration point was calculated. The half-maximal toxicity concentration (IC50) of cardiomyocytes was obtained by fitting a four-parameter logistic equation. Each experiment was performed in six replicates, independently three times. A positive control group cell viability ≤30% was considered valid, and a Z′ factor ≥0.5 was acceptable.

[0066] For gene mutation risk detection, the AMES test pre-culture method was used under conditions without S9 metabolic activation: histidine-deficient Salmonella Typhimurium strains TA98, TA100, TA1535, TA1537 and tryptophan-deficient Escherichia coli WP2uvrA strains were pre-cultured in liquid medium with a series of concentrations (0.3, 1, 3, 10, 30, 100 μg / well) of candidate compounds, and then mixed with top agar and plated on minimum glucose agar plates. After incubation at 37°C for 48 hours, the number of revertant colonies was counted. A positive result was defined as a revertant colony count that was 3 times or more higher than that of the solvent control group and showed a dose-response relationship. At the same time, corresponding positive mutagen control groups and upper limits of concentration for no colony toxicity criteria were set.

[0067] In terms of blood-brain barrier penetration verification experiments, the apparent permeability coefficient of candidate compounds was re-determined independently, following the established in vitro Transwell blood-brain barrier model operation procedure. Three parallel Transwell chambers were used for the measurement, and the mean and relative standard deviation were calculated to ensure that the relative standard deviation was ≤20%.

[0068] The data on liver microsomal metabolic half-life, intrinsic clearance rate, metabolic stability grade, cardiomyocyte half-toxicity concentration, AMES test regression mutation rate and positive / negative determination results for each strain, and independent validation apparent penetration coefficient were summarized by compound number. The Z-score normalization method was used to normalize each toxicity indicator to form a toxicity-penetration dataset that includes metabolic stability parameters, cardiotoxicity risk parameters, gene mutation risk parameters, and blood-brain barrier penetration validation parameters.

[0069] S43. Standardize the multi-dimensional efficacy data, toxicity data, and blood-brain barrier penetration test data respectively, and integrate them according to the compound sample identifier to generate a core feature parameter set.

[0070] Specifically, before standardization, the completeness and quality of the acquired multi-dimensional efficacy data, toxicity data, and blood-brain barrier penetration verification experimental data were checked: high performance liquid chromatography-tandem mass spectrometry was used to verify the purity of the stock solutions of each compound in the candidate compound sample set, and all data records corresponding to compounds with a purity of less than 95% were removed; at the same time, dynamic light scattering technology was used to determine the hydrated particle size of each compound in the experimental buffer, and data corresponding to compounds with a hydrated particle size exceeding 100 nm were removed to eliminate deviations in efficacy or toxicity data caused by aggregate formation; for each batch of experimental data, the positive control indicators were checked one by one to see if they met the preset effective standards, and all data of batches with substandard positive controls were removed, with the reasons for invalidity noted.

[0071] After the data verification was passed, the multi-dimensional efficacy data were standardized. The half-maximal inhibitory concentration (IC50) of each compound against the five targets of TNF-α, Caspase-3, NR2B subunit, NMDA receptor, and P2X7 receptor was taken as the negative logarithm to base 10 to obtain the pIC50. 50 The cell protection rates at each concentration point in the oxygen-glucose deprivation-reperfusion neural cell model were standardized using the Z-score method. The standardized protection rate score was calculated by dividing the difference between the protection rate of each compound and the mean protection rate of the entire batch of compounds by the standard deviation. Toxicity data were also standardized by taking the natural logarithm of the liver microsomal metabolic half-life and taking the negative logarithm of the cardiomyocyte half-toxicity concentration (pTC) to obtain the pTC. 50 The ratio of the number of reversion mutations of each strain in the AMES test to the number of reversion mutations of the corresponding solvent control group was first calculated, and then normalized by taking the natural logarithm to form the mutation risk index; for the blood-brain barrier penetration verification experimental data, the apparent permeability coefficient was standardized by taking the negative logarithm to the base 10.

[0072] After standardization, the compounds are linked and integrated according to their unique identifiers: using the compound number as the primary key, the standardized multi-target pICs are linked together. 50 Value, standardized protection rate score, logarithmic metabolic half-life, pTC 50 Values, mutation risk indices, and standardized apparent permeability coefficients were aligned and merged column by column to generate a structured data matrix with one row for each compound and one column for each parameter. During the integration process, UV-Vis spectrophotometry was used to determine the absorbance ratio of each compound's test solution at 260 nm and 280 nm. If the A260 / A280 ratio exceeded the range of 1.8-2.0, the corresponding data row for that compound was marked "Test solution purity questionable," triggering a purity retest. Simultaneously, surface plasmon resonance technology was used at low flow rates to determine the non-specific binding response values ​​of each compound to serum albumin, and these response values ​​were added to the data matrix as supplementary feature dimensions. After integration, the data matrix was exported in comma-separated value format. Two researchers independently checked the original experimental records and standardized conversion calculations for each data point. If data mismatches or calculation errors were found, the original experimental data was traced back for correction and re-standardization. After verification, the data matrix was officially named the core feature parameter set.

[0073] S5. It presets efficacy threshold rules, toxicity rejection rules, and penetration lower limit rules, and constructs a multi-index comprehensive decision evaluation model to comprehensively evaluate the core feature parameter set and output the comprehensive efficacy rating, target adaptation results, and safety risk identification of each candidate compound.

[0074] In this embodiment of the invention, a multi-index comprehensive decision-making and evaluation model is constructed by pre-setting efficacy threshold rules, toxicity rejection rules, and penetration lower limit rules, and comprehensively evaluating the core feature parameter set to output the comprehensive efficacy rating, target adaptation results, and safety risk identification of each candidate compound, including the following steps: S51. Combining the multi-target half-maximal inhibitory concentration, neuronal protection rate, metabolic stability parameters, cardiotoxicity risk parameters, gene mutation risk parameters, and blood-brain barrier penetration verification data from the core feature parameter set, establish efficacy threshold rules, toxicity rejection rules, and penetration lower limit rules, and construct a multi-index evaluation matrix that includes efficacy indicators, toxicity indicators, and penetration indicators. Specifically, the efficacy threshold rule is set as follows: candidate compounds must simultaneously meet the following conditions, targeting at least three key targets: NMDA receptor, Caspase-3, and TNF-α at pIC. 50 All values ​​are not lower than 6.0, and the standardized protection rate score of neural cells is greater than 0; The toxicity rejection rule is set as follows: a candidate compound is directly deemed unsafe if it exhibits any of the following: pTC in cardiomyocytes. 50 A value below 4.0, or a mutation risk index above 1.5, or a logarithmic value of the liver microsomal metabolic half-life ln(T1 / 2) less than -2.30 (corresponding to a half-life less than 10 minutes). The lower limit rule for penetration is set as follows: when the blood-brain barrier standardized apparent permeability coefficient pPapp value of a candidate compound is lower than 6.0, it is judged that the penetration does not meet the standard and is eliminated. Construct a multi-index evaluation matrix: using the candidate compound number that has passed data validation and standardization as the row index, and the NMDA receptor pIC as the index. 50 Caspase-3pIC 50 TNF-αpIC 50 NR2B subunit pIC 50 P2X7 receptor pIC 50 Standardized protection rate score, logarithmic metabolic half-life, pTC 50 Ten indicators, including the value, mutation risk index, and pPapp value, are used as columns to construct an original evaluation matrix X with m rows and 10 columns, where m is the number of candidate compounds. Then, each column of matrix X is processed dimensionlessly using the range standardization method to map the values ​​of each indicator to the interval [0, 1], resulting in a standardized evaluation matrix X'.

[0075] S52. The multi-index evaluation matrix is ​​standardized, and a comprehensive scoring algorithm based on weighted distance is used to calculate the comprehensive performance score of each candidate compound. Then, based on the target equilibrium dissociation constant and the preset ideal solution of the target, the target fit is calculated. Specifically, the constructed multi-index evaluation matrix is ​​first standardized without dimensions. Since the dimensions and range of variation of each index are different, a vector normalization method is used to divide each element in the matrix by the square root of the sum of the squares of all elements in its column to obtain a normalized matrix. Then, within the framework of a weighted distance-based comprehensive scoring algorithm, the entropy weight method is used to objectively assign weights to the ten evaluation indicators. First, the proportion of each compound value under each indicator is calculated based on the normalized matrix. Then, the information entropy value of each indicator is calculated. Finally, the objective weight of each indicator is determined through the complementary relationship of information entropy, forming a weight vector containing ten weights. Then, the normalized matrix is ​​weighted using this weight vector, that is, each indicator value is multiplied by the corresponding weight to obtain a weighted normalized matrix.

[0076] Next, the positive and negative ideal solutions for each index are determined in the weighted standardization matrix. For the NMDA receptor pIC 50 Caspase-3pIC 50 TNF-αpIC 50 NR2B subunit pIC 50 P2X7 receptor pIC 50 Standardized protection rate score, pTC 50 For the eight benefit-related indicators, such as the value of the ideal value and the pPapp value, the larger the value, the better. For the positive ideal solution, the maximum value of the column is taken, and for the negative ideal solution, the minimum value is taken. For the two cost-related indicators, the logarithmic value of the metabolic half-life and the mutation risk index, the smaller the value, the better. For the positive ideal solution, the minimum value of the column is taken, and for the negative ideal solution, the maximum value is taken.

[0077] After obtaining the positive and negative ideal solutions, the weighted Euclidean distance to the positive ideal solution and the weighted Euclidean distance to the negative ideal solution are calculated for each candidate compound. The weighted Euclidean distance is calculated by squaring the difference between each index value of the compound and the corresponding index value of the ideal solution, multiplying by the weight, summing the results, and then taking the square root. Finally, according to the principle of the superior-inferior solution distance method, the comprehensive performance score of each compound is calculated, that is, the distance of the compound to the negative ideal solution is divided by the sum of its distances to the positive and negative ideal solutions, and the score value is between zero and one, with the closer to one indicating that the compound has better comprehensive performance.

[0078] In terms of target fit calculation, the equilibrium dissociation constants of candidate compounds and five cerebral ischemia-related targets, obtained by surface plasmon resonance technology, were extracted and uniformly converted into negative logarithms pK. D Value, a pre-defined ideal target spectrum, that is, the optimal binding strength vector expected for each target, such as the ideal pK of the NMDA receptor. DThe values ​​were 8.0 for Caspase-3, 7.5 for TNF-α, 7.5 for the NR2B subunit, and 7.0 for the P2X7 receptor, corresponding to high activity levels with equilibrium dissociation constants not exceeding ten nanomolars or tens of nanomolars, respectively. Then, the measured pK for each compound was calculated. D The cosine similarity between the vector and the ideal target spectrum vector is such that the closer the cosine value is to one, the better the multi-target binding characteristics of the compound match the ideal target spectrum.

[0079] S53. The toxicity rejection rule and the lower limit of penetration rule are used to screen candidate compounds with a single vote. For compounds that are not rejected, the comprehensive rating score and target fit ranking are output based on the comprehensive efficacy score and target fit. Specifically, a veto screening process is first implemented, retrieving toxicity data and blood-brain barrier penetration verification data for each candidate compound from the core characteristic parameter set, and comparing them with preset toxicity veto rules and penetration lower limit rules. The toxicity veto rule employs three parallel judgments: first, examining the compound's pTC in cardiomyocytes. 50 First, the compound's mutation risk index is checked against its hepatic microsomal metabolic half-life. If the value is below 4.0, it indicates a high risk of cardiotoxicity and is immediately rejected. Second, the compound's mutation risk index is checked against its hepatic microsomal metabolic half-life. If the value is above 1.5, it indicates a potential risk of gene mutation induction and is immediately rejected. Third, the compound's hepatic microsomal metabolic half-life is checked against its hepatic microsomal metabolic half-life. If the half-life is less than 10 minutes, it indicates that the compound will be cleared too quickly in vivo and cannot maintain an effective therapeutic concentration, and is also immediately rejected. For the lower limit of penetration rule, the compound's blood-brain barrier normalized apparent permeability coefficient (pPapp) is checked against its hepatic microsomal metabolic half-life. If the value is below 6.0, it indicates that the compound's ability to penetrate the blood-brain barrier and enter brain tissue is insufficient, making it difficult to achieve an effective exposure at the target site, and is immediately rejected. The above judgment process is performed compound by compound and rule by rule. Any compound that triggers any rejection rule is immediately removed from the candidate list, and the specific reason for rejection is recorded. Compounds that pass the screening without triggering any rejection rules will be designated as compounds to be rated and will proceed to the subsequent rating and ranking stage.

[0080] For compounds to be rated, a comprehensive rating and ranking are conducted based on the calculated overall efficacy score and target fit. The overall rating score is calculated by weighting the overall efficacy score and target fit according to preset weights. The preset weights are set based on the importance of the two indicators in drug screening decisions; for example, the overall efficacy score is weighted at 0.6, and the target fit is weighted at 0.4. This yields the overall rating score for each compound to be rated. The grading criteria for the overall efficacy score are as follows: compounds with an overall rating score of 0.80 or higher are rated as "Grade A," indicating excellent overall performance, possessing both potent efficacy and good target matching characteristics; compounds with an overall rating score between 0.60 and 0.80 are rated as "Grade B," indicating good overall performance and suitable for inclusion in the candidate pool; compounds with an overall rating score below 0.60 are rated as "Grade C," indicating average overall performance and recommended for reserve observation.

[0081] Target fit ranking is performed by directly sorting compounds in descending order of target fit affinity. Compounds with higher target fit affinity indicate that their multi-target binding characteristics are closer to the pre-defined ideal target spectrum, demonstrating superior performance in terms of target selectivity and affinity balance. After ranking, the overall rating of each compound and its target fit ranking are output one-to-one. Finally, the entire screening and rating process is reviewed to check whether the eliminated compounds truly meet the rejection rule trigger conditions and whether the scoring calculation process for the retained compounds is accurate. Once confirmed to be error-free, the overall rating score and target fit ranking results for the compounds to be rated are officially generated.

[0082] S54. Match the comprehensive rating score with the efficacy threshold rule to determine the comprehensive efficacy rating level, and review it in conjunction with the results of in vitro validation experiments to generate the target adaptation results and safety risk labels for each candidate compound.

[0083] Specifically, the first step is to match and verify the overall rating score of each compound against the efficacy threshold rule. The efficacy threshold rule requires that candidate compounds simultaneously meet the following criteria: pIC of at least two of the three key targets: NMDA receptor, Caspase-3, and TNF-α. 50 The value must be no less than 6.0, and the standardized protection rate of nerve cells must be greater than 0. After item-by-item comparison, compounds that meet the conditions will maintain their original rating level; those that fail to meet both conditions will have their overall efficacy rating level downgraded by one level, such as A to B, B to C, and the names of the targets that did not meet the standards and their corresponding values ​​will be marked in the results.

[0084] Subsequently, a three-tiered review was conducted based on the results of in vitro validation experiments. The first tier review, for Category A and B compounds, checked the goodness of fit of the dose-response curve in the multi-target binding inhibition experiment. If the goodness of fit was below 0.95, the rating was downgraded and the data fit was marked as questionable. The second tier review compared the apparent penetration coefficient of the Transwell blood-brain barrier model with the validation experimental data for consistency. Data with a deviation exceeding 30% between two independent measurements were marked as requiring further confirmation of penetration. The third tier review checked whether the protection rate of each batch of positive controls in the oxygen-glucose deprivation-reperfusion neural cell model was within three standard deviations of the historical mean. Data deviating from the historical mean were marked as indicating questionable batch quality control. The review results were independently verified by two people before the final rating was determined.

[0085] The target fit results were sorted according to the target fit fit score. A fit score above 0.85 was marked as a high fit, between 0.70 and 0.85 as a medium fit, and below 0.70 as a low fit. The measured pK values ​​for each compound were also listed. D Advantageous target combinations that achieve a value of 80% or more of the ideal solution threshold are recommended.

[0086] Safety risk labels are generated based on various toxicity test data. Myocardial cell pTC 50 Values ​​between 4.0 and 5.0 are marked as indicating cardiotoxicity requiring attention, while values ​​above 5.0 are marked as indicating low cardiotoxicity risk. A mutation risk index above 1.0 is marked as indicating mutagenicity requiring attention, while a value below 1.0 is marked as indicating low mutagenicity risk. A liver microsomal metabolic half-life shorter than 20 minutes is marked as metabolically unstable, 20 to 60 minutes as moderately metabolically stable, and longer than 60 minutes as good metabolically stable. A serum albumin nonspecific binding response value exceeding the standard is marked as high protein binding. The final output is a complete evaluation report including compound number, overall efficacy rating, target suitability level, list of advantageous targets, and various safety risk labels.

[0087] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for evaluating the activity of multi-target candidate compounds in cerebral ischemia based on in vitro detection, characterized in that, Includes the following steps: S1. Obtain the compound samples to be screened and preset the drug target association detection benchmark database; S2. Based on the drug target association detection benchmark database, the compounds to be screened are initially ranked according to their blood-brain barrier penetration. The compounds ranked first are then subjected to apparent permeability coefficient determination using the in vitro Transwell blood-brain barrier model. Compounds with unsatisfactory measured values ​​are removed. The remaining compounds are then subjected to molecular structure characterization to obtain measured feature vectors. S3. The measured feature vectors are compared with the drug target correlation detection benchmark database in multiple dimensions to obtain compounds with potential multi-target efficacy. For the initially screened compounds, the equilibrium dissociation constants of their interactions with multiple brain ischemia-related targets are determined one by one using surface plasmon resonance technology. Then, a sample set of candidate compounds is obtained based on the measured results and activity screening. S4. For the compounds in the candidate compound sample set, in vitro multi-target binding inhibition experiments were performed to determine the half-maximal inhibitory concentration (IC50). The cell protection rate was determined using an oxygen-glucose deprivation-reperfusion neural cell model. Liver microsomal metabolic stability, cardiomyocyte toxicity, gene mutation risk, and blood-brain barrier penetration were also tested. The experimental data were collected, standardized, and integrated to obtain the core feature parameter set. S5. It presets efficacy threshold rules, toxicity rejection rules, and penetration lower limit rules, and constructs a multi-index comprehensive decision evaluation model to comprehensively evaluate the core feature parameter set and output the comprehensive efficacy rating, target adaptation results, and safety risk identification of each candidate compound.

2. The method for evaluating the activity of multi-target candidate compounds for cerebral ischemia based on in vitro detection according to claim 1, characterized in that, The drug target association detection benchmark database is used to perform initial screening and ranking of compounds to be screened based on their blood-brain barrier penetration. The top-ranked compounds are then subjected to apparent permeability coefficient determination using an in vitro Transwell blood-brain barrier model. Compounds with substandard measured values ​​are removed. The remaining compounds are then subjected to molecular structure characterization to obtain measured feature vectors, including the following steps: S21. Input the molecular structure information of the compound sample into the pre-built blood-brain barrier penetration prediction model in the drug target association detection benchmark database, calculate the blood-brain barrier penetration prediction value of each compound and sort it. S22. For compounds that ranked high in the initial screening, the apparent permeability coefficient of each compound was determined using an in vitro Transwell blood-brain barrier model. S23. Remove compounds with an apparent permeability coefficient lower than a preset threshold to obtain the remaining compounds after secondary screening; S24. Perform molecular structure feature processing on the retained compounds after rescreening. Convert the SMILES string of the compound into a molecular graph structure containing atomic features and chemical bond features. Then, use a graph neural network to encode the molecular graph structure to obtain a preliminary molecular representation vector. Then, input the preliminary molecular representation vector into a deep learning representation network for feature extraction and dimension mapping, and output a fixed-dimensional measured feature vector.

3. The method for evaluating the activity of multi-target candidate compounds in cerebral ischemia based on in vitro detection according to claim 1, characterized in that, The process of comparing measured feature vectors with a drug target correlation detection benchmark database in multiple dimensions to obtain compounds with potential multi-target efficacy, and then using surface plasmon resonance technology to determine the equilibrium dissociation constants of the initially screened compounds with multiple cerebral ischemia-related targets, and finally obtaining a candidate compound sample set based on measured results and activity screening, includes the following steps: S31. Perform feature alignment matching between the measured feature vector and the pre-stored multimodal fusion feature vector in the drug target association detection benchmark database, and calculate the molecular structure feature similarity between the compound sample and the benchmark database sample. S32. Based on the target pathway experimental correlation information in the drug target correlation detection benchmark database, infer the target pathway coupling degree corresponding to the compound sample, and combine the molecular structure feature similarity with the target pathway coupling degree to screen out compounds with potential multi-target efficacy. S33. For the screened compounds, surface plasmon resonance technology is used to conduct in vitro multi-target binding experiments, determine the equilibrium dissociation constant of the compound with each target, determine the actual binding activity based on the equilibrium dissociation constant, and screen to obtain a sample set of candidate compounds.

4. The method for evaluating the activity of multi-target candidate compounds in cerebral ischemia based on in vitro detection according to claim 1, characterized in that, The compounds in the candidate compound sample set were subjected to in vitro multi-target binding inhibition experiments to determine the half-maximal inhibitory concentration (IC50). Cell protection rates were measured using an oxygen-glucose deprivation-reperfusion neural cell model. Liver microsomal metabolic stability, cardiomyocyte toxicity, gene mutation risk, and blood-brain barrier penetration were also tested. The collected experimental data, after standardization, were integrated to obtain the core feature parameter set, including the following steps: S41. For the compounds in the candidate compound sample set, in vitro multi-target binding inhibition experiments were performed to detect the half-maximal inhibitory concentrations of the compounds on multiple targets related to cerebral ischemia. Cell activity was detected using an oxygen-glucose deprivation-reperfusion neural cell model to obtain neural cell protection rate data, which were then integrated to obtain multi-dimensional pharmacodynamic data. S42. For the compounds in the candidate compound sample set, liver microsomal metabolic stability test, cardiomyocyte toxicity test and gene mutation risk test are performed respectively. The corresponding metabolic stability parameters, cardiotoxicity risk parameters and gene mutation risk parameters are obtained as toxicity data. At the same time, the blood-brain barrier penetration verification experimental data of the compounds are collected. S43. Standardize the multi-dimensional efficacy data, toxicity data, and blood-brain barrier penetration test data respectively, and integrate them according to the compound sample identifier to generate a core feature parameter set.

5. The method for evaluating the activity of multi-target candidate compounds in cerebral ischemia based on in vitro detection according to claim 1, characterized in that, The preset efficacy threshold rules, toxicity rejection rules, and penetration limit rules, along with the construction of a multi-index comprehensive decision-making evaluation model, comprehensively evaluate the core feature parameter set and output the comprehensive efficacy rating, target adaptation results, and safety risk identification of each candidate compound, including the following steps: S51. Combining the multi-target half-maximal inhibitory concentration, neuronal protection rate, metabolic stability parameters, cardiotoxicity risk parameters, gene mutation risk parameters, and blood-brain barrier penetration verification data from the core feature parameter set, establish efficacy threshold rules, toxicity rejection rules, and penetration lower limit rules, and construct a multi-index evaluation matrix that includes efficacy indicators, toxicity indicators, and penetration indicators. S52. The multi-index evaluation matrix is ​​standardized, and a comprehensive scoring algorithm based on weighted distance is used to calculate the comprehensive performance score of each candidate compound. Then, based on the target equilibrium dissociation constant and the preset ideal solution of the target, the target fit is calculated. S53. The toxicity rejection rule and the lower limit of penetration rule are used to screen candidate compounds with a single vote. For compounds that are not rejected, the comprehensive rating score and target fit ranking are output based on the comprehensive efficacy score and target fit. S54. Match the comprehensive rating score with the efficacy threshold rule to determine the comprehensive efficacy rating level, and review it in conjunction with the results of in vitro validation experiments to generate the target adaptation results and safety risk labels for each candidate compound.

6. The method for evaluating the activity of multi-target candidate compounds in cerebral ischemia based on in vitro detection according to claim 2, characterized in that, The step of inputting the molecular structure information of the compound sample into a pre-built blood-brain barrier penetration prediction model in a drug target association detection benchmark database, calculating and sorting the blood-brain barrier penetration prediction values ​​of each compound includes the following steps: S211. Using the blood-brain barrier penetration experimental data stored in the drug target association detection benchmark database as the training set, and training a blood-brain barrier penetration prediction model based on deep learning algorithm. S212. Extract the molecular fingerprint features of the compounds to be screened, and input the molecular structure features into the blood-brain barrier penetration prediction model to calculate the blood-brain barrier penetration prediction value of each compound. S213. Sort the blood-brain barrier penetration prediction values ​​from high to low, and select compounds whose prediction values ​​exceed the set threshold as the initial screening compounds.

7. The method for evaluating the activity of multi-target candidate compounds in cerebral ischemia based on in vitro detection according to claim 2, characterized in that, The determination of the apparent permeability coefficient of the compounds that ranked high in the initial screening using an in vitro Transwell blood-brain barrier model includes the following steps: S221. Brain microvascular endothelial cells were seeded in Transwell chambers and cultured until a tight monolayer was formed to establish an in vitro blood-brain barrier model. S222. Add the compounds that passed the initial screening to the upper chamber of the Transwell chamber and incubate them under culture conditions for a predetermined time. S223. Take samples from the lower chamber, determine the concentration of the compounds, and calculate the apparent permeability coefficient of each compound.

8. The method for evaluating the activity of multi-target candidate compounds in cerebral ischemia based on in vitro detection according to claim 2, characterized in that, The process of removing compounds with an apparent permeability coefficient lower than a preset threshold to obtain the remaining compounds after rescreening includes the following steps: S231. Compare the preset lower limit of penetration with the apparent permeability coefficient of the compound; S232. Based on the comparison results, compounds with apparent permeability coefficients below the lower limit threshold are removed, and the remaining compounds are recorded as rescreened retained compounds.

9. The method for evaluating the activity of multi-target candidate compounds for cerebral ischemia based on in vitro detection according to claim 2, characterized in that, The process of performing molecular structure characterization on the retained compounds after rescreening involves converting the SMILES string of the compound into a molecular graph structure containing atomic and chemical bond features, encoding the molecular graph structure using a graph neural network to obtain a preliminary molecular representation vector, and then inputting the preliminary molecular representation vector into a deep learning representation network for feature extraction and dimension mapping to output a fixed-dimensional measured feature vector. This process includes the following steps: S241. Perform molecular structure standardization preprocessing on the compounds retained after repeated screening, convert their SMILES strings into molecular graph structures with atoms as nodes and chemical bonds as edges, and assign chemical attribute features to the nodes and edges. S242. Input the molecular graph structure into a pre-trained graph neural network and encode it through message passing and neighborhood aggregation to obtain the initial molecular representation vector. S243. Input the preliminary molecular representation vector into the deep learning representation network for feature dimensionality reduction and mapping, and output a fixed-dimensional measured feature vector.

10. The method for evaluating the activity of multi-target candidate compounds for cerebral ischemia based on in vitro detection according to claim 3, characterized in that, The step of performing feature alignment matching between the measured feature vector and the pre-stored multimodal fusion feature vector in the drug target association detection benchmark database, and calculating the molecular structure feature similarity between the compound sample and the benchmark database sample includes the following steps: S311. Normalize the measured feature vector and the multimodal fusion feature vector pre-stored in the benchmark database to make them the same scale. S312. Calculate the cosine similarity between the normalized measured feature vector and each multimodal fused feature vector, and use it as the molecular structure feature similarity. S313. Based on the similarity of molecular structure features and the target pathway association information in the benchmark database, the multi-target efficacy prediction score of the compound samples is calculated by weighting, and compounds with scores higher than the threshold are selected as compounds with potential multi-target efficacy.