An organ chip and deep learning-based method for predicting the efficacy of antitumor compounds
By combining organ-on-a-chip and deep learning technologies, an organ-on-a-chip sample library and a wet experimental database were constructed, which solved the problems of insufficient accuracy and ethical issues in drug efficacy evaluation in existing technologies. This enabled high-throughput and automated efficacy prediction, improving the accuracy and clinical relevance of drug evaluation.
Patent Information
- Application Number
- CN202310708965.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-15
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2043-06-15
AI Technical Summary
In existing technologies, animal models and two-dimensional cell culture techniques have insufficient accuracy and ethical issues in drug efficacy evaluation, and machine learning lacks sufficient data support, making it impossible to effectively combine organ-on-a-chip technology for high-throughput drug efficacy prediction.
By combining organ-on-a-chip and deep learning technologies, an organ-on-a-chip sample library and a wet experimental database are constructed. Through feature stitching models and image preprocessing techniques, high-throughput and automated drug efficacy prediction is achieved, and multimodal data obtained from organ-on-a-chip are used to predict drug effects.
This technology enables high-throughput, automated drug efficacy prediction at the organ level in human subjects, reducing ethical risks, improving the accuracy and clinical relevance of prediction results, and saving experimental costs.
Smart Images

Figure CN116597916B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of computer and biological medicine cross science and technology, and particularly relates to a method for predicting the anti-tumor compound prognosis efficacy based on an organ chip and deep learning. BACKGROUND
[0002] The evaluation of drug efficacy is the most core link in the drug development process; at present, the traditional animal test model, two-dimensional cell culture technology and the technology combined with machine learning are still mainly used for the evaluation of drug efficacy. The animal test model has many inconsistencies between the results of drugs in humans and animals due to the differences in genome and physiology, such as metabolic rate, organ size and structure, immune system, metabolic pathway, etc. In recent years, the accuracy of animal models has been increasingly questioned, and the protection of animal rights has been increasingly valued. The necessity of animal models in the drug development pipeline has gradually decreased. Although the two-dimensional cell culture technology is based on cell lines or human primary cell culture, it lacks some important cell types existing in the real microenvironment in vivo, and lacks the necessary cell-cell cooperation and interaction at the organ level in vivo, especially in the study of tumor cells, which lacks the necessary consideration of tumor microenvironment interaction, resulting in that the two-dimensional cell culture model cannot completely simulate the development of diseases at the organ level in vivo, affecting the reliability of the results.
[0003] Compared with the two-dimensional cell culture which can only provide the expression information of single-layer cells and is difficult to simulate complex tissue structure and function, the organ chip and organoid technology can form effective three-dimensional structure cells, thereby presenting the expression characteristics closer to the organ level in human body, and can better simulate the cell types, cell-cell interaction, tissue structure and cell differentiation in the organ in vivo. For the prediction of drug effect, the metabolism, transport, toxicity and the like of the organ in vivo under the action of the drug can be restored as much as possible. Therefore, the use of organ chips or organoid chips for the research of the efficacy of candidate compounds in the drug development pipeline has important significance for obtaining more reliable test data and reducing the number of animal experiments.
[0004] Thanks to the enhancement of data processing and computing power, the technology in the field of machine learning has developed rapidly, and has been applied in the direction of drug development. Although the intervention of machine learning technology effectively improves the efficiency of drug screening, it is mainly combined with the two-dimensional cell culture technology with low throughput data for drug screening, and the machine learning and early deep learning model lack sufficient explainability and high-throughput biological and clinical data support. The low degree of openness of clinical and experimental data eventually leads to the limited data throughput available for machine learning, and the model accuracy is not as expected.
[0005] Compared with two-dimensional cell culture technology, the data type obtained by organ chip and organoid chip is more complex, including metabolomics data, activity data and time series image data of various modalities. At present, there is no effective data processing method for complex and diverse high-throughput data, especially high-content image data. Therefore, organ chip technology and deep learning technology cannot be combined for compound drug efficacy prediction and evaluation. SUMMARY
[0006] The purpose of the present application is to overcome the defects in the prior art and provide an organ chip and deep learning based anti-tumor compound prognosis drug efficacy prediction method, which combines organ chip technology and deep learning technology for anti-tumor drug efficacy prediction, realizes the integration of dry and wet experiments, and realizes high-throughput, automation, and avoidance of ethical risks while as close as possible to the human body organ level system drug effect.
[0007] To achieve the above purpose, the technical solutions adopted by the present application are as follows:
[0008] An organ chip and deep learning based anti-tumor compound prognosis drug efficacy prediction method, comprising the following steps:
[0009] Step 1, constructing an organ chip sample library
[0010] The organ chip sample library stores sample information of tumor organoids of different sources and corresponding clinical information thereof;
[0011] Step 2, constructing a wet experiment database
[0012] Select organ chip samples and at least one compound (derived from a treatment scheme) corresponding thereto for wet experiment from the organ chip sample library, determine the reference concentration of the compound in the wet experiment process according to the public information of the corresponding compound in the organ chip sample database, and set up a concentration gradient based on the reference concentration of the compound to process the organ chip, obtain wet experiment data results, and construct a wet experiment database; the wet experiment database stores wet experiment information, and the wet experiment information includes wet experiment basic information and wet experiment data results;
[0013] Step 3, data preprocessing, feature extraction and vectorization: classify the information in the organ chip database and the wet experiment information in the wet experiment database into text, numerical value and image, and divide them into text type data, numerical value type data and image type data; then, the text type data, numerical value type data and image type data are preprocessed, feature extracted and vectorized to convert them into a form recognizable by a computer;
[0014] Step 4, construction of feature splicing model;
[0015] The text type data features, numerical type data features and image type data features are combined using a feature splicing mechanism to construct a feature splicing model;
[0016] Step 5, training the feature splicing model, constructing a lethality effectiveness prediction model, and performing validation evaluation;
[0017] Step 6, training the feature splicing model, constructing an EC50 prediction model, and performing validation evaluation;
[0018] Step 7, using the constructed lethality effectiveness prediction model and EC50 prediction model to predict lethality effectiveness and EC50 value.
[0019] As a further technical feature, in step 1, the tumor organoids of different sources include one or more of tumor organoids extracted from primary clinical samples after desensitization treatment, organoids obtained from other sample banks, specific types of tumor organoids differentiated from stem cells, and organoids constructed from cell lines;
[0020] As a further technical feature, the wet experiment data results include one or more of omics detection data results, activity detection data results, and image data results; wherein the image data further includes bright field three-dimensional imaging data with time information and / or fluorescent staining image data with time information.
[0021] As a further technical feature, when performing wet experiments, the compound reference concentration is obtained by converting the drug concentration information in the organ chip sample database;
[0022] As a further technical feature, if the compound for wet experiment is a marketed drug, the blood drug peak data (Cmax) under the conventional dose can be obtained in the organ chip database, and the concentration is used as the reference concentration of the compound added in the culture medium during organ chip culture. The blood drug peak data (Cmax) under the conventional dose is the maximum drug concentration that can be reached by the compound in the plasma within the plasma protein binding or in the plasma during the drug action time period after taking the drug at the standard dose.
[0023] If the compound for wet experiment cannot obtain the blood drug peak data (Cmax) under the conventional dose from the organ chip database, the reference concentration of the compound needs to be calculated based on the conventional drug concentration recommended by its prescription;
[0024] As a further technical feature, in step 3, the text type data includes information data in the organ chip database and wet experiment basic information in the wet experiment database; the numerical type data includes one or several of the omics detection data results and the activity detection data results in the wet experiment data result information; and the image type data includes image data results in the wet experiment data result information.
[0025] As a further technical feature, the preprocessing, feature extraction and vectorization of the text type data specifically include: first, manually preprocessing the obvious abnormal input, missing and other unreasonable values in the text form information; then using a natural language tool to clean up the mood words, punctuation marks, and non-standard words in the text; then using scikit-learn to extract features, structuring the feature structure of the text information in the sample library, and processing discrete features through One-Hot encoding and processing regular frequency text type features through Word2Vec, to convert the text data into numerical vectors that can be processed by the model.
[0026] As a further technical feature, the preprocessing, feature extraction and vectorization of the numerical type data results include: first, using the Z-Score conversion of the z_score_scaler.fit_transform() function in the StandardScaler library to standardize the data, and then normalizing the numerical type features to vectorize them, wherein part of the feature data is subjected to range classification operation and converted into a grade classification label.
[0027] As a further technical feature, the preprocessing, feature extraction and vectorization of the image type data results include: first, image preprocessing of the image type data results, then constructing an organ image pre-training model, and using the decoder module of the constructed organ image pre-training model to extract image feature data from the preprocessed image data and vectorize them.
[0028] As a further technical feature, the image preprocessing of the image type data results includes one or more of image cropping and circle selection, image size adjustment, image rotation and flipping, image dark corner processing, image color adjustment, and standardized pixel value.
[0029] As a further technical feature, image cropping and circle selection: using the cv2.crop() function in the OpenCV library to crop unnecessary image areas and remove unnecessary areas or non-interesting areas that are not related to the task.
[0030] As a further technical feature, image size adjustment: using the cv2.resize() function of the OpenCV library to adjust the size of the original image, set to a unified target size, remove the differences in image collector settings under different experimental batches, or the image size differences introduced by the image cropping step;
[0031] As a further technical feature, image rotation and flipping: rotate and flip the image through the functions of cv2.getRotationMatrix2D(), cv2.warpAffine(), and cv2.flip(), to increase the robustness of the model;
[0032] As a further technical feature, image dark corner processing: use the background correction tool in ImageJ to make the brightness distribution of the image more uniform, and reduce the influence of the natural light source of the bright field image in the experimental operation;
[0033] As a further technical feature, image color adjustment: adjust the brightness and contrast of all images through the cv2.convertScaleAbs() function;
[0034] As a further technical feature, standardize the pixel value: map the pixel value of the image to a selected range through the cv2.normalize() function, and standardize the image pixel value;
[0035] As a further technical feature, feature extraction and vectorization of image data: construct an organoid image pre-training model, complete training and verification, and use the three-dimensional convolutional neural network in the encoding layer of the organoid image pre-training model to extract lower-dimensional feature representations through multi-layer convolution and pooling operations, realize feature extraction and vectorization, which specifically includes:
[0036] Step b-1, construction of organoid image pre-training model training data set
[0037] Take all image data after image type data preprocessing in step 3 as the initial training set, randomly mask the preprocessed image data in the initial training set, and delete the pixel vector data of the mask corresponding substructure in the initial training set, obtain the remaining image data after removing the mask part of the substructure, then construct the training data set by combining the remaining image data after removing the mask part of the substructure with the mask substructure pixel vector data, which is used for training the organoid image pre-training model;
[0038] Step b-2, construction of organoid image pre-training model
[0039] The organoid image pre-training model includes a decoder module and an encoder module;
[0040] First, the three-dimensional space vector image processed by the mask is subjected to convolution operation using a convolution kernel sliding in three dimensions, combined with a maximum pooling layer to filter the maximum feature value as the output, to reduce the spatial size of the feature map layer by layer, retain the main features, and perform 3D transpose convolution in the decoding layer to enlarge the spatial size to the original size before the pooling process; then the data before pooling is directly spliced with the data after deconvolution using a skip connection to transfer the low-level feature information to the decoder to assist in restoring details and position information, and finally, the same size image result as the input is output through the last convolution layer, thereby completing the design of the organoid image pre-training model.
[0041] Then, the substructure information after random mask processing is used as the target, the remaining image data after removing the mask part of the substructure in step b-1 is used as the input item, and the corresponding mask substructure pixel vector data is used as the output item. The designed organoid image pre-training model is trained, and the model training learning rate, batch, and iteration number are set as hyperparameters. The model optimizer is used for model iteration to obtain the final organoid image pre-training model.
[0042] Step b-3, using the encoder module in the constructed organoid image pre-training model to extract features from the preprocessed image data
[0043] The encoder module in the organoid image pre-training model is used to extract lower-dimensional feature representations of the preprocessed image data using multi-layer convolution and pooling operations in the three-dimensional image-based convolutional neural network.
[0044] As a further technical feature, the feature splicing mechanism includes one or more of Cross Attention cross attention mechanism, early connection fusion mechanism, hierarchical combination mechanism, and bottleneck fusion mechanism.
[0045] As a further technical feature, in step 5, the training and validation evaluation of the killing effectiveness model includes: determining the killing effectiveness of the compounds with different concentration gradients in the wet experiment database of step 2 according to the clinical data information in the organ chip sample database of step 1, and adding the judgment result of the killing effectiveness as a label to the wet experiment information of the corresponding wet experiment; then the data set of the wet experiment information with the added killing effectiveness label is divided into a killing effectiveness prediction model training set and a killing effectiveness prediction model validation set, the feature splicing model constructed in step 4 is trained and validated, and a killing effectiveness prediction model is obtained.
[0046] As a further technical feature, in step 5, the training and validation evaluation of the killing effectiveness model specifically includes:
[0047] Step 5-1, determine the killing effectiveness of the compounds in different concentration gradients in the step 2 wet experiment database according to the clinical data information in the step 1 organ chip sample database, and add the determination result of the killing effectiveness as a label to the wet experiment information of the corresponding wet experiment; the determination result of the killing effectiveness is positive or negative;
[0048] Step 5-2, divide the data set of the wet experiment information added with the killing effectiveness label into a killing effectiveness prediction model training set and a killing effectiveness prediction model validation set, take the features of the wet experiment information in the killing effectiveness prediction model training set as an input vector, take the features of the killing effectiveness label as an output vector, train the feature splicing model constructed in step 4, and use the optimization algorithm of stochastic gradient descent to set the hyperparameter learning rate, update the weight and bias parameters of the model after multiple iterations of feedback until the model converges, and form a preliminary killing effectiveness prediction model;
[0049] Step 5-3, input the features of the wet experiment information in the killing effectiveness prediction model validation set into the preliminary killing effectiveness prediction model as an input vector, predict the killing effectiveness result, and establish a confusion matrix by comparing the predicted killing effectiveness result with the label of the corresponding wet experiment information, and calculate the performance evaluation index parameters of accuracy, precision, recall, F-score and ROC curve to verify the performance of the model, and finally obtain the killing effectiveness prediction model.
[0050] As a further technical feature, in step 6, the training and validation evaluation of the EC50 prediction model includes: obtaining the EC50 value of the compound by using the data directly accumulated by the wet experiment or the data predicted by the effective killing model in step 5, and adding the obtained EC50 value to the wet experiment information corresponding to the wet experiment; then divide the wet experiment information added with the EC50 value corresponding to the data directly accumulated by the wet experiment into an EC50 prediction model training set and an EC50 prediction model validation set; and supplement the wet experiment information added with the EC50 value corresponding to the effective killing model prediction data to the EC50 prediction model training set, and finally obtain the EC50 prediction model after training and verification of the feature splicing model constructed in step 4.
[0051] As a further technical feature, in step 6, the training and validation evaluation of the EC50 prediction model includes:
[0052] Step 6-1, obtaining of EC50; there are two methods for obtaining of EC50 value, method one: using the data accumulated in wet experiment, regression analysis is carried out on the corresponding relationship between the concentration gradient of the compound and the cell activity detection data, and the EC50 value of the compound is calculated; method two: at least six concentration gradients of the compound are input into the killing effectiveness prediction model constructed in step 5 to predict the killing effectiveness, and then regression analysis is carried out on the concentration gradient and the label of the prediction output result, and the EC50 value of the compound is calculated;
[0053] Step 6-2, the EC50 value obtained in step 6-1 is added to the wet experiment information corresponding to the wet experiment;
[0054] Step 6-3, the wet experiment information corresponding to the data accumulated in the wet experiment and added with the EC50 value is divided into an EC50 prediction model training set and an EC50 prediction model verification set; and the wet experiment information added with the EC50 value corresponding to the effective killing model prediction data is supplemented to the EC50 prediction model training set, and the EC50 prediction model training set is supplemented; then, the features of the wet experiment information in the EC50 prediction model training set are used as input vectors, and the features of the labels corresponding to the wet experiment data are used as output vectors, the feature splicing model constructed in step 4 is trained, and the optimization algorithm of stochastic gradient descent is adopted, the learning rate is set as a hyperparameter, the weight and bias parameters of the model are updated after multiple iterations of feedback until the model converges, and a preliminary EC50 prediction model is formed;
[0055] Step 6-4, the features of the wet experiment data in the EC50 prediction model verification set are used as input vectors to input into the preliminary EC50 prediction model, and the EC50 value is predicted, and the EC50 value predicted by the preliminary EC50 prediction model is compared with the EC50 value in the corresponding label by using mean square error for performance evaluation, and the final EC50 prediction model is obtained after optimization of the model.
[0056] Compared with the prior art, the beneficial effects of the present application are that:
[0057] 1、The present application realizes the effective combination of dry and wet experiments by establishing an organ chip sample database and incorporating the clinical information of the compound into the database, and closely links the wet experiment with the clinical information, thereby improving the clinical relevance and accuracy of the prediction results.
[0058] 2. The application establishes a self-supervised pre-training task based on random image masks when constructing an organoid image pre-training model for image extraction. The data itself is fully based on the automatic and streamlined capture and pre-processing module of high-content organoid images, without the need for additional manual label annotation. Through random image masking, a self-supervised model training target is designed. This model training method can fully utilize all collected image data without affecting other image analysis purposes, without the need for additional human cost, and efficiently extracts image features through the organoid image pre-training model.
[0059] 3. The model architecture of the organoid image pre-training model based on high-content images uses a three-dimensional convolution kernel to slide in three directions and adds a ReLU nonlinear activation function to process the image through a three-dimensional convolution layer. After multiple three-dimensional convolution layers, a pooling layer is used to reduce the dimension, the image feature size is reduced and concentrated, and the image data dimension is gradually reduced through repeated iteration, thereby effectively extracting low-dimensional features of the image. Then, through transpose convolution, the image size is gradually restored, and the image before pooling in the same dimension is integrated in each level step of the transpose convolution through the way of skip connection, effectively improving the transmission of information in different levels. Finally, the same dimension output realizes the design and reasoning target of the organoid image pre-training model. This model architecture design can effectively extract low-dimensional features from the relatively sparse three-dimensional structure images of organoids and organ chips.
[0060] 4. The use of the organoid image pre-training model based on high-content images divides the design of the organoid image pre-training model into an encoding layer of downward convolution features and a decoding layer of upward transpose convolution. After the organoid image pre-training model completes the training of the downstream prediction task based on random masks, the performance of the model can be considered as the encoding layer effectively extracting low-dimensional features of the image data, and the decoding layer effectively amplifying the low-dimensional features and returning to the original image information dimension. When using the organoid image pre-training model, the downstream decoding layer of the upward transpose convolution is discarded, and only the encoding layer for data low-dimensional feature extraction is used. Such preprocessing effectively realizes the feature extraction of image data. Compared with using open-source general image large models or other methods to extract features from images, this method applies a special organoid image pre-training model for high-content images, effectively avoiding the situation that general large models in the field of computer vision are not targeted for organoid feature capture, and the low-dimensional features of open-source large models do not match the cell images.
[0061] 5、The application adopts multi-modal time sequence data to construct a prediction model, fully utilizes the high-throughput dynamic data capture characteristics of organ chip technology, fully utilizes the multi-omics multi-modal data capture instrument developed based on the organ chip and organoid platform, and comprehensively describes the state and characteristics of the organoid by using multi-dimensional data. By splicing and fusing the features extracted from the image data and the detection numerical data, a comprehensive feature vector is formed, the numpy.concatenate() function in the Python Numpy library can be used for feature fusion step, the image and numerical features of different data types are spliced in the same deep model, and the information of the multi-modal data is captured in parallel for inference and prediction, thereby further improving the inference performance of the model.
[0062] 6、The classification label is used as the output item in the killing effectiveness prediction model, which has higher stability and is suitable for decision-making scenarios with weak background experience.
[0063] 7、In the construction of the EC50 prediction model data set, based on the results of the killing effectiveness prediction model at each drug concentration, the EC50 is calculated by regression, and the output of the first model is used as part of the input data of the new model. Compared with the traditional wet experiment method to obtain the drug concentration, the data of the wet experiment is reduced, and the cost is saved.
[0064] 8、In the prediction of EC50, all experimental result data and wet experiment basic information data are used as input, and EC50 value is used as output to train the model, and the prediction result is further refined and specific. Compared with the original binary yes / no label application domain relying on drug public information, it is expanded to the real world wet experiment numerical data closely combined with three-dimensional organ phenotype level effect, and provides valuable prediction reference for the research of three-dimensional cell system such as organoid and organ chip.
[0065] 9、The application combines the output types of the killing effectiveness prediction model and the EC50 prediction model, so that the application has different use scenarios, different downstream inference targets can be used to build models according to the data capture resource situation and purpose of building the model, and the accuracy of the expected result, so that the established model, process and tool have stronger universality and are suitable for more actual scenes. BRIEF DESCRIPTION OF DRAWINGS
[0066] Figure 1 It is a drug hole image capture diagram of gefitinib plus drug under four times magnification observation in an embodiment of the application;
[0067] Figure 2 It is a principle diagram of random mask of organoid image pre-training model in an embodiment of the application;
[0068] Figure 3 An architecture diagram of an organoid image pre-training model in an embodiment of the present application;
[0069] Figure 4 An image before and after the substructure of the image is restored in the organoid image pre-training model in an embodiment of the present application;
[0070] Figure 5 A classification principle diagram of a Cross Attention cross attention mechanism in an embodiment of the present application;
[0071] Figure 6 A performance level diagram of a killing effectiveness model on a training data set in an embodiment of the present application;
[0072] Figure 7 A performance level diagram of a killing effectiveness model on a validation data set in an embodiment of the present application;
[0073] Figure 8 A regression principle diagram of a Cross Attention cross attention mechanism in an embodiment of the present application;
[0074] Figure 9 An EC50 regression curve diagram of gefitinib, lapatinib and erlotinib in an embodiment of the present application. DETAILED DESCRIPTION
[0075] The present application will be further described in detail below in combination with embodiments.
[0076] Embodiment 1:
[0077] A method for predicting the prognostic efficacy of an antitumor compound based on an organ chip and deep learning, comprising the following steps:
[0078] Step 1, constructing an organ chip sample library
[0079] The organ chip sample library stores sample information of tumor organoids of different sources and corresponding clinical information, which is derived from previously published open source information;
[0080] The different sources of tumor organoids include tumor organoids extracted from primary clinical samples after desensitization treatment, organoids obtained from other sample libraries, specific types of tumor organoids differentiated from stem cells, organoids constructed from cell lines, etc.;
[0081] The sample information includes sample and pretreatment information, sample culture information, sample drug information, chip information, etc.
[0082] The sample and its pre-processing information include sample library association information, sample collection time, sample morphology, sample quality, and pre-processing method.
[0083] The sample culture information includes, in time sequence, sample growth state, cell morphology, cell density, cell survival rate, passage number, culture medium name, batch number, formula, and type and concentration of additives, culture temperature condition, humidity condition, and CO2 concentration.
[0084] The sample drug addition information includes drug name, main drug component three-dimensional structure, batch number, source, and drug addition concentration information, drug prescription, and clinical effectiveness.
[0085] The organ chip information includes organ chip source, cutting method, size and shape, production batch, structural characteristics, and washing and sterilization pre-processing method.
[0086] The clinical information includes patient information after desensitization treatment, clinical detection index, treatment scheme, and sampling information, etc. The patient information includes age, gender, family history, previous medical history, previous treatment history, previous living habit, tumor diagnosis time, and diagnosis symptom information. The clinical detection index includes tumor type, stage, location, size, imaging examination result, pathological examination result, patient gene detection data, gene mutation data, and tumor marker index. The treatment scheme includes analysis report of the above clinical index, discussion conclusion of the multidisciplinary team, patient's willingness, participation in clinical trial and new drug experiment information. The sampling information includes sampling time, sampling specimen type (surgical tissue, blood drawing, saliva, biopsy, etc.), quantity, informed consent, and laboratory sample collection processing method.
[0087] In this embodiment, H1975 non-small cell lung cancer cell sample is selected from the organ chip sample library for subsequent operation. The original source of the H1975 non-small cell lung cancer cell sample data is the National Organ Chip Scientific Data Center-Organ Chip Database (OOCDB) (http:117.73.8.164:8083).
[0088] Step 2, constructing a wet experiment database
[0089] An organ chip sample and at least one compound (derived from a treatment scheme) corresponding to the organ chip sample are selected from the organ chip sample library for wet experiment. In the wet experiment process, the reference concentration of the compound in the wet experiment process is first determined according to the public information of the corresponding compound in the organ chip sample database, and then the organ chip is processed by using the concentration gradient based on the reference concentration of the compound to obtain the wet experiment data result and construct a wet experiment database. The wet experiment database stores wet experiment information, and the wet experiment information includes wet experiment basic information and wet experiment data result.
[0090] When performing wet experiments, the reference concentration of the compound is converted from the drug concentration information in the organ chip sample database;
[0091] If the compound for wet experiment is a marketed drug, the blood peak concentration data (Cmax) under the conventional dose can be obtained in the organ chip database, and the reference concentration of the compound added in the culture medium during the organ chip culture is used as the reference concentration of the compound added in the culture medium during the organ chip culture; the blood peak concentration data (Cmax) under the conventional dose is the maximum drug concentration that can be reached in the plasma after taking the drug at a standard dose and within the drug action time period.
[0092] If the compound for wet experiment cannot obtain the blood peak concentration data (Cmax) under the conventional dose from the organ chip database, the reference concentration of the compound needs to be converted based on the conventional drug concentration recommended by its prescription;
[0093] Taking an oral drug as an example, the blood concentration data corresponding to the compound is obtained by multiplying the standard dose of the drug by the bioavailability of the drug, and the reference drug concentration of the compound at the organ chip level is obtained by following the above method. The calculation formula of bioavailability is as follows:
[0094]
[0095] Wherein, D IV is the intravenous injection dose, which can be equivalent to the blood peak concentration data;
[0096] D po is the oral dose, i.e. the standard measurement of the compound prescription;
[0097] [AUC] IV is the area under the curve of intravenous injection dose, [AUC] po is the area under the curve of oral dose;
[0098] F is the oral bioavailability, and the main drug ingredients of the marketed oral compound will be disclosed.
[0099] The wet experiment basic information includes the type of compound, the type of organ chip, the drug concentration, the culture method of organ chip, etc. during the wet experiment;
[0100] The wet experiment data results include one or more of omics detection data results, activity detection data results, and image data results; wherein the image data includes bright field three-dimensional imaging data with time information and / or fluorescent staining image data with time information.
[0101] Among them, the fluorescence staining image data: the cells in the organ chip are fluorescently labeled with selected markers, and the fluorescence three-dimensional conformation of the organoid under the organ chip culture system is batch collected at fixed time intervals by setting an automatic collection program through a high-content imaging instrument with a confocal imaging module.
[0102] Activity detection data: the cells in the organ chip are detected for activity, and the method includes MTT, CCK8, ATP, etc. The optical density value of each sample is determined at a specific wavelength by an enzyme labeler or a spectrophotometer. If a standard curve is made, the corresponding record is made, and the relative activity rate / survival rate of the sample calculated by the optical density value of each sample and the standard curve is recorded.
[0103] Omics detection data: the excretion of the organ chip exosome or the excretion culture medium is collected and enriched by the organ chip microfluidic control system, separated and extracted by filtration chromatography method, and the detection of metabolomics biomarkers based on liquid chromatography, gas chromatography and ELISA. The detection indexes include but are not limited to ATP, lactic acid, choline, nucleotides, antioxidants, oxidative product metabolite indicators. After preliminary analysis of metabolomics data, the data is recorded in units of samples.
[0104] In this embodiment, the construction of the wet experiment database specifically includes:
[0105] Step one, H1975 non-small cell lung cancer cell line is frozen and recovered, and human tumor organoid culture medium (human lung cancer organoid culture kit-AVATARGET KLU0010001) is used for organoid culture and subculture. The human tumor organoid culture medium contains special optimized growth factors, nutrients, small molecule inhibitors, Matrigel glue and other active ingredients corresponding to lung cancer cells;
[0106] Step two, after at least three generations of stable growth state of organoids using lung cancer self-gravity drug sensitive organ chip culture, under the lung organ chip system to detect the stability of the morphology and growth state of organoids during culture. After 7 days of culture and stable state of organoids, take out 5*10^6 or more organoid samples for enzymolysis, centrifugation, resuspension, PBS washing and other conventional cell extraction steps for counting quality control, freeze in liquid nitrogen and keep at-80℃ environment, and complete whole exome sequencing (Whole Exome Sequencing, WES) within 2 days, confirm that the frozen and thawed organoid samples are consistent with the EGFR mutation in the priori knowledge of H1975; If the mutation map is inconsistent during the resuscitation and culture of the organ chip organoids, the sample is excluded and the subsequent experimental operation is not performed; The consistency of the remaining mutation map is evaluated by mutation frequency calculation, and all mutation consistency categories are summarized to confirm that the organoids after resuscitation and culture still maintain the genome and mutation typing of the original H1975 cell strain, and confirm the value and verifiability of the organ chip model;
[0107] Step three, based on the mutation information of organoids, according to the existing lung cancer corresponding clinical guideline first-line treatment drugs, respectively using EGFR-TKI, corresponding EGFR mutation target tyrosine kinase inhibitor, drug treatment to organoids; Specifically, the organ chip is divided into four experimental groups, respectively adding Gefitinib, Lapatinib, Erlotinib, and DMSO as a control; Based on the calculation formula of bioavailability:
[0108]
[0109] Where, D IV is the intravenous dose, which can be equivalent to the blood drug peak data;
[0110] D po is the oral dose, i.e. the standard measurement of compound prescription;
[0111] [AUC] IV is the area under the curve of intravenous dose, [AUC] po is the area under the curve of oral dose;
[0112] F is the oral bioavailability, and the main drug components of the marketed oral compound will be disclosed.
[0113] The standard dose of the drug is used orally, multiplied by the bioavailability of the drug, to obtain the blood concentration data (Cmax) of the compound, that is, the maximum drug concentration that the compound in the plasma can reach after taking the drug at a standard dose within the plasma protein binding or in the plasma during the drug action time period. The median concentration of the drug added to the culture medium is used as the reference concentration; Table 1 shows the drug concentration information of the compounds in this example;
[0114] Table 1
[0115]
[0116]
[0117] Based on the median concentration after conversion, the drug concentration gradient of the three compounds in this example is set as a queue of 8 lengths, specifically 0 μM, 0.01 μM, 0.03 μM, 0.1 μM, 0.3 μM, 1 μM, 3 μM, 10 μM. Each concentration is tested in three parallel wells as a biological repeat.
[0118] Step four, during the whole culture period, the organoid samples on the organ chip in each experimental group are imaged at fixed time intervals of 2 hours; specifically, the high-content imaging device is used to image each organ chip sample under bright field, and the height gradient midpoint of the automatic focusing height is used as the height gradient midpoint, and the gradient is expanded up and down the Z axis with a step of 10 μm. A total of 16 organoid sample images are sampled with the automatic focusing point as the midpoint. Figure 1 The drug hole image capture of gefitinib at different time nodes under four-fold magnification observation is shown;
[0119] Step five, at the end of the culture period, the cells in the organ chip are subjected to ATP-based cell activity detection. The Firefly luciferase ATP cell activity detection kit is used, and the multifunctional enzyme label instrument with Luminometer function is used for chemiluminescence detection. By comparing with the prepared standard curve, the relative activity rate of the optical density value and the standard curve is obtained, and then the corresponding cell activity data of the sample is further obtained;
[0120] Step 3, data preprocessing, feature extraction and vectorization: the information in the organ chip database and the wet experiment information in the wet experiment database are classified and categorized according to text, numerical value and image, and divided into text type data, numerical value type data and image type data; then the text type data, numerical value type data and image type data are preprocessed, feature extracted and vectorized, so as to be converted into a form recognizable by a computer;
[0121] The text type data includes information data in the organ chip database and basic information of wet experiments in the wet experiment database.
[0122] The numerical type data includes one or more of omics detection data results and activity detection data results in the wet experiment data result information; and the image type data includes image data results in the wet experiment data result information.
[0123] In the embodiment, the text of the H1975 non-small cell lung cancer cell sample information collected in step 1 is preprocessed by using a natural language tool NLTK (v3.8.0) to clean up adverbs, punctuation marks, and non-standard words in the text; and then scikit-learn (v1.2.2) is used for feature extraction to structure the features of the text information in the sample library.
[0124] The information in the organ chip database and the basic information of wet experiments in the wet experiment database are presented and stored in the form of text, so when the pre-processing, feature extraction and vectorization are performed, first, the unreasonable values such as obvious abnormal input and missing in the text form information are manually pre-processed; then the natural language tool such as NLTK is used to clean up the adverbs, punctuation marks, and non-standard words in the text; and then scikit-learn is used for feature extraction to structure the feature structure of the text information in the sample library, and convert the text data into a numerical vector that can be processed by the model. For example, the One-Hot encoding can be used to process the discrete features in the clinical data, each feature corresponds to a dimension, and the value is 0 or 1, so as to convert the discrete data into a binary vector; or for example, for text features, the natural language processing pre-training model of Bert or the word embedding method of Word2Vec is used to capture the relationship between words and grammar in the text, and convert the text data into a numerical vector that can be processed by the model.
[0125] In the embodiment, the text of the H1975 non-small cell lung cancer cell sample information collected in step 1 is preprocessed by using a natural language tool NLTK (v3.8.0) to clean up adverbs, punctuation marks, and non-standard words in the text; and then scikit-learn (v1.2.2) is used for feature extraction to structure the features of the text information in the sample library.
[0126] (2) Preprocessing, feature extraction and vectorization of numerical data results: numerical data results include omics test data results and activity test data results;
[0127] For numerical data results, first, Z-Score conversion of the z_score_scaler.fit_transform() function in the StandardScaler library is used for data standardization, and then normalization processing is performed to vectorize numerical features (such as encoding in the range of [0, 1]) and / or convert them into hierarchical classification labels (such as low, medium and high hierarchical classification labels).
[0128] In this embodiment, Principal component analysis (PCA) linear dimension reduction is used for dimension reduction processing of numerical data. According to the principle of principal component analysis, the principal components are sorted according to the proportion of variance in the data from high to low. The principal components with high ranking represent that they contain more differences in the data, so the data in the first 50 principal components are retained to effectively reduce the model calculation complexity, filter out the redundant data with low correlation and refine the core feature signals, and improve the efficiency and robustness of the deep learning model training.
[0129] (3) Preprocessing, feature extraction and vectorization of image data results:
[0130] Image data results are first preprocessed, and then an organoid image pre-training model is constructed, and the constructed organoid image pre-training model is used to extract image feature data from the preprocessed image data and vectorize them.
[0131] a. Image preprocessing of image data results includes image cropping and circle selection, image size adjustment, image rotation and flipping, image dark corner processing, image color adjustment, standardized pixel value, etc.
[0132] Among them, image cropping and circle selection: for images of different experimental batches, use functions such as cv2.crop() function in OpenCV library to crop unnecessary image areas, remove unnecessary areas or non-task-related non-interest areas;
[0133] Image size adjustment: use functions such as cv2.resize() function in OpenCV library to adjust the size of the original image, set a uniform target size, remove differences in image collector settings under different experimental batches, or image size differences introduced by the image cropping step, increase the consistency of the input image of the model, and facilitate subsequent feature extraction;
[0134] Image rotation and flipping: Rotate and flip the original image by functions such as cv2.getRotationMatrix2D(), cv2.warpAffine(), cv2.flip(), etc. to increase the diversity of image data, and to improve the robustness of subsequent data-based models in the case of uneven sample labels.
[0135] Image dark corner processing: Process the brightness difference between the center and edge of the image due to different natural light conditions. Use the background correction tool in ImageJ to make the brightness distribution of the image more uniform.
[0136] Image color adjustment: Adjust the brightness and contrast of all images using functions such as cv2.convertScaleAbs(), and adjust in different directions according to the performance of the model to increase the heterogeneity of image data and improve the diversity of data sources, or reduce the differences between images to improve consistency.
[0137] Standardized pixel value: Map the pixel value of the image to a selected range using functions such as cv2.normalize() to standardize the image pixel value and reduce the impact of extreme values in the pixel value range on model performance.
[0138] b. Feature extraction and vectorization of image data: After constructing and training the organoid image pre-training model, use the three-dimensional convolutional neural network in the encoding layer of the organoid image pre-training model to extract lower-dimensional feature representations through multi-layer convolution and pooling operations, achieving feature extraction and vectorization. Specifically, it includes:
[0139] Step b-1, training data set construction
[0140] Take all image data after image class data preprocessing in step 3 as the initial training set, randomly mask the preprocessed image data in the initial training set, and delete the pixel vector data of the masked substructure in the initial training set. Remove the remaining image data of the masked substructure, then combine the remaining training data set with the masked substructure pixel vector data to form a training data set for training the organoid image pre-training model.
[0141] Step b-2, construction of organoid image pre-training model
[0142] The organoid image pre-training model includes a decoder module and an encoder module.
[0143] First, the three-dimensional space vector image processed by the mask is subjected to convolution operation using a convolution kernel sliding in three dimensions, combined with a maximum pooling layer to filter the maximum feature value as the output, to reduce the spatial size of the feature map layer by layer, retain the main features, and perform 3D transpose convolution in the decoding layer to enlarge the spatial size to the original size before the pooling process; then the data before pooling is directly spliced with the data after deconvolution using a skip connection to transfer the low-level feature information to the decoder to assist in restoring details and position information, finally, the same size image result as the input is output through the last convolution layer, thereby completing the design of the organoid image pre-training model.
[0144] Then, as shown in Figure 2 , the remaining image data of the substructure after removing the mask part in step b-1 is taken as the input item, and the corresponding mask substructure pixel vector data is taken as the output item, the designed organoid image pre-training model is trained, and the model training learning rate, batch and iteration number and other hyperparameters are set, the model optimizer such as Adam and stochastic gradient descent is used for model iteration, and the final organoid image pre-training model is obtained.
[0145] Specifically, the same hole Z-axis 16 image views are treated as three-dimensional array processing, and for a three-dimensional array with a total pixel size of [1280, 1024, 16], a substructure with a pixel size of [16, 16, 4] is used for cutting, and a total of [80*64*4] substructures are generated in different dimensions. As shown by the dashed line of the substructure in Figure 1 , 20% of the substructures are randomly masked using a random number generator, i.e. the vector data contained therein is saved as the true label, while the vector data contained therein is deleted in the training data set, and the label of this structure is modified and marked as a three-dimensional space vector to be learned. At this point, the random mask processing is complete, and the task of the image feature extraction model is considered as a self-supervised task of predicting the real vector data contained in the masked substructure vector based on the three-dimensional data structure that is not masked.
[0146] Step b-3, using the encoder module in the constructed organoid image pre-training model to extract features from the preprocessed image data
[0147] Using the encoder module in the organoid image pre-training model, the preprocessed image data is subjected to multi-layer convolution and pooling operation based on three-dimensional image convolutional neural network to extract lower-dimensional feature representation and vectorization of the image.
[0148] In this embodiment, based on the fact that the organoid image pre-training model can capture low-dimensional features in three-dimensional cell image structures, the organoid image pre-training model is decomposed, and only the encoding module of the organoid image pre-training model is used. Multi-layer convolution and pooling operations in a three-dimensional image-based convolutional neural network are then used to extract lower-dimensional feature representations from the preprocessed image data. Specifically, such as... Figure 3 As shown in the encoder module on the left side of the pre-trained organoid image model, the Z-axis segmented image is first processed as a 3D image. Convolution operations are performed by sliding convolution kernels along three dimensions. Then, a max pooling layer is used to select the largest feature value as the output, thereby reducing the spatial size of the feature map layer by layer, preserving the main features, and achieving the reduction of image dimensionality and the extraction of low-dimensional features. Through three pooling layers, the data feature size is finally reduced to a small-dimensional feature of [160, 128, 256]. In the decoding layer, 3D transposed convolution is performed to enlarge the spatial size and restore it to the original size before pooling. Then, skip connections are used to directly concatenate the data before pooling with the data after deconvolution, passing low-level feature information to the decoder to help restore details and positional information. Finally, the last convolutional layer outputs an image of the same size as the input, namely [1280, 1024, 16].
[0149] This model architecture was used to train the established training dataset with a learning rate of 0.0015, a batch size of 4, and 200 epochs. Stochastic gradient descent was used as the optimizer to iterate the model and complete the substructure information after random masking. Figure 4 An inference example of the decoding module for downstream tasks in an organoid imaging pre-trained model is demonstrated. Figure 4 The left side of the image shows the true image of the Z=13 section of the masked [40,32,1] substructure in the cell image, while the right side shows the image generated by the organoid imaging pre-trained model based on the predicted Z=13 interface of the masked substructure. Since the downstream training target of the organoid imaging pre-trained model, namely the decoding layer, is not actually used in subsequent workflows, the performance evaluation of the organoid imaging pre-trained model can be based on the subjective reconstruction of the masked substructure. Figure 4 As can be seen, the organoid image pre-training model can effectively restore the masked substructure by capturing low-dimensional spatial information and extracting information by associating the information of adjacent structures of the masked substructure.
[0150] Step 4: Construction of the feature splicing model: Use the feature splicing mechanism to combine textual data features, numerical data features and image data features to construct the feature splicing model;
[0151] In the present embodiment, since the information in the organ chip database, the wet experiment basic information in the wet experiment database belong to text type data; the activity detection data and the omics detection data in the wet experiment database belong to numerical type data, and the image type result data in the wet experiment database belong to image type data, therefore, in the construction process of the feature splicing model, the present embodiment adopts the Cross Attention cross attention mechanism to combine the text type data features, the numerical type data features and the image type data features, and constructs the feature splicing model;
[0152] The Cross Attention cross attention mechanism is used for feature splicing and construction of two modal feature vectors. This dual-flow architecture performs cross-attention mechanism-based hidden layer feature association on the features extracted from (1) image data and (2) experimental detection values and other text data. The numerical and text vectors are used as Q for attention weighting of the image data vectors, and vice versa. The image vectors are used as Q for weighting of the text and numerical vectors, so as to realize cross-modal perception interaction in the respective constructed decoding layers. After output from the decoding layer, all features enter the pooling layer together, and the maximum value pooling method is used, and then the Softmax activation function layer is entered, to output the probability of binary classification or multi-classification result.
[0153] Step 5, training the feature splicing model, constructing the killing effectiveness prediction model, and performing verification and evaluation
[0154] According to the clinical data information in the organ chip sample database in step 1, the killing effectiveness of the compounds with different concentration gradients in the wet experiment database in step 2 is determined, and the judgment result of the killing effectiveness is added to the wet experiment information of the corresponding wet experiment as a label; then the data set of the wet experiment information with the added killing effectiveness label is divided into a killing effectiveness prediction model training set and a killing effectiveness prediction model verification set, the feature splicing model constructed in step 4 is trained and verified, and a killing effectiveness prediction model is obtained; specifically including:
[0155] Step 5-1, according to the clinical data information in the organ chip sample database in step 1, the killing effectiveness of the compounds with different concentration gradients in the wet experiment database in step 2 is determined, and the judgment result of the killing effectiveness is added to the wet experiment information of the corresponding wet experiment as a label; the judgment result of the killing effectiveness is positive or negative (i.e. effective or ineffective);
[0156] When the killing effectiveness is determined according to the clinical data information, the data generated in the experimental batch reaching the drug concentration corresponding to the conventional drug dose after conversion by the pre-step is a positive label; in the experimental batch not reaching the converted concentration, and in the case of compounds not within the approved indications of the drug, it is a negative label; the judgment results of the killing effectiveness of some compounds in the present embodiment are shown in Table 2.
[0157] Table 2
[0158]
[0159] Step 5-2, divide the data set of wet experiment information added with lethality effectiveness label into a lethality effectiveness prediction model training set and a lethality effectiveness prediction model validation set, take the features of the wet experiment information in the lethality effectiveness prediction model training set as an input vector, take the features of the lethality effectiveness label as an output vector, train the feature splicing model constructed in step 4, and use the random gradient descent optimization algorithm, set the hyperparameter learning rate, update the weight and bias parameters of the model after multiple iterations of feedback until the model converges, and form a preliminary lethality effectiveness prediction model;
[0160] As shown in Figure 5 In this embodiment, Cross Attention cross attention mechanism is used for feature splicing construction of two modal feature vectors. This dual-flow architecture associates the hidden layer features based on the cross attention mechanism of the features extracted from the image data and the experimental detection values and other text data, uses the numerical and text vectors as Q for attention weighting of the image data vectors, and vice versa uses the image vectors as Q for weighting of the text and numerical vectors, to realize cross-modal perception interaction in the respective constructed decoding layers. After output from the decoding layer, all features enter the pooling layer together, use the maximum value pooling method, and then enter the Softmax activation function layer, output the probability that the lung cancer organoids formed by the specific compound at this dose on the EGFR+ non-small cell lung cancer cell strain will produce effective killing; through this probability and the threshold selected by humans, the classification label of effective killing or ineffective killing can be obtained; set the learning rate = 0.001, batch = 2 batches, perform Epoch = 500 times, and use the random gradient descent as the optimizer to iterate the model;
[0161] The performance level of the trained preliminary lethality effectiveness prediction model is shown in Figure 6 The performance of the preliminary lethality effectiveness prediction model is evaluated by the area under the curve, and the preliminary lethality effectiveness prediction model reaches a performance level of AUC = 0.765 under the condition that the threshold is variable. According to the optimal threshold selection method of Youden Index in the following formula, 0.465 is selected as the probability threshold of the classifier, and the performance of the model is Specificity = 0.633, Sensitivity = 0.810;
[0162]
[0163] Step 5-3, input the features of the wet experiment information in the validation set of the lethality prediction model into the preliminary lethality prediction model as an input vector, predict the lethality result, and establish a confusion matrix by comparing the predicted lethality result with the label of the corresponding wet experiment information, and calculate the accuracy, precision, recall, F-score and ROC curve performance evaluation index parameters to verify the performance of the model, and finally obtain the lethality prediction model.
[0164] wherein the F-score comprehensively considers the precision and recall, and the formula is:
[0165]
[0166] wherein β represents the weight of the trade-off between the recall and the precision, and is the weight multiple of the recall to the precision. For the same weight, the F1 score can be used to evaluate the performance, that is, the recall and the precision are considered equally important, and the formula of the F1 score is:
[0167]
[0168] In this embodiment, Figure 7 The performance of the lethality prediction model validation set in the model is F1Score = 2 * (0.857 * 0.731) / (0.857 + 0.731) = 0.789, and the lethality prediction model has a prediction ability;
[0169] Step 6, train the feature splicing model, build an EC50 prediction model, and perform verification and evaluation, including: obtaining the EC50 value of the compound by using the data directly accumulated by the wet experiment or the data predicted by the effective lethality model, and adding the obtained EC50 value to the wet experiment information corresponding to the wet experiment; then dividing the wet experiment information corresponding to the wet experiment directly accumulated data and added with the EC50 value into an EC50 prediction model training set and an EC50 prediction model validation set; and supplementing the wet experiment information corresponding to the effective lethality model prediction data and added with the EC50 value to the EC50 prediction model training set, finally obtaining the EC50 prediction model after training and verifying the feature splicing model built in step 4;
[0170] Step 6-1, obtaining EC50; there are two methods to obtain the EC50 value:
[0171] Method one: through the data directly accumulated by the organ chip wet experiment, the concentration gradient of the compound is regressed with the corresponding relationship of the cell activity detection data to calculate the EC50 value of the compound; method two: through the killing effectiveness prediction model constructed in step 5, at least 6 concentrations of the compound are input to predict the killing effectiveness, and then the concentration gradient and the label of the prediction output result are regressed to calculate the EC50 value of the compound.
[0172] Compared with the traditional two-dimensional cell level EC50, a large number of concentration gradient experiments are needed to draw the dose and reaction curve to infer the EC50. Combined with the previous killing quantification model and deep learning technology, the EC50 of the compound can be predicted according to a small amount (more than twice) of wet experiment data of the unknown compound at specific concentration doses on the organ chip platform.
[0173] Step 6-2, adding the EC50 value obtained in step 6-1 to the wet experiment result data in the wet experiment database in step 2 as a label;
[0174] Step 6-3, data set division: the wet experiment information corresponding to the data directly accumulated by the wet experiment and added with the EC50 value is divided into an EC50 prediction model training set and an EC50 prediction model verification set; and the wet experiment information added with the EC50 value corresponding to the effective killing model prediction data is supplemented to the EC50 prediction model training set, and the EC50 prediction model training set is supplemented; then, the features of the wet experiment data in the EC50 prediction model training set are taken as the input vector, and the features of the label corresponding to the wet experiment data are taken as the output vector, the feature splicing model constructed in step 4 is trained, and the optimization algorithm of stochastic gradient descent (Stochastic Gradient Descent, SGD) is adopted, the learning rate is set as the hyperparameter, the weight and bias parameters of the model are updated after multiple iterations of feedback until the model converges, and a preliminary EC50 prediction model is formed;
[0175] As Figure 8As shown, in this embodiment, the Cross Attention mechanism is used for feature splicing and construction of two modal feature vectors. This dual-flow architecture associates the features extracted from image data and experimental detection values and other text data based on the cross-attention mechanism in the hidden layer. The numerical and text vectors are used as Q for image data vector attention weighting, and vice versa. The image vector is used as Q for text and numerical vector weighting to achieve cross-modal perception interaction in the respective decoding layer. After output from the decoding layer, all features are jointly input into the pooling layer using the maximum value pooling method, and then into the final fully connected layer. The identity linear activation function is used in the fully connected layer without introducing nonlinear transformation. Finally, a single-dimensional continuous numerical value is output after processing the input multi-dimensional features. This numerical output is the regression result. The output is not dose-limited, and the specific compound can produce active inhibition effect on EGFR+ non-small cell lung cancer cell lines to form lung cancer organoids. When exactly half of the cells produce active inhibition effect, the value of the concentration of the specific compound is output. Through the half maximal effective concentration value, the anti-tumor effect of the compound can be predicted and evaluated. The learning rate is set to 0.001, the batch is 2 batches, the Epoch is 500 times, and the stochastic gradient descent is used as the optimizer to iterate the model;
[0176] Step 6-4, the features of the wet experiment information in the EC50 prediction model verification set are input into the preliminary EC50 prediction model as input vectors, the EC50 value is predicted, and the mean squared error (MSE) is used to compare the EC50 value predicted by the preliminary EC50 prediction model with the EC50 value in the corresponding label and evaluate the performance, and finally obtain the EC50 prediction model.
[0177] The EC50 values in the EC50 prediction model verification set are obtained through wet experiment data;
[0178] In this embodiment, as shown in Figure 9 The cell activity ratios of the three compounds under different concentration gradients in the experimental results of the sample units in the EC50 prediction model verification set are applied in the regression analysis, and the wet experiment IC50 data of the three compounds are obtained as {0.135 μM, 0.096 μM, 0.937 μM};
[0179] When evaluating the performance of the regression model, the comparison of the model predicted EC50 and the true label value and the performance evaluation are evaluated by the mean squared error (MSE), as shown in the following formula. The smaller the value of MSE, the smaller the average error between the predicted results of the model and the true value, and the better the performance of the model;
[0180]
[0181] With reference to the performance of the model obtained in the model verification data set in this example, the MSE of the model is 0.1437, and the model has prediction ability;
[0182] Step 7, using the constructed killing effectiveness prediction model and EC50 prediction model to predict the killing effectiveness and EC50 value.
[0183] The input of the killing effectiveness prediction model is the characteristics of the wet experimental data, and the output is the predicted killing effectiveness judgment result;
[0184] The input of the EC50 prediction model is the characteristics of the wet experimental data, and the output is the predicted EC50 value.
[0185] The above-described embodiments are only preferred embodiments of the present application, and are not exhaustive of the feasible implementations of the present application. Any obvious modifications made by those skilled in the art without departing from the principles and spirits of the present application should be considered to be included in the protection scope of the claims of the present application.
Claims
1. A method for predicting the prognostic efficacy of antitumor compounds based on organ-on-a-chip and deep learning, characterized in that, include: Step 1: Construct an organ-on-a-chip sample library The organ-on-a-chip sample library stores sample information of tumor organoids from different sources and their corresponding clinical information; Step 2: Construct a wet experiment database Organ-on-a-chip samples for wet experiments and at least one corresponding compound are selected from the organ-on-a-chip sample library. The information of the corresponding compound in the organ-on-a-chip sample library determines the baseline concentration of the compound during the wet experiment. Based on the baseline concentration of the compound, a concentration gradient is established to process the organ-on-a-chip, obtain wet experiment data results, and construct a wet experiment database. The wet experiment database stores wet experiment information, which includes basic wet experiment information and wet experiment data results. In step 2, the wet experimental data results include one or more of the following: omics detection data results, activity detection data results, and image data results; wherein, the image data includes bright-field three-dimensional imaging data with time information and / or fluorescence staining image data with time information. Image preprocessing, feature extraction, and vectorization are performed on the image data results. Feature extraction and vectorization of image data includes: constructing an organoid image pre-training model, and after completing training and validation, using multi-layer convolution and pooling operations of the three-dimensional convolutional neural network in the coding layer of the organoid image pre-training model to extract lower-dimensional feature representations of the image, thereby realizing feature extraction and vectorization; Step 3: Data preprocessing, feature extraction, and vectorization: The information in the organ-on-a-chip database and the wet experiment information in the wet experiment database are classified into text, numerical, and image data. Then, the text data, numerical data, and image data are preprocessed, feature extracted, and vectorized respectively to transform them into a computer-recognizable form. In step 3, the text data includes information data from the organ-on-a-chip database and basic wet experiment information from the wet experiment database; Step 4: Construction of the feature splicing model; A feature concatenation mechanism is used to combine textual data features, numerical data features, and image data features to construct a feature concatenation model; the feature concatenation mechanism adopts the Cross Attention mechanism. Step 5: Train the feature splicing model to build a kill effectiveness prediction model and perform validation and evaluation; Step 6: Train the feature splicing model to build the EC50 prediction model and perform validation and evaluation; Step 7: Use the constructed kill effectiveness prediction model and EC50 prediction model to predict kill effectiveness and EC50 value.
2. The method for predicting the prognostic efficacy of antitumor compounds based on organ-on-a-chip and deep learning according to claim 1, characterized in that, In step 1, the tumor organoids from different sources include one or more of the following: tumor organoids extracted from primary clinical samples after desensitization treatment, tumor organoids obtained from other sample libraries, tumor organoids of specific types induced from stem cells, and tumor organoids constructed from cell lines.
3. The method for predicting the prognostic efficacy of antitumor compounds based on organ-on-a-chip and deep learning according to claim 1, characterized in that, In step 3, numerical data includes one or more of the omics detection data results and activity detection data results in the wet experimental data results information; image data includes the image data results in the wet experimental data results information.
4. The method for predicting the prognostic efficacy of antitumor compounds based on organ-on-a-chip and deep learning according to claim 1, characterized in that, The preprocessing, feature extraction, and vectorization of textual data specifically include: first, manually preprocessing unreasonable values in the textual information, including obvious abnormal inputs and missing values; then, using natural language tools to clean the text of modifiers, punctuation marks, and non-standard words; next, using scikit-learn for feature extraction, structuring the textual information features, processing discrete features through One-Hot encoding, and processing regular frequency textual features through Word2Vec, transforming the textual data into numerical vectors that the model can process; The preprocessing, feature extraction, and vectorization of numerical data results include: firstly, data standardization is performed using the Z-Score transformation of the fit_transform() function in the StandardScaler library, and then the numerical features are vectorized after normalization; some feature data undergo range classification operations and are converted into hierarchical classification labels.
5. The method for predicting the prognostic efficacy of antitumor compounds based on organ-on-a-chip and deep learning according to claim 1, characterized in that, Image preprocessing for image-type data results includes one or more of the following: image cropping and selection, image resizing, image rotation and flipping, image vignetting, image color adjustment, and normalization of pixel values. Feature extraction and vectorization of image data, specifically including: Step b-1: Construction of the training dataset for the organoid image pre-training model Using all the image data after image data preprocessing in step 3 as the initial training set, the preprocessed image data in the initial training set is randomly masked, and the pixel vector data of the substructure corresponding to the mask is deleted from the initial training set to obtain the remaining image data after removing the masked substructure. Then, the remaining image data after removing the masked substructure and the pixel vector data of the masked substructure are combined to form a training dataset for training the organoid image pre-training model. Step b-2: Construction of organoid imaging pre-training model First, the masked 3D spatial vector image is subjected to convolution operations by sliding the convolution kernel across three dimensions. Combined with max pooling layers to select the maximum feature value as the output, the spatial size of the feature map is reduced layer by layer while retaining the main features. Then, 3D transposed convolution is performed in the decoding layer to enlarge the spatial size and restore it to the original size before pooling. Next, skip connections are used to directly concatenate the data before pooling with the data after deconvolution, passing low-level feature information to the decoder to help restore details and positional information. Finally, the last convolutional layer outputs an image with the same size as the input, thus completing the design of the organoid image pre-training model. Then, with the goal of completing the substructure information after random masking, the remaining image data of the substructure after removing the mask in step b-1 is used as the input and the corresponding mask substructure pixel vector data is used as the output. The designed organoid image pre-training model is trained, and the hyperparameters of model training learning rate, batch and iteration number are set. The model optimizer is used to iterate the model to obtain the final organoid image pre-training model. Step b-3: Use the encoder module in the constructed organoid image pre-training model to extract features from the preprocessed image data; Using the encoder module in the organoid image pre-training model, multi-layer convolution and pooling operations in a 3D image-based convolutional neural network are used to extract lower-dimensional feature representations of the images and quantize them.
6. The method for predicting the prognostic efficacy of antitumor compounds based on organ-on-a-chip and deep learning according to claim 1, characterized in that, In step 5, the training and validation evaluation of the lethality model includes: determining the lethality of compounds of different concentration gradients in the wet experiment database in step 2 based on the clinical data information in the organ-on-a-chip sample database in step 1, and adding the judgment results of lethality as labels to the wet experiment information of the corresponding wet experiment; then dividing the dataset of wet experiment information with lethality labels into a lethality prediction model training set and a lethality prediction model validation set, training and validating the feature splicing model constructed in step 4, and obtaining the lethality prediction model.
7. The method for predicting the prognostic efficacy of antitumor compounds based on organ-on-a-chip and deep learning according to claim 6, characterized in that, Step 5, the training and validation evaluation of the lethality model, specifically includes: Step 5-1: Determine the killing effectiveness of compounds of different concentration gradients in the wet test database in Step 2 based on the clinical data information in the organ-on-a-chip sample database in Step 1, and add the judgment result of killing effectiveness as a tag to the wet test information of the corresponding wet test; the judgment result of killing effectiveness is positive or negative. Step 5-2: Divide the dataset of wet experiment information with added kill effectiveness labels into a kill effectiveness prediction model training set and a kill effectiveness prediction model validation set. Use the features of the wet experiment information in the kill effectiveness prediction model training set as the input vector and the features of the kill effectiveness labels as the output vector. Train the feature concatenation model constructed in Step 4 and use the stochastic gradient descent optimization algorithm. Set the hyperparameter learning rate and update the model's weights and bias parameters after multiple iterations and feedback until the model converges, forming a preliminary kill effectiveness prediction model. Step 5-3: Input the features of the wet experiment information in the validation set of the kill effectiveness prediction model into the preliminary kill effectiveness prediction model as input vectors to predict the kill effectiveness results. Then, establish a confusion matrix by comparing the predicted kill effectiveness results with the labels of the corresponding wet experiment information, and calculate the performance evaluation index parameters such as accuracy, precision, recall, F score and ROC curve to verify the model performance, and finally obtain the kill effectiveness prediction model.
8. The method for predicting the prognostic efficacy of antitumor compounds based on organ-on-a-chip and deep learning according to claim 1, characterized in that, Step 6, the training and validation evaluation of the EC50 prediction model, includes: obtaining the EC50 value of the compound using data directly accumulated from wet experiments or data predicted by the effective kill model in step 5, and adding the obtained EC50 value to the corresponding wet experiment information; then dividing the wet experiment information corresponding to the directly accumulated wet experiment data with added EC50 values into an EC50 prediction model training set and an EC50 prediction model validation set; and supplementing the wet experiment information corresponding to the effective kill model prediction data with added EC50 values into the EC50 prediction model training set. Finally, after training and validating the feature splicing model constructed in step 4, the EC50 prediction model is obtained.
9. The method for predicting the prognostic efficacy of antitumor compounds based on organ-on-a-chip and deep learning according to claim 8, characterized in that, Step 6, the training and validation evaluation of the EC50 prediction model, specifically includes: Step 6-1, Obtaining EC50; There are two methods to obtain the EC50 value. Method 1: Using the data accumulated from wet experiments, perform regression analysis on the correspondence between the compound concentration gradient and cell activity detection data to calculate the EC50 value of the compound. Method 2: Using the killing effectiveness prediction model constructed in Step 5, use at least 6 concentration gradients of the compound as input conditions to predict the killing effectiveness, and then perform regression analysis on the concentration gradient and the label of the prediction output to calculate the EC50 value of the compound. Step 6-2: Add the EC50 value obtained in Step 6-1 as a tag to the wet test information of the corresponding wet test; Step 6-3: Divide the wet experiment information corresponding to the directly accumulated data of the wet experiment with EC50 values into an EC50 prediction model training set and an EC50 prediction model validation set; supplement the EC50 prediction model training set with wet experiment information corresponding to the effective kill model prediction data with EC50 values; then, using the features of the wet experiment information in the EC50 prediction model training set as the input vector and the features of the labels corresponding to the wet experiment data as the output vector, train the feature concatenation model constructed in Step 4, and use the stochastic gradient descent optimization algorithm, set the hyperparameter learning rate, and update the model's weights and bias parameters after multiple iterations until the model converges, forming a preliminary EC50 prediction model. Step 6-4: Use the features of the wet experimental data in the EC50 prediction model validation set as input vectors to input into the preliminary EC50 prediction model to predict the EC50 value. Then, use the mean square error to compare and evaluate the EC50 value predicted by the preliminary EC50 prediction model with the EC50 value in the corresponding label. After optimizing the model, the final EC50 prediction model is obtained.
Citation Information
Patent Citations
System for deducing cancer risk probability by using multi-modal risk factors
CN113539493A
Performing pharmacodynamics evaluations using microfluidic devices
US20220074924A1